A method for predicting intracavitary dissemination from lung cancer histopathological images

By applying the image analysis method of twin hybrid encoder in the histopathological image analysis of lung cancer, the problem of high subjectivity and misdiagnosis rate of STAS diagnosis in the prior art is solved, and more accurate and reliable prediction of intra-air cavity dissemination is achieved.

CN118507062BActive Publication Date: 2025-05-16HUNAN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410584578.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-12
Publication Date
2025-05-16
Estimated Expiration
2044-05-12

AI Technical Summary

Technical Problem

The prior art has problems such as strong subjectivity, time-consuming and high misdiagnosis rate when diagnosing lung cancer intra-air dissemination (STAS), especially in the diagnosis of frozen sections with low diagnostic accuracy and repeatability.

Method used

The image analysis method based on twin hybrid encoder is adopted to predict the intra-air cavity dissemination by pre-processing, feature extraction, spatial topology map construction and feature extraction of lung cancer histopathological images.

Benefits of technology

This improves the diagnostic accuracy and repeatability of STAS, reduces the false positive rate, and realizes the ability to more accurately predict intra-air cavity dissemination from lung cancer histopathological images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118507062B_ABST
    Figure CN118507062B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for predicting intracavitary dissemination from lung cancer tissue pathology images, comprising: S1, image preprocessing of whole slide images; S2, feature extraction of preprocessed images based on pretrained models; S3, construction of a spatial topology map based on the features of the slide images; S4, feature extraction of the spatial topology map based on a twin hybrid encoder; S5, prediction of lung cancer tissue pathology images using the features extracted in step S4 to obtain intracavitary dissemination results. The present invention learns the features of the spatial topological structure of tissue pathology images through a twin graph encoder, and infers the training model through feature sharing and jump connections to achieve prediction of intracavitary dissemination from lung cancer tissue pathology images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence and biomedical engineering technology, and more specifically, to a method for predicting intracavitary dissemination from lung cancer tissue pathology images. Background Art

[0002] The phenomenon of lung cancer intra-air space dissemination was first discovered and described in 2013. Onozato et al. found through three-dimensional reconstruction technology that there are connected tumor cell islands outside the boundaries of the main lung cancer lesions. The phenomenon of these tumor cell islands occupying the air space of the lung parenchyma is closely related to the recurrence of the patient's tumor. In lung adenocarcinoma, Kadota and colleagues reported a pathological phenomenon of tumor STAS, which is defined as the spread of lung tumor cells through the air space into the lung parenchyma adjacent to the tumor. Subsequently, the World Health Organization's classification of lung adenocarcinoma in 2015 described intra-air space dissemination as a pattern of tumor invasion. Multiple global research groups have published data on more than 3,500 patients and found that STAS exists in 15% to 69% of patients with lung adenocarcinoma, and is strongly associated with recurrence and poor survival. STAS consists of microclusters, solid nests, or single cells outside the tumor that enter the air space of the surrounding lung parenchyma.

[0003] Histopathology slides are currently the gold standard for STAS diagnosis. However, WSI has super-large pixels, and manual evaluation of STAS by pathologists is a highly subjective, time-consuming, labor-intensive process prone to missed or misdiagnosis, and is not suitable for large-scale histopathological diagnosis. The diagnosis of intraoperative frozen sections (FS) can assist clinicians in making decisions during surgery, including guiding doctors to adjust the scope and method of surgery during surgery, timely evaluating the surgical margin, and providing an important reference for the designation of subsequent treatment plans. According to clinical studies, lobectomy may have better survival outcomes than sublobectomy in patients with STAS-positive T1 stage. At present, the accuracy and repeatability of FS in intraoperative diagnosis of STAS are still relatively low, and the false positive rate is high, which makes AI controversial in predicting frozen sections. Paraffin sections are also one of the basic means of pathological examination. They can be used to study the pathogenesis, pathophysiological processes, molecular biological characteristics, etc. of the disease, providing a scientific basis for the prevention, treatment and control of the disease. Secondly, through the observation and analysis of paraffin sections, doctors can provide patients with appropriate treatment plans, such as surgery, radiotherapy, chemotherapy, etc.

[0004] Benefiting from the rapid development of artificial intelligence, deep learning has been widely used in image analysis tasks, providing state-of-the-art performance for various image classification tasks. In computational pathology, image analysis based on deep learning has shown comparable performance to pathologists in various tasks. For example, semi-supervised training based on deep learning classifies prostate tissue pathology images and generalizes to completely different data sets. The classification accuracy of liver cancer tissue pathology images based on deep learning reached 96%, which is close to the ability of pathologists with 5 years of experience. Recent studies have shown that the Transformer based on the attention mechanism can achieve global feature extraction, and can also enable the features of each position to interact with the features of other positions, helping the model to better capture the relationship between different regions in the image, thereby improving the classification performance. For example, the multi-instance learning network based on the Swim Transformer can accurately classify colorectal cancer WSI using the label of WSI. The average accuracy of the breast cancer tissue pathology image classification model based on color deconvolution and Transformer architecture reached 93.02%. However, compared with traditional methods based on convolutional neural networks, the effect of Transformer on local feature extraction may be poor and the computational complexity is relatively high.

[0005] Given that most STAS are distributed next to the main tumor, it is necessary to construct a spatial topological map of WSI to assist in the diagnosis of STAS. The spatial topological map takes an independent patch as a point, uses the edge between two points to construct the relationship between patches, and integrates the global structural information into the structure of the graph, so that the model can comprehensively consider the characteristics of WSI, not just the local area of ​​​​WSI. Secondly, the spatial topological map can be used to visualize the spatial relationship and structural characteristics of different regions in WSI, making the classification results of the model more interpretable and understandable. Here, the diagnosis of STAS can be predicted by image analysis based on the twin graph encoder. Both frozen sections and paraffin sections were included in the internal validation set and the external validation set for testing. Summary of the invention

[0006] The purpose of the present invention is to provide a method for predicting intracavitary dissemination from lung cancer tissue pathology images, so as to overcome the defects of the prior art.

[0007] In order to achieve the above object, the technical solution adopted by the present invention is as follows:

[0008] A method for predicting intracavitary dissemination from lung cancer tissue pathology images comprises the following steps:

[0009] S1. Preprocessing the whole slide image;

[0010] S2, extracting features from the preprocessed image based on the pre-trained model;

[0011] S3, constructing a spatial topological map based on the features of the slide image;

[0012] S4, feature extraction of spatial topology map based on twin hybrid encoder;

[0013] S5. Use the features extracted in step S4 to predict the lung cancer tissue pathology image to obtain the result of intracavitary dissemination.

[0014] Furthermore, the step S1 specifically includes: digitizing the slide image, segmenting the tissue area, and detecting the background and blurred area.

[0015] Furthermore, the step S2 specifically includes: using the pre-trained models of KimiaNet and CTransPath to extract feature representations with dimensions of 1024 and 768 for each patch.

[0016] Furthermore, the KimiaNet pre-trained model is a pre-trained model with 7 million weights, which is fine-tuned and trained based on the DenseNet framework. The local connection and sliding window operation of the convolutional layer of the KimiaNet pre-trained model enable the network to capture local features; the CTransPatht pre-trained model is based on the Swin Transformer architecture, which consists of three convolutional layers and four Swin Transformer modules. The convolutional layers in CTransPath can extract local features in image data, and Swin Transformer can capture global features and long-distance dependencies.

[0017] Furthermore, in step S3, all patches in the slide image are represented in the form of feature vectors, and the K-nearest neighbor algorithm is used to construct a spatial topological graph G, G = {v, e}, v represents the node set of the graph, e represents the edge set, A is the adjacency matrix, and X is the feature of the node.

[0018] Furthermore, in step S4, the twin hybrid encoder is composed of two symmetrical encoders, and the output end of the twin hybrid encoder is the features output by the two encoders, which are concatenated and averaged and then output. Each encoder includes a GCNConv, a SAGEConv, two ReLU layers, a Dropout operation, an MLP network and a Rescale layer. GCNConv initializes the convolution kernel weight matrix W, multiplies the node feature X by the weight matrix W to obtain the linear transformation result Z=XW of the node, and uses the adjacency matrix A to propagate the node features; SAGEConv uses the neighbor node features of G to update the feature representation of the node.

[0019] Furthermore, the use of the adjacency matrix A to propagate the features of the nodes specifically includes:

[0020] For node i, the weighted sum of the features of its neighboring nodes is added to the node's own features to obtain a new feature representation. The process is:

[0021]

[0022] In the formula, H (l-1) Represents the feature representation of the l-1th layer node is the adjacency matrix A plus the self-connectivity matrix, is a diagonal matrix, where σ is the ReLU activation function;

[0023] Furthermore, the SAGEConv uses the neighbor node features of G to update the feature representation of the node, specifically including:

[0024] For each node v, select its neighbor node set N(v), aggregate its neighbor node set using mean pooling, and aggregate the features of node v It is concatenated with the aggregated neighbor node features, and then goes through a linear transformation and a nonlinear activation function to obtain the updated representation H of node v. (l+1) , the process can be expressed as:

[0025]

[0026] In the formula, W is the weight matrix, Concat concatenates the node's own features with the features of neighboring nodes, Aggregate is an aggregation operation, and the features of the Graph are subjected to a ReLU layer and a Dropout operation, and then an MLP network and a Rescale layer for feature extraction and dimensionality reduction.

[0027] Compared with the prior art, the advantages of the present invention are: the present invention learns the features of the spatial topological structure of tissue pathology images through a twin graph encoder, infers the training model through feature sharing and jump connections, and realizes the prediction of intracavitary dissemination from lung cancer tissue pathology images. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0029] Figure 1 The flowchart of the method for predicting intracavitary dissemination from lung cancer tissue pathology images of the present invention.

[0030] Figure 2 This is the experimental result of the present invention predicting STAS after five-fold cross-validation training.

[0031] Figure 3 It is the interpretability analysis of the model of the present invention. DETAILED DESCRIPTION

[0032] The preferred embodiments of the present invention are described in detail below in conjunction with the accompanying drawings so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making a clearer and more definite definition of the protection scope of the present invention.

[0033] See also Figure 1 As shown, this embodiment discloses a method for predicting intracavitary dissemination from lung cancer tissue pathology images, comprising the following steps:

[0034] Step S1, performing image preprocessing on the whole slide image.

[0035] Step S2: extract features from the preprocessed image based on the pretrained model.

[0036] Step S3: constructing a spatial topology map based on the features of the slide image.

[0037] Step S4: extract features of the spatial topology map based on the twin hybrid encoder.

[0038] Step S5: predicting the lung cancer tissue pathology image using the features extracted in step S4 to obtain the result of intracavitary dissemination.

[0039] In this embodiment, the step S1 specifically includes: digitizing the slide image, segmenting the tissue area, and detecting the background and blurred area.

[0040] The step S2 specifically includes: in order to improve the robustness and generalization of the PRESORT model, the pre-trained models of KimiaNet and CTransPath are used to extract feature representations with dimensions of 1024 and 768 for each patch.

[0041] The KimiaNet pre-trained model is a pre-trained model with 7 million weights, which is fine-tuned and trained based on the DenseNet framework. Unlike other convolutional neural networks (CNNs), it has learned millions of pathological patches to effectively extract the features of different patches. The local connection and sliding window operation of the convolutional layer of the KimiaNet pre-trained model enable the network to capture local features, and gradually expand the receptive field through multi-layer convolution operations, so that the network can extract hierarchical features, thereby achieving effective representation of the input data. Each patch embedding is unified into a 1024-dimensional vector representation by the global pool and fully connected output.

[0042] The CTransPatht pre-trained model is based on the Swin Transformer architecture, which consists of three convolutional layers and four Swin Transformer modules. The convolutional layers in CTransPath can extract local features in image data, and Swin Transformer can capture global features and long-distance dependencies. Combining the two, we can fully utilize the advantages of CNN in local feature extraction, while utilizing the advantages of Swim transformer in global feature capture to improve the representation ability of the model. This embodiment uses CTransPath to extract the embedded features of patches, and each patch embedding is unified into a 768-dimensional vectorized representation. Secondly, in order to avoid the model falling into absolute bias, this embodiment uses two pre-trained models to extract features for further analysis.

[0043] In this embodiment, all patches in the slide image in step S3 are represented in the form of feature vectors. The cells and tissue structures in the WSI have a relative position relationship on a two-dimensional plane. This embodiment uses the K-nearest neighbor algorithm (K=9) to construct a spatial topological graph G, which is composed of points and edges. G={v,e}, v represents the node set of the graph, e represents the set of edges, A is the adjacency matrix, and X is the feature of the node. The spatial topological graph G can help extract the features in the WSI and analyze the spatial relationship between these features. These features may include cell morphology, nuclear morphology, cell distribution, etc. By analyzing the relationship between these features, the tissue structure and pathological characteristics can be better understood.

[0044] In step S4, the twin hybrid encoder is composed of two symmetrical encoders. The output end of the twin hybrid encoder is the features output by the two encoders, which are concatenated and averaged before being output. Each encoder includes a GCNConv, a SAGEConv, two ReLU layers, a Dropout operation, an MLP network and a Rescale layer. GCNConv initializes the convolution kernel weight matrix W, multiplies the node feature X by the weight matrix W to obtain the linear transformation result Z=XW of the node, and uses the adjacency matrix A to propagate the node features; SAGEConv uses the neighbor node features of G to update the feature representation of the node.

[0045] The use of the adjacency matrix A to propagate the features of the nodes specifically includes:

[0046] For node i, the weighted sum of the features of its neighboring nodes is added to the node's own features to obtain a new feature representation. The process is:

[0047]

[0048] In the formula, H (l-1) Represents the feature representation of the l-1th layer node is the adjacency matrix A plus the self-connectivity matrix, is a diagonal matrix, where σ is the ReLU activation function;

[0049] The SAGEConv uses the neighbor node features of G to update the feature representation of the node, specifically including:

[0050] For each node v, select its neighbor node set N (v) , use mean pooling to aggregate its neighbor node set, and aggregate the feature of node v It is concatenated with the aggregated neighbor node features, and then goes through a linear transformation and a nonlinear activation function to obtain the updated representation H of node v. (l+1) , the process can be expressed as:

[0051]

[0052] In the formula, W is the weight matrix, Concat concatenates the node's own features with the features of neighboring nodes, Aggregate is an aggregation operation, and the features of the Graph are subjected to a ReLU layer and a Dropout operation, and then an MLP network and a Rescale layer for feature extraction and dimensionality reduction.

[0053] Importantly, the twin hybrid encoder proposed in this embodiment has a skip connection and weight sharing operation. In the encoder that processes the 768-dimensional feature vector, a skip connection is used to pass the original input information to the back end, which not only retains the original input information but also improves the gradient flow, helping to alleviate the gradient vanishing problem and overfitting problems. The twin hybrid encoder shares the weights of the lower encoder with the upper encoder to promote the PRESORT model to learn more robust feature representations. Secondly, through feature sharing, the model can learn a general feature representation, rather than a feature representation for a specific position, thereby improving the generalization performance of the model.

[0054] The training of all diagnostic models is based on WSI sample-level research, not patient-level research. The PRESORT model uses an internal validation set to verify the performance of the model through a five-fold cross-validation method. In this cross-validation method, 1,190 WSIs are divided into five groups, four groups are trained in each fold, and the remaining group is used for validation. After the five-fold cross-validation, the five best prediction models will be obtained. Most importantly, this embodiment uses an external validation set to test the generalization performance of the model. The model predicts in a data set containing 40 frozen sections to assist clinicians in making clinical decisions quickly, and predicts in a data set containing 100 paraffin sections to assist pathologists in completing pathological diagnosis reports.

[0055] The PRESORT model is trained using the RMSprop optimizer, with a learning rate of 0.001, alpha=0.9, and a weight decay rate of 0.9. The batch size is set to 1. The main reason is that the number of patches in each WSI is variable, and the adjacency matrix of the graph is different in size, so it is impossible to stack graphs of different sizes into batches of equal size output. Secondly, the PRESORT model will occupy a large amount of graphics card memory during inference, and increasing the batch size will cause memory overflow. The epoch is set to 200, and the training data will be traversed once per epoch.

[0056] The present invention is further described by the following examples:

[0057] In this example, the PRESORT model is trained and tested on a dataset of lung cancer patients' pathological sections to predict the positive and negative of STAS (e.g. Figure 1 ). First, this embodiment uses a five-fold cross-validation method to divide the internal validation set (including permanent sections and paraffin sections) into five groups, using four groups of data for training in each training phase and the remaining group for testing in each test phase, and finally obtaining five optimal STAS prediction models. This embodiment found that the model of the third fold cross-validation obtained an intra-domain test AUROC of about 0.92 ( Figure 2 A), the fourth fold model obtained an in-domain test PRC of 0.95 ( Figure 2 B). The five-fold cross-validation model achieved an average in-domain test AUROC of 0.87 ( Figure 2 A) and an average of 0.91 in-domain test PRC ( Figure 2 B). Secondly, the model of the third fold cross validation achieved the best prediction effect, with the values ​​of accuracy, precision, recall, f1, specificity, and AUC being 0.85, 0.84, 0.84, 0.84, 0.81, and 0.92 respectively. Finally, by looking at the confusion matrix ( Figure 2 D) It was found that the five models trained by the five-fold cross-validation method also achieved good prediction accuracy. Therefore, this embodiment found that the PRESORT model can accurately predict the positive and negative of STAS. It must be mentioned that this embodiment also conducted relevant comparative experiments and ablation experiments (Appendix 1), which proved that the PRESORT model is currently the best model for predicting the positive and negative of STAS. It must be mentioned that the model trained by the fifth-fold cross-validation may be due to the imbalance of categories in the training data and the test data, which makes it difficult for the model to learn good feature representations during the training and verification process. Secondly, there is randomness in the division of data and the initialization of the model during the cross-validation process.

[0058] Next, in order to verify the generalization performance of the PRESORT model, this embodiment uses an external validation dataset, including 40 frozen sections and 100 paraffin sections, to test the PRESORT model. This embodiment uses the PRESORT model to obtain an in-domain test AUROC of 0.83 on frozen sections ( Figure 2 E) and the in-domain test PRC of 0.85 ( Figure 2 F). The PRESORT model achieved an intra-domain test AUROC of 0.883 on paraffin sections ( Figure 2 E) and the in-domain test PRC of 0.85 ( Figure 2 F). This example found that the performance of the PRESORT model in predicting frozen section STAS is low, which may be due to the low participation of frozen sections in the model training process. Secondly, it may also be due to the poor quality of frozen sections, the destruction of cell structure, and the small sample size, which leads to large errors in the model's interpretation of frozen sections. However, by observing the confusion matrix ( Figure 2 G), and also achieved an accuracy of 80%. The prediction performance of STAS in paraffin sections is better, which may be due to the stable quality of the sections, the complete structure of the cells and the large number of paraffin sections in the internal dataset that were extracted by the PRESORT model.

[0059] This embodiment hopes that the model tends to be interpretable and provides methods and results for pathologists. To this end, this embodiment first allows pathologists to delineate the main tumor area and STAS area (positive and negative STAS) in WSI. Figure 3 A). In STAS-positive WSI, microemulsified, solid nested, and single-cell STAS are clearly distributed next to the main tumor. Then, this example visualizes the final contribution of each patch obtained by the PRESORT model and returns these contributions to the WSI based on the location of the patch and whether they contribute to the identification of the tumor area ( Figure 3 B).

[0060] In STAS-negative WSI, patches with higher contributions are concentrated in the main tumor area of ​​WSI. In STAS-positive WSI, patches with higher contributions are not only in the main tumor area but also outside the main tumor. These color markings help pathologists to find out which areas are important areas of concern through global observation. At the same time, the categories with relatively low contributions may be the interstitial areas and normal tissue areas of WSI. Because STAS often appears at the edge of the main tumor, pathologists should pay more attention to the areas at the edge of the main tumor and the connected interstitial areas, which can save pathologists time in reading films and improve the efficiency of reading films. In this embodiment, 9 patches with relatively high contributions are screened out through the PRESORT model. They provide cellular-level features including the morphology of the nucleus, the state of mitosis, etc., which provide important data support for studying the growth laws and metastasis characteristics of tumors.

[0061] Although the embodiments of the present invention are described in conjunction with the accompanying drawings, the patent owner may make various variations or modifications within the scope of the appended claims. As long as they do not exceed the protection scope described in the claims of the present invention, they should be within the protection scope of the present invention.

Claims

1. A method for predicting intracavitary dissemination from lung cancer tissue pathology images, characterized in that: The following steps are involved: S1. Preprocessing the whole slide image; S2, extracting features from the preprocessed image based on the pre-trained model; S3, constructing a spatial topological map based on the features of the slide image; S4, feature extraction of spatial topology map based on twin hybrid encoder; S5, using the features extracted in step S4 to predict the lung cancer tissue pathology image, and obtaining the result of intracavitary dissemination; In step S3, all patches in the slide image are represented in the form of feature vectors, and the K-nearest neighbor algorithm is used to construct a spatial topology graph G. G ={v,e}, v represents the node set of the Graph, e represents the edge set, A is the adjacency matrix, and X is the feature of the node; In step S4, the twin hybrid encoder is composed of two symmetrical encoders. The output end of the twin hybrid encoder is the features output by the two encoders, which are concatenated and averaged before being output. Each encoder includes a GCNConv, a SAGEConv, two ReLU layers, a Dropout operation, an MLP network and a Rescale layer. GCNConv initializes the convolution kernel weight matrix W, multiplies the node feature X by the weight matrix W to obtain the linear transformation result Z=XW of the node, and uses the adjacency matrix A to propagate the node features. SAGEConv uses the neighbor node features of G to update the feature representation of the node; The use of the adjacency matrix A to propagate the features of the nodes specifically includes: For node i, the weighted sum of the features of its neighboring nodes is added to the node's own features to obtain a new feature representation. The process is: In the formula, H (l-1) represents the feature representation of the l-1th layer node, is the adjacency matrix A plus the self-connectivity matrix, is a diagonal matrix, where σ is the ReLU activation function; The SAGEConv uses the neighbor node features of G to update the feature representation of the node, specifically including: For each node v, select its neighbor node set N(v), aggregate its neighbor node set using meanpooling, and aggregate the features of node v It is concatenated with the aggregated neighbor node features, and then goes through a linear transformation and a nonlinear activation function to obtain the updated representation H of node v. (l+1) , the process can be expressed as: In the formula, W is the weight matrix, Concat concatenates the node's own features with the features of neighboring nodes, Aggregate is an aggregation operation, and the features of the Graph are subjected to a ReLU layer and a Dropout operation, and then an MLP network and a Rescale layer for feature extraction and dimensionality reduction.

2. The method for predicting intracavitary dissemination from lung cancer tissue pathology images according to claim 1, characterized in that: The step S1 specifically includes: digitizing the slide image, segmenting the tissue area, and detecting the background and blurred area.

3. The method for predicting intracavitary dissemination from lung cancer tissue pathology images according to claim 1, characterized in that: The step S2 specifically includes: using the pre-trained models of KimiaNet and CTransPath to extract feature representations with dimensions of 1024 and 768 for each patch.

4. The method for predicting intracavitary dissemination from lung cancer tissue pathology images according to claim 3, characterized in that: The KimiaNet pre-trained model is a pre-trained model with 7 million weights that is fine-tuned and trained based on the DenseNet framework. The local connection and sliding window operation of the convolutional layer of the KimiaNet pre-trained model enable the network to capture local features; the CTransPatht pre-trained model is based on the Swin Transformer architecture, which consists of three convolutional layers and four Swin Transformer modules. The convolutional layers in CTransPath can extract local features in image data, and Swin Transformer can capture global features and long-distance dependencies.

Citation Information

Patent Citations

  • Optical remote sensing image slice level change detection method and device

    CN117372842A

  • Lung cancer WSI air cavity dissemination pathological diagnosis device based on deep learning

    CN117423447A

  • Cervical cell sample age and health degree prediction method

    CN117635551A