Pathological full-slice image-based data enhancement method and system, and storage medium
By using the conditional diffusion model to generate new WSI samples and image block representations in pathological full-slice image processing, the data incompleteness problem caused by insufficient number of WSIs is solved, and the accuracy and robustness of survival predictions are improved.
Patent Information
- Application Number
- CN202510130338.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-06-13
AI Technical Summary
When processing pathological full-slice images, the prior art cannot effectively solve the data incompleteness problem caused by insufficient number of WSIs, resulting in limited training of survival prediction models and inability to fully realize the data potential.
Using a data augmentation method based on the conditional diffusion model, new WSI samples and image block representations are generated through WSI-level and patch-level conditional diffusion models, introducing possible but unobserved missing information to fill the data gap.
Improve the accuracy and robustness of survival prediction, and enhance the generalization ability and prediction performance of the model by introducing new biological information and diversified training data.
Smart Images

Figure CN120147148A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of survival prediction tasks, and in particular, to a data augmentation method and system based on whole slide pathology images. Background Art
[0002] In survival prediction tasks, whole slide pathology images (WSIs) can serve as core data. These are digital tissue section images generated by high-throughput scanning devices and can comprehensively reflect the tissue and cell characteristics of patients. As an important tool for pathological analysis, WSIs are widely used to study key features such as tissue heterogeneity, cell morphology, and lesion areas, providing data support for precision medicine and personalized treatment.
[0003] However, in practical applications, there may be a problem of uneven distribution of WSIs among patients, that is, some patients have a large number of WSIs, while some patients only have a small number of available WSIs. This unevenness may be caused by limited sampling conditions (such as tissue size and lesion area distribution) or individual patient differences (such as disease progression). In patients with a small number of WSIs, the survival prediction results may be inaccurate, thus misleading the treatment decisions of patients. For example, the survival prediction model may underestimate the invasiveness of tumors or misevaluate the survival risk of patients, thereby affecting the selection of treatment plans and prognosis judgment, ultimately endangering the health of patients. Moreover, this may also lead to unfair distribution of medical resources, making patients with less data unable to fully benefit from advanced deep learning technologies, further exacerbating the inequality of medical services. Therefore, it is necessary to study the problem of insufficient WSI quantity to improve the fairness and accuracy of survival prediction.
[0004] To address the problem of insufficient whole-slide images of patients, some classic existing data augmentation techniques can be adopted, such as resampling and AugDiff (Diffusion based feature augmentation). These methods can alleviate the impact of the imbalance in the number of WSIs among patients on the training of the survival prediction model by adjusting the distribution of WSIs and generating diverse WSIs. The resampling method can balance the number of samples among different categories by oversampling or undersampling patient samples, thereby adjusting the distribution of the training data and reducing the interference of imbalance on the training process: in the case of oversampling, the patient category with fewer WSIs will increase its quantity by replicating existing samples or synthesizing new samples; while in the case of undersampling, the patient category with more WSIs will randomly remove some samples to avoid overfitting of the model to these categories. The AugDiff method combines image-level and region-level data augmentation techniques to generate diverse samples by introducing random noise and perturbations, which not only effectively increases the scarce sample volume of some patients but also enhances the robustness of the model by introducing more noise and variations. Specifically, AugDiff perturbs or transforms local regions on the original WSI images of patients to generate a series of new WSIs. These samples not only retain the key information of the original images but also improve the model's adaptability to data diversity through varying degrees of changes. By adopting these classic resampling and AugDiff methods, the existing technology can alleviate the problem of WSI quantity imbalance to a certain extent, provide more balanced and diverse training data for the survival prediction model, and thus improve the generalization ability and accuracy of the model.
[0005] However, existing data augmentation methods, such as resampling and AugDiff, mainly rely on modifying existing image regions to generate enhanced whole-slide images. Since these methods mainly generate new samples by replicating, perturbing, or adjusting existing tissue and tumor microenvironment regions, although effectively expanding the quantity and diversity of samples, these methods fail to introduce truly new information, and the enhanced samples are essentially similar to the original images. That is, the enhanced WSIs generated by existing data augmentation methods are still limited to the expansion of the original data and do not predict unobserved or missing image patterns. Since these methods can only enhance within the existing knowledge range and cannot generate "missing" WSIs that may exist but have not been observed, the problem of current data incompleteness has not been substantially solved.
[0006] Considering that resampling and AugDiff cannot provide additional, uncaptured biological information or image patterns, especially in the case of scarce sample numbers, this may lead to the training of survival prediction models still being restricted by data incompleteness and unable to fully exploit the potential of existing data. Therefore, there is an urgent need for a new data augmentation method for whole-slide pathology images that can introduce missing information that may exist but has not been observed when the number of patients' WSIs is insufficient, thereby improving the accuracy of survival prediction. Summary of the Invention
[0007] In view of this, the embodiments of the present invention provide a data augmentation method and system based on whole-slide pathology images, which can introduce missing biological information or image patterns that may exist but have not been observed into the prediction input data when the number of patients' WSIs is insufficient, thereby improving the accuracy of survival prediction.
[0008] One aspect of the present invention provides a data augmentation method based on whole-slide pathology images, the method comprising the following steps:
[0009] Obtain at least one whole-slide pathology image (WSI), divide the obtained WSI into a plurality of non-overlapping image patches, extract the corresponding WSI representation from the WSI, and extract the corresponding image patch representations from each image patch;
[0010] Obtain the conditional input of the pre-trained first conditional diffusion model based on the extracted WSI representation, and input a pre-set noise image into the pre-trained first conditional diffusion model to output a predicted noise WSI representation and a predicted noise tissue type distribution;
[0011] Determine the tissue types included in the predicted noise tissue type distribution, obtain the conditional input of the pre-trained second conditional diffusion model based on the class embedding vectors corresponding to the determined tissue types and the predicted noise WSI representation, and input each image patch representation into the pre-trained second conditional diffusion model respectively to output the corresponding predicted noise image patch representations.
[0012] In some embodiments of the present invention, obtaining the conditional input of the pre-trained first conditional diffusion model based on the extracted WSI representation includes:
[0013] If the number of obtained WSIs is one, use the extracted WSI representation as the conditional input of the pre-trained first conditional diffusion model;
[0014] If the number of obtained WSIs is multiple, fuse the extracted WSI representations through average pooling to obtain an average WSI representation, and use the average WSI representation as the conditional input of the pre-trained first conditional diffusion model.
[0015] In some embodiments of the present invention, obtaining the conditional input of the pre-trained second conditional diffusion model based on the category embedding vector corresponding to the determined tissue type and the predicted noise WSI representation includes: combining the predicted noise WSI representation and the category embedding vector of the determined tissue type through the Hadamard product, and using the obtained combined constraint as the conditional input of the pre-trained second conditional diffusion model.
[0016] In some embodiments of the present invention, the pre-trained first conditional diffusion model is trained in the following manner:
[0017] Obtain the training WSI and the corrected WSI belonging to the same patient, divide the corrected WSI into multiple non-overlapping image patches, and extract the corresponding WSI representations from the training WSI and the corrected WSI;
[0018] Use the cell recognition model to classify the tissue types of the image patches corresponding to the corrected WSI to obtain the tissue type distribution corresponding to the corrected WSI;
[0019] Based on the WSI representation corresponding to the training WSI, obtain the conditional input of the initial first conditional diffusion model, input the pre-set noise image into the initial first conditional diffusion model, output the training-predicted noise WSI representation and the training-predicted noise tissue type distribution, and compare the WSI representation and the tissue type distribution corresponding to the corrected WSI with the output results of the initial first conditional diffusion model, so as to adjust the parameters of the initial first conditional diffusion model to obtain the pre-trained first conditional diffusion model;
[0020] The pre-trained second conditional diffusion model is trained in the following manner:
[0021] Divide the WSI used to train the second conditional diffusion model into multiple non-overlapping image patches, and extract the corresponding image patch representations from each image patch;
[0022] Use the WSI representation corresponding to the WSI used to train the second conditional diffusion model as the conditional input of the first conditional diffusion model, obtain the conditional input of the initial second conditional diffusion model based on the output results of the first conditional diffusion model, and input the extracted image patch representations into the initial second conditional diffusion model respectively;
[0023] Compare the output results of the initial second conditional diffusion model with the input image patch representations, so as to adjust the parameters of the initial second conditional diffusion model to obtain the pre-trained second conditional diffusion model.
[0024] In some embodiments of the present invention, using the cell recognition model to classify the tissue types of the image patches corresponding to the corrected WSI includes:
[0025] The cells in the image patches corresponding to the corrected WSI are recognized by using a pre-trained cell recognition model, and the tissue type corresponding to each image patch is determined according to the majority statistical strategy, and a class label for identifying the tissue type is assigned to each image patch whose tissue type is determined; wherein, the class label is the class embedding vector of the tissue type corresponding to the image patch, and the cell recognition model is trained based on the PanNuke dataset including tissue type classification.
[0026] In some embodiments of the present invention, the tissue type distribution corresponding to the corrected WSI is determined by the following method:
[0027] Based on the number of image patches corresponding to each tissue type, the class labels of the image patches obtained by splitting the corrected WSI are normalized, so as to obtain the tissue type distribution of the corrected WSI.
[0028] In some embodiments of the present invention, both the WSI representation and the image patch representation include the morphological features and arrangement features of cells and tissues; and
[0029] Determining the tissue types included in the predicted noise tissue type distribution includes:
[0030] Determining the tissue types whose tissue type values are not zero from the predicted noise tissue type distribution; wherein, the tissue types include neoplastic, dead, inflammatory, non-neoplastic epithelium, connective tissue or unlabeled.
[0031] In some embodiments of the present invention, splitting the obtained WSI into a plurality of non-overlapping image patches includes: dividing the obtained WSI image into a plurality of non-overlapping image patches by means of a sliding window strategy and the Otsu algorithm.
[0032] Another aspect of the present invention provides a data augmentation system based on a whole slide image of pathology, including a processor, a memory and a computer program / instructions stored on the memory, and the processor is used to execute the computer program / instructions, and when the computer program / instructions are executed, the system implements the steps of the method described in any of the above embodiments.
[0033] Another aspect of the present invention provides a computer-readable storage medium, on which computer program / instructions are stored, and when the computer program / instructions are executed by a processor, the steps of the method described in any of the above embodiments are implemented.
[0034] The data augmentation method and system based on whole-slide pathology images proposed by the present invention can utilize a WSI-level conditional diffusion model to generate new WSIs based on the patient's original WSIs, and combine the newly generated WSIs with the patient's original patch representations through a patch-level conditional diffusion model, thereby generating new patch representations for the patient. Further, considering the prognostic information contained in the tissue type distribution changes, the present application adds the tissue type distribution as a condition in the process of newly generating patch representations. The present application can not only provide uncaptured missing information and improve the accuracy of survival prediction, but also enhance the robustness of the survival prediction model by using the noise and variations introduced by the diffusion model.
[0035] Additional advantages, objects, and features of the present invention will be partially described below, and will become partially apparent to those of ordinary skill in the art after studying the following, or can be learned from the practice of the present invention. The objects and other advantages of the present invention can be realized and obtained by the structures specifically pointed out in the specification and the drawings.
[0036] Those skilled in the art will understand that the objects and advantages that can be achieved by the present invention are not limited to the above specifically described, and the above and other objects that the present invention can achieve will be more clearly understood according to the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not limit the present invention. In the drawings:
[0038] Figure 1 is a schematic flowchart of a data augmentation method based on whole-slide pathology images in an embodiment of the present invention.
[0039] Figure 2 is a schematic diagram of the training process of a first conditional diffusion model and a second conditional diffusion model in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] To make the objects, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the embodiments and the drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention.
[0041] Here, it should also be noted that in order to avoid obscuring the present invention with unnecessary details, only the structures and / or processing steps closely related to the solution of the present invention are shown in the drawings, and other details less related to the present invention are omitted.
[0042] It should be emphasized that the term "comprising / including" as used herein refers to the presence of features, elements, steps, or components, but does not exclude the presence or addition of one or more other features, elements, steps, or components.
[0043] In the following, embodiments of the present invention will be described with reference to the accompanying drawings. In the drawings, the same reference numerals represent the same or similar components, or the same or similar steps.
[0044] In the survival prediction task of whole slide images (WSIs) in pathology, resampling and AugDiff are often used to address the issue of incomplete data. Although existing data augmentation techniques can expand the number of samples, they are limited to modifying and expanding existing image regions and cannot generate truly new WSI samples, resulting in the training of the survival prediction model still being restricted by the original data. That is, the existing method of data augmentation by only modifying existing image regions cannot effectively solve the challenges brought by data missing, especially in the case of data imbalance or lack of key samples.
[0045] To address the limitations of existing methods in terms of data incompleteness and sample scarcity, the present invention proposes an innovative data augmentation method that can generate completely new WSI samples from scratch to fill in the missing or unknown parts in the existing data. Different from the traditional data augmentation method that only modifies existing samples, the present invention can generate new WSI samples by simulating the possible missing data patterns in the real world based on the tissue type distribution of the existing data, thus effectively solving the current possible problem of data incompleteness.
[0046] Figure 1 The flowchart of the data augmentation method based on whole slide images in pathology proposed in this application is shown as Figure 1 shown. This method includes steps S110 to S130. The data augmentation method of this application mainly includes the following three steps: Step S110, in the data preparation stage, perform WSI modeling at multiple scales; Step S120, in the stage of generating WSI features, use a WSI-level conditional diffusion model (i.e., the first conditional diffusion model) to generate completely new WSI-level features; Step S130, taking the WSI-level features generated in Step S120 as the condition and introducing the corresponding tissue type constraints, use a patch-level conditional diffusion model (i.e., the second conditional diffusion model) to generate the image patch features corresponding to the original WSI of the patient. Finally, the image patch features generated in Step S130 can be input into the survival prediction model to predict the survival time of the patient. By generating completely new WSI features in Step S120 and integrating them into the image patch features generated from the original WSI in this application, new pathological information is added, which helps to solve the problem of insufficient WSI data for some patients.
[0047] Step S110: Obtain at least one whole slide image (WSI) of pathology, divide the obtained WSI into multiple non-overlapping image patches (patches), extract the corresponding WSI representation (also referred to as WSI features) from the obtained WSI, and extract the corresponding patch representations (also referred to as patch features) from each patch. Among them, a patch can represent a local area and is usually used for feature extraction and analysis.
[0048] The reasons for feature extraction in different dimensions in this application are as follows: Due to the huge size of the WSI (in gigabytes), it is not feasible to perform survival prediction using the WSI in the image space and it needs to be mapped to the feature space; moreover, since the current survival prediction models all use patch representations as inputs, the method of generating patches in the image space for survival prediction is relatively indirect.
[0049] In the data preparation stage, mainly use feature extraction to prepare the input data for the WSI-level conditional diffusion model and the patch-level conditional diffusion model. Through feature extraction, the WSI representation and patch representation corresponding to the original whole slide image of the patient can be obtained at the WSI level and the patch level respectively (that is, extract the WSI-level features and patch-level features from the obtained WSI), for subsequent hierarchical diffusion. The process of feature extraction in step S110 is as follows: Step S111, divide the obtained original WSI of the patient into multiple non-overlapping image patches; Step S112, map the obtained WSI and the image patches obtained in step S111 from the image space to a shared representation space (that is, the feature space).
[0050] In some embodiments of the present invention, in step S111, multiple non-overlapping and non-background image patches can be cropped from each obtained WSI through a sliding window strategy and the Otsu algorithm (i.e., the OTSU threshold method). In this application, other algorithms or methods can be used to divide the image patches from the WSI, as long as there is no overlap between the image patches obtained from the same WSI. Moreover, the sizes of the image patches obtained from the same WSI should be as consistent as possible, the number of image patches obtained from different WSIs can be different, and the sizes of the image patches divided from different WSIs can also be inconsistent.
[0051] Furthermore, for the obtained WSI, through feature extraction, the WSI can be mapped to a representation vector (i.e., the WSI representation), and one WSI can obtain a corresponding WSI representation; for each image patch, through feature extraction, the image patch can be mapped to the representation space, and one image patch can obtain a corresponding patch representation. Therefore, the number of patch representations obtained after step S112 is the same as the number of image patches obtained by dividing the obtained WSI.
[0052] As an example, the WSI representation may include the morphological and arrangement features of cells in the WSI, as well as the morphological and arrangement features of the tissue; the patch representation may include the morphological and arrangement features of cells in the patch, as well as the morphological and arrangement features of the tissue. Additionally, as Figure 2 shown, the corresponding WSI representation can be extracted from the WSI using a pre-trained first feature extractor f w , and the corresponding patch representation can be extracted from each patch using a pre-trained second feature extractor f e . Here, the first and second feature extractors can be neural network models for learning the latent representation of pathological data. In addition to using feature extractors, the WSI representation and patch representation can also be obtained by manually extracting features. The present invention does not specifically limit the extraction methods of the WSI representation and patch representation.
[0053] Diffusion models generate new images by gradually adding noise to an image and then removing the noise through a reverse diffusion process. Different from traditional diffusion models, conditional diffusion models (CDMs, also known as conditional denoising diffusion probability models) introduce additional conditional information during the generation process, including the forward diffusion process and the reverse diffusion process. In the forward diffusion process, noise is gradually added to a clear image to obtain a sequence of increasingly blurred images; in the reverse diffusion process, a learned neural network gradually removes the noise from the noisy image to restore the original image, and the reverse diffusion process also incorporates conditional inputs to ensure that the generated images meet the conditional requirements. This application can utilize conditional diffusion models for data generation and enhancement. For example, the conditional diffusion model can be a DDIM (Denoising diffusion implicit models) model. The present invention does not specifically limit the specific type of conditional diffusion model, as long as it can generate high-quality image data under given conditions.
[0054] Step S120: Obtain the conditional input of the pre-trained first conditional diffusion model based on the extracted WSI representation, and input a pre-set noisy image into the pre-trained first conditional diffusion model, which can output a predicted noisy WSI representation and a predicted noisy tissue type distribution.
[0055] As Figure 2As shown, considering the WSI representation corresponding to the obtained WSI comprehensively, the present application can obtain the conditional input of the pre-trained first conditional diffusion model based on the extracted WSI representation, including: if the number of obtained WSIs is one, the extracted WSI representation is used as the conditional input of the pre-trained first conditional diffusion model; if the number of obtained WSIs is multiple, an average WSI representation is obtained by performing average pooling (mean-pooling) on the extracted WSI representations, and the obtained average WSI representation is used as the conditional input of the pre-trained first conditional diffusion model. That is, if there are multiple WSIs of a patient, the WSI representations corresponding to each obtained WSI can be fused by average pooling to obtain an average WSI representation including the features of the patient's original WSI. The average pooling mentioned in the present application specifically refers to weighted average, and other methods of comprehensively representing WSIs can also be used, and the present invention is not limited thereto.
[0056] In some embodiments of the present invention, the WSI-level conditional diffusion model is trained based on the conditional diffusion model, and the training process is as follows: Obtain the training WSI and the corrected WSI belonging to the same patient, divide the corrected WSI into multiple non-overlapping image patches, and extract the corresponding WSI representations from the training WSI and the corrected WSI; Use the cell recognition model to classify the tissue types of the image patches corresponding to the corrected WSI to obtain the tissue type distribution corresponding to the corrected WSI; Obtain the conditional input of the initial first conditional diffusion model based on the WSI representation corresponding to the training WSI, input the pre-set noise image into the initial first conditional diffusion model, and the training-predicted noise WSI representation and the training-predicted noise tissue type distribution can be output, and compare the output results of the initial first conditional diffusion model with the WSI representation and the tissue type distribution corresponding to the corrected WSI, so as to adjust the parameters of the initial first conditional diffusion model; Iteratively train until the first training condition is met (i.e., the training termination condition of the initial first conditional diffusion model), and the pre-trained first conditional diffusion model can be obtained.
[0057] As an example, since the WSI-level conditional diffusion model is used to generate a new WSI representation based on the patient's original WSI, to improve the accuracy of the WSI-level conditional diffusion model, the training WSI and the calibration WSI in this application need to belong to the same patient to reduce errors. Moreover, to improve the model training effect, the WSI data of multiple patients can be included in the training set, and each patient in the training set includes at least two WSIs (during training, one whole slide image is selected as the calibration WSI, and one or more whole slide images are selected as the training WSIs). In addition, the training objective of the initial first conditional diffusion model in this application can be to minimize the difference between the predicted noise output by the model and the real noise, that is, to minimize the difference between the training-prediction noise WSI representation output by the initial first conditional diffusion model and the WSI representation corresponding to the calibration WSI, and to minimize the training-prediction noise tissue type distribution output by the initial first conditional diffusion model and the tissue type distribution corresponding to the calibration WSI. The first training condition in this application can be determined according to the training objective of the model. For example, the value of the loss function is less than a set threshold. The present invention does not specifically limit the training objective and training conditions of the initial first conditional diffusion model.
[0058] In some embodiments of the present invention, the tissue type of the image patches corresponding to the calibration WSI is classified by using a cell recognition model, including: using a pre-trained cell recognition model to recognize the cells in the image patches corresponding to the calibration WSI, determining the tissue type corresponding to each image patch according to the majority statistical strategy, and assigning a class label for identifying the tissue type to each image patch with the determined tissue type.
[0059] Specifically, the tissue types mentioned in this application can be six types defined in the PanNuke dataset (the cell differences between various types are obvious), including neoplastic, dead, inflammatory, non-neoplastic epithelial, connective tissue or unlabeled. Therefore, to ensure that the cell recognition model can recognize the cells (or cell nuclei) in the image patches based on the above six-class tissue types, the cell recognition model in this application can be trained based on the PanNuke dataset including tissue type classification. Moreover, since an image patch may include multiple cells, the cell recognition model can determine the tissue types of multiple cells in the image patch, and there may be a situation where the tissue types of the cells are different. Considering that the purpose of this application is to assign a tissue type to an image patch, the majority statistical strategy can be used to count the number of cells belonging to each tissue type in the image patch, and the tissue type with the largest number of cells is determined as the tissue type corresponding to the image patch, that is, the tissue type of each image patch corresponding to the calibration WSI is determined based on the principle of majority voting. The above-mentioned cell recognition model can be the HoverNet model. The present invention does not specifically limit the type of the cell recognition model, as long as it can recognize the type of cells or cell nuclei in the image patch.
[0060] Combining the cell recognition model and the majority statistical strategy can assign a tissue type to each image patch (the image patch corresponding to the corrected WSI). After determining the tissue type of the image patch, a class label can also be assigned to each image patch according to the determined tissue type to identify the tissue type. Moreover, in this application, the class embedding vector of the tissue type corresponding to the image patch can be used as the class label. The class embedding vectors corresponding to each tissue type are preset constants, which can be custom vectors in the same dimension, and for different tissue types, the class embedding vectors are also different (that is, the class embedding vectors of each tissue type are uniquely represented). This application does not specifically limit the representation form of the class embedding vector. The purpose of customizing the class embedding vector for each tissue type in this application is to introduce the tissue type distribution and implement diffusion in the patch-level conditional diffusion model. Therefore, the class label mentioned in this application needs to be a constant. In addition, the class label can be a one-hot label or other labels used to represent the tissue type corresponding to the image patch. The present invention is not limited thereto.
[0061] Furthermore, the tissue type distribution corresponding to the corrected WSI is determined in the following manner:
[0062] Based on the number of image patches corresponding to each tissue type, the class labels of the image patches obtained by dividing the corrected WSI are normalized to obtain the tissue type distribution of the corrected WSI. Among them, the value of each tissue type in the tissue type distribution is the proportion of the number of image patches of each tissue type in the corrected WSI. For example, the corrected WSI can be divided into 20 image patches, the class labels of 2 image patches are A, and the class labels of 18 image patches are B. Then the value of the tissue type with the class label A in the tissue type distribution is 0.1, and the value of the tissue type with the class label B in the tissue type distribution is 0.9.
[0063] Specifically, the WSI-level conditional diffusion model includes two processes: the forward process and the reverse process, which are as follows:
[0064] (1) Forward process: Although the predicted noise WSI representation and the predicted noise tissue type distribution output by the WSI-level conditional diffusion model are in different distribution spaces, due to their similar structures, the independent forward diffusion process in the WSI-level conditional diffusion model can be described by a unified formula. For example, the forward diffusion process can be represented as a Gaussian distribution, and its noise schedule controls the variance of the noise, which can be set differently for the output WSI representation and tissue type distribution. That is, the image input in the previous stage is weighted and fused with a Gaussian distribution sample for noise addition processing. To improve efficiency, the forward diffusion process can be derived through the reparameterization trick.
[0065] (2) Reverse process: Considering the strong correlation between WSI representations and tissue type distributions, the present application generates new WSI representations and corresponding tissue type distributions simultaneously during the reverse diffusion process of the WSI-level conditional diffusion model. Specifically, according to the dimensions of the WSI features and the tissue type distributions finally output by the pre-trained first conditional diffusion model, noise is sampled from a Gaussian distribution (the dimension of the noise is the sum of the dimensions of the WSI features and the tissue type distributions that the model will output), and the sampled noise is divided into noise for WSI representations and noise for tissue type distributions according to the dimensions; denoising processing is performed on these two noises, so as to generate a noisy WSI representation and a tissue type distribution with the original WSI features of the patient as a condition during the reverse process.
[0066] As an example, the process of obtaining the newly generated WSI representation (i.e., the predicted noise WSI representation) and the corresponding predicted noise tissue type distribution of the WSI representation using the pre-trained first conditional diffusion model in the present application is as follows: After obtaining the WSI representation of the original WSI of the patient in step S110, all the WSI representations are fused to obtain an average WSI representation, and the average WSI representation is used as the conditional input of the pre-trained first conditional diffusion model, and a noise image is input. The pre-trained first conditional diffusion model continuously repeats the reverse process and the forward process to output the predicted noise WSI representation and the predicted noise tissue type distribution corresponding to the predicted noise WSI representation.
[0067] In step S120, a pre-trained WSI-level conditional diffusion model can be used to generate new WSI representations and corresponding tissue type distributions. Based on the output result of the WSI-level conditional diffusion model, the conditional input of the patch-level conditional diffusion model can be obtained, so as to introduce new biological information into the image patch representation of the patient.
[0068] The reasons for introducing the tissue type distribution in the data augmentation method proposed in the present application are as follows: The image patch features within the WSI follow distributions corresponding to different tissue types, and the changes in the tissue type distributions between WSIs can convey different prognostic information. For example, a higher proportion of tumor patches may indicate that the tumor is more invasive and indicate a poorer prognosis. Therefore, to reflect the distributions of various tissue types in the whole-slide image, the present application uses the tissue type distribution corresponding to the newly generated WSI features output by the WSI-level conditional diffusion model and introduces it as a constraint into the process of regenerating the image patch representation using the patch-level conditional diffusion model.
[0069] Step S130 can be divided into: Step S131, determining the tissue types included in the predicted noise tissue type distribution; Step S132, obtaining the conditional input of the pre-trained second conditional diffusion model based on the category embedding vectors corresponding to the determined tissue types and the predicted noise WSI representation, and respectively inputting each image patch representation into the pre-trained second conditional diffusion model to output the corresponding predicted noise image patch representation.
[0070] In some embodiments of the present invention, determining the tissue types included in the predicted noise tissue type distribution includes: determining the tissue types with non-zero values of tissue type from the predicted noise tissue type distribution. For example, in the predicted noise tissue type distribution output by the pre-trained first conditional diffusion model, there may be only 5 tissue categories with non-zero values, so the tissue types determined in Step S131 are also only these 5 types.
[0071] In this application, both the tissue type distribution corresponding to the corrected WSI and the predicted noise tissue type distribution output by the WSI-level conditional diffusion model can be used to reflect the proportion of the number of image patches of each tissue type in the WSI. Its abscissa is the tissue type, and the ordinate is the proportion of the number of image patches of this tissue type.
[0072] In the patch-level conditional diffusion model, in order to ensure the consistency between image patches, the "missing" WSI representation generated by the WSI-level conditional diffusion model is used to impose constraints on the denoising process of the patch-level conditional diffusion model; and, in order to introduce the prognostic information contained in the tissue type distribution at the WSI level, the category embedding vector of each image patch in the predicted noise tissue type distribution generated by the WSI-level conditional diffusion model is used as a condition. Therefore, based on the predicted noise WSI representation and the predicted noise tissue type distribution obtained in Step S120, obtaining the condition of the pre-trained second conditional diffusion model includes: determining the tissue types included in the predicted noise tissue type distribution, and obtaining the conditional input of the pre-trained second conditional diffusion model based on the category embedding vectors corresponding to the determined tissue types and the predicted noise WSI representation.
[0073] In some embodiments of the present invention, obtaining the conditional input of the pre-trained second conditional diffusion model based on the category embedding vectors corresponding to the determined tissue types and the predicted noise WSI representation includes: combining the predicted noise WSI representation and the category embedding vectors of the determined tissue types through the Hadamard product, and using the obtained combined constraint as the conditional input of the pre-trained second conditional diffusion model. That is, after obtaining the output result of the WSI model, it is combined through the Hadamard product to ensure that the generated image patches conform to the distribution of the target tissue type. The above-mentioned method of using the Hadamard product for constraint combination is only an example, and it can also be other matrix multiplications such as dot product or Kronecker product. The present invention is not limited thereto.
[0074] In step S132, the patch-level conditional diffusion model can regenerate the image patch representation (possibly with noise) based on the original image patch representation, which not only includes the image patch features of the patient's original WSI data, but also introduces the WSI representation newly generated by the WSI-level conditional diffusion model. This process is achieved by separately adding noise and denoising the image patch representations obtained in each step S110, and the image patch representation generated by the patch-level conditional diffusion model in this application contains information on the tissue type distribution. Specifically, since the patch-level conditional diffusion model is trained based on the conditional denoising diffusion probability model, its forward and backward processes can follow formulas similar to those of the WSI-level conditional diffusion model. However, the conditions in the backward process of the second conditional diffusion model pre-trained in this application and the first conditional diffusion model pre-trained are different. The patch-level conditional diffusion model denoises according to the predicted noise WSI representation and the class embedding vector sum, and the WSI-level conditional diffusion model denoises according to the patient's original WSI features. That is, the patch-level conditional diffusion model performs step-by-step noise addition and denoising processing, and regenerates the image patch representation through iteration, so as to ensure the consistency of the image patches and the effective retention of the tissue type distribution. Since in step S132, the multiple image patch representations obtained in step S110 are separately input into the patch-level conditional diffusion model, the corresponding predicted noise image patch representations output are corresponding to the input image patch representations.
[0075] In some embodiments of the present invention, the second conditional diffusion model pre-trained is obtained by the following method:
[0076] The WSI used to train the second conditional diffusion model is segmented into multiple non-overlapping image patches, and the corresponding image patch representations are extracted from each image patch; the WSI representation corresponding to the WSI used to train the second conditional diffusion model is used as the condition to input the first conditional diffusion model (the pre-trained first conditional diffusion model or the first conditional diffusion model under training), the conditional input of the initial second conditional diffusion model is obtained based on the output result of the first conditional diffusion model, and the extracted image patch representations are respectively input into the initial second conditional diffusion model; the output result of the initial second conditional diffusion model is compared with the input patch representation, so as to adjust the parameters of the initial second conditional diffusion model; iterative training is performed until the second training condition is met to obtain the pre-trained second conditional diffusion model.
[0077] The WSI for training the first conditional diffusion model and the WSI for training the second conditional diffusion model can be the WSI data in the same training set or the WSI data in different training sets. Different from the training data of the first conditional diffusion model, it is not necessary to distinguish between training WSI and calibration WSI in the training data of the second conditional diffusion model, but the training effect of the second conditional diffusion model is limited by the training effect of the first conditional diffusion model. In addition, in this application, the initial first conditional diffusion model and the initial second conditional diffusion model can be trained simultaneously, or the second conditional diffusion model can be trained after the first conditional diffusion model is trained (i.e., after obtaining the pre-trained first conditional diffusion model). This application does not specifically limit the training order.
[0078] As an example, the training objective of the patch-level conditional diffusion model is to minimize the difference between the noise predicted by the denoiser and the true noise. That is, the training objective of the initial second conditional diffusion model is to minimize the difference between the input image patch representation and the output image patch representation. Similarly, this application does not specifically limit the second training condition. In this application, the training methods of the initial first conditional diffusion model and the initial second conditional diffusion model can adopt existing training steps, and the present invention does not specifically limit this. In addition, this application does not specifically set the time steps of the first conditional diffusion model and the second conditional diffusion model, which can be determined according to the training objective.
[0079] This application does not need to be executed in the specific order mentioned above. It is only necessary to extract the WSI representation from the original WSI of the patient before executing step S120, and before executing step S130, obtain the image patch representation and the predicted noise WSI representation and the predicted noise tissue type distribution output by the pre-trained first conditional diffusion model in step S120.
[0080] The data augmentation method based on pathological whole slide images proposed in this application can provide new and uncaught biological information, which helps to expand the diversity of data. The image patch representation generated by the method of this application can provide more comprehensive and diverse training data for the survival prediction model, thereby improving the performance, robustness and generalization ability of the model in tasks such as survival prediction, and greatly enhancing the application potential of the survival prediction model in complex and noisy data.
[0081] When performing the survival prediction task, a group of patients P = {W 1 , W 2 ,..., W N} can be given, where N represents the number of patients. And, each patient p has a set of whole slide images, which can be represented as where n represents the number of WSIs associated with patient p, W pA set of whole-slide images for performing survival prediction tasks for patient p.
[0082] When training a survival prediction model, each patient is assigned a corresponding follow-up label (t p , δ p ), where t p represents the survival time or censoring time of patient p, and the binary indicator δ p can be used to represent the censoring situation of patient p, that is, whether the survival result of patient p is observed before the end of the study. If patient p dies or is lost to follow-up before the end of the survival study, censoring can be determined. δ p = 1 indicates censoring, and δ p = 0 indicates that the survival result of patient p has been observed. When using the survival prediction model for prediction, based on the set of whole-slide images W p of each patient in a group of patients, the corresponding patch representations of each patient can be obtained through the data augmentation method proposed in this application, and these patch representations are input into a pre-trained survival prediction model to predict the most likely survival time t p of each patient, so as to obtain the risk ranking of each patient in a given group of patients, for example, patient A risk > patient B risk > patient C risk. After obtaining the risk results output by the survival prediction model, the concordance index (C-index) can be used to evaluate the risk results, that is, to measure the consistency between the predicted risk and the actual survival time. In addition, although the more whole-slide images, the higher the survival prediction accuracy may be, when generating missing WSIs and performing survival prediction, all WSIs before patient censoring can be selected, or only some WSIs can be selected, and the present invention does not make specific limitations on this.
[0083] In a specific embodiment of the present invention, the steps of the data augmentation method based on pathological whole-slide images proposed in this application are as follows:
[0084] Step S01: Assume that a patient has a high-resolution pathological whole-slide image (WSI) containing billions of pixels. First, the WSI can be segmented into smaller patches through a sliding window strategy and the OTSU threshold algorithm. For each WSI, a pre-trained feature extractor can be used to map the WSI to a representation vector, and at the same time, each patch can also be mapped to the representation space to obtain a set of patch representations. In addition, during the training process, the HoverNet model can be used to classify the tissue type of each patch and assign a one-hot encoded tissue type label to each patch.
[0085] Step S02: The goal of the WSI-level conditional diffusion model is to generate a brand-new WSI representation as the consistency constraint for image patches and generate the corresponding predicted noise tissue type distribution. Since it is constructed based on the conditional denoising diffusion probability model, the WSI-level conditional diffusion model can generate a brand-new WSI representation and the predicted noise tissue type distribution through two processes: forward and backward.
[0086] Step S03: To ensure the consistency between image patches, constraints can be imposed on the image patch consistency condition (introducing the WSI representation generated in Step S02); and to maintain the tissue type distribution at the WSI level, the application category embedding vector of each image patch can be used as a condition. Under the constraint conditions generated by the WSI-level diffusion model, these constraints are combined through the Hadamard product. At this time, the patch-level conditional diffusion model can generate the image patch representation for simulating the "missing" WSI, and these image patch representations conform to a specific distribution. Specifically, a set of image patches can be first cropped from the given WSI and features can be extracted; since the patch-level diffusion model is also constructed based on the conditional denoising diffusion probability model, its forward and backward processes are similar to those of the WSI-level conditional diffusion model, and the key difference between the two lies in the different conditional inputs during the denoising process.
[0087] Aiming at the problem that the existing data augmentation methods cannot effectively solve the problems of data incompleteness and sample scarcity, the present invention proposes an innovative data augmentation method that can generate new WSI samples, rather than just modifying or expanding the existing samples. This method not only effectively fills the data gap, simulates the missing data patterns that may appear in the real world, provides more comprehensive and diverse training data for survival prediction, but also significantly improves the accuracy and robustness of survival prediction: First, the present invention combines patch-level and WSI-level feature extraction in the data preparation stage, which can accurately capture tissue information at different scales. By extracting image features at different scales, the ability to identify different tissue types is enhanced. Second, image patch classification is introduced in the process of generating image patch features, which ensures that the distribution of different tissue types in the WSI is effectively retained. And through the generation of image patch consistency constraints and tissue type distribution, the present invention further improves the adaptability and prediction ability of the model in a complex data environment. Finally, the present invention generates new WSI representations and image patch representations based on the conditional denoising diffusion probability model with high efficient generation ability and noise elimination ability. And in the reverse diffusion process of the patch-level model, multi-level conditional information is combined to accurately restore the "missing" WSI and ensure the quality of the image patch features and the correctness of the tissue type distribution, thereby further improving the accuracy of survival prediction.
[0088] In summary, the present invention provides an efficient and reliable data augmentation method by generating "missing" WSIs and combining multi-scale features and conditional denoising diffusion models. It not only solves the problem that existing methods cannot handle incomplete data, but also significantly improves the prediction accuracy and stability of the survival prediction model, contributing to the development of medical image analysis technology, especially in the applications of survival prediction and disease diagnosis.
[0089] Correspondingly, the present invention also provides a data augmentation system based on whole slide pathology images. The system includes a computer device, which includes a processor and a memory. The memory stores computer programs / instructions, and the processor is configured to execute the computer programs / instructions stored in the memory. When the computer programs / instructions are executed by the processor, the system implements the steps of the method described above.
[0090] An embodiment of the present invention also provides a computer-readable storage medium, on which computer programs / instructions are stored. When the computer programs / instructions are executed by a processor, the steps of the foregoing edge computing server deployment method are implemented. The computer-readable storage medium can be a tangible storage medium, such as random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, floppy disks, hard disks, removable storage disks, CD-ROMs, or any other form of storage medium known in the art.
[0091] Those of ordinary skill in the art should understand that the exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to implement in hardware or software depends on the specific application and design constraints of the technical solution. Professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, etc. When implemented in software, the elements of the present invention are programs or code segments used to execute the required tasks. The programs or code segments can be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave on a transmission medium or a communication link.
[0092] It should be clear that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present invention.
[0093] In the present invention, features described and / or illustrated for one embodiment can be used in the same or similar manner in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.
[0094] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A data enhancement method based on pathological full-slice images, characterized in that: The method comprises the following steps: Acquire at least one pathological whole-slice image WSI, segment the acquired WSI into a plurality of non-overlapping image blocks, and extract a corresponding WSI representation from the WSI, and extract a corresponding image block representation from each of the image blocks; Based on the extracted WSI representation, a conditional input of a pre-trained first conditional diffusion model is obtained, and a preset noise image is input into the pre-trained first conditional diffusion model to output a predicted noise WSI representation and a predicted noise tissue type distribution; Determine the tissue type contained in the predicted noise tissue type distribution, obtain the conditional input of the pre-trained second conditional diffusion model based on the category embedding vector corresponding to the determined tissue type and the predicted noise WSI representation, and input each image block representation into the pre-trained second conditional diffusion model respectively, and output the corresponding predicted noise image block representation.
2. The method according to claim 1, characterized in that: The step of obtaining a conditional input of a pre-trained first conditional diffusion model based on the extracted WSI representation includes: If the number of WSIs obtained is one, the extracted WSI representation is used as the conditional input of the pre-trained first conditional diffusion model; If the number of acquired WSIs is multiple, the extracted WSI representations are fused by average pooling to obtain an average WSI representation, and the average WSI representation is used as the conditional input of the pre-trained first conditional diffusion model.
3. The method according to claim 1, characterized in that: The conditional input of the pre-trained second conditional diffusion model is obtained based on the category embedding vector of the determined tissue type and the predicted noise WSI representation, including: merging the predicted noise WSI representation and the category embedding vector corresponding to the determined tissue type through the Hadamard product, and using the obtained merging constraint as the conditional input of the pre-trained second conditional diffusion model.
4. The method according to claim 1, characterized in that The pre-trained first conditional diffusion model is trained in the following way: Acquire a training WSI and a correction WSI belonging to the same patient, segment the correction WSI into a plurality of non-overlapping image blocks, and extract corresponding WSI representations from the training WSI and the correction WSI respectively; The cell recognition model is used to classify the tissue type of the image block corresponding to the corrected WSI, and the tissue type distribution corresponding to the corrected WSI is obtained; Based on the WSI representation corresponding to the training WSI, a conditional input of an initial first conditional diffusion model is obtained, a preset noise image is input into the initial first conditional diffusion model, and a training-prediction noise WSI representation and a training-prediction noise tissue type distribution are obtained as outputs, and the WSI representation and the tissue type distribution corresponding to the correction WSI are compared with the output results of the initial first conditional diffusion model, thereby adjusting the parameters of the initial first conditional diffusion model to obtain a pre-trained first conditional diffusion model; The pre-trained second conditional diffusion model is trained in the following way: Segmenting the WSI used for training the second conditional diffusion model into a plurality of non-overlapping image blocks, and extracting corresponding image block representations from each image block; The WSI corresponding to the WSI used to train the second conditional diffusion model is represented as a conditional input to the first conditional diffusion model, the conditional input of the initial second conditional diffusion model is obtained based on the output result of the first conditional diffusion model, and the extracted image block representations are respectively input into the initial second conditional diffusion model; The output result of the initial second conditional diffusion model is compared with the input image block representation, thereby adjusting the parameters of the initial second conditional diffusion model to obtain a pre-trained second conditional diffusion model.
5. The method according to claim 4, characterized in that The method of using the cell recognition model to classify the tissue type of the image block corresponding to the corrected WSI includes: A pre-trained cell recognition model is used to identify cells in the image block corresponding to the corrected WSI, the tissue type corresponding to each image block is determined according to a majority statistical strategy, and a category label for identifying the tissue type is assigned to each image block with a determined tissue type; wherein the category label is a category embedding vector of the tissue type corresponding to the image block, and the cell recognition model is trained based on a PanNuke dataset including tissue type classification.
6. The method according to claim 5, characterized in that The tissue type distribution corresponding to the corrected WSI is determined in the following manner: Based on the number of image blocks corresponding to each tissue type, each image block obtained by segmenting the corrected WSI is normalized to obtain the tissue type distribution of the corrected WSI.
7. The method according to claim 1, characterized in that The WSI representation and the image block representation both include morphological features and arrangement features of cells and tissues; as well as The determining of the tissue type contained in the predicted noise tissue type distribution comprises: Tissue types having non-zero values of tissue types are determined from the predicted noise tissue type distribution; wherein the tissue types include neoplastic, dead, inflammatory, non-neoplastic epithelial, connective tissue, or unlabeled.
8. The method according to claim 1, characterized in that: The method of dividing the acquired WSI into a plurality of non-overlapping image blocks includes: dividing the acquired WSI image into a plurality of non-overlapping image blocks by using a sliding window strategy and an Otsu algorithm.
9. A data enhancement system based on pathological full-slice images, comprising a processor, a memory, and a computer program / instruction stored in the memory, characterized in that: The processor is used to execute the computer program / instructions, and when the computer program / instructions are executed, the system implements the steps of the method according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Cited By
Pathological image slice feature enhancement method and system based on variance guidance
CN121526911A
A variance-guided pathological image slice feature enhancement method and system
CN121526911B