A weakly supervised nucleus classification method based on domain adaptation
Patent Information
- Application Number
- CN202411324122.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-23
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-09-23
AI Technical Summary
本发明与现有的技术相比具有明显的进步和优势,可以有效的解决现有细胞核分类技术中标注数据少,数据集之间差异大等诸多问题,可以同时达到更加准确完善的细胞核分类效果
[0005]为了实现上述目的,本发明提供了一种基于域适应的弱监督细胞核分类方法,包括以下步骤:
Smart Images

Figure CN119323783B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of cell nucleus segmentation and classification technology in medical images, specifically to a weakly supervised cell nucleus classification method based on domain adaptation. Background Technology
[0002] Digital pathology images are considered crucial for clinical diagnosis and treatment. High-resolution images created by scanning tissue sections provide clear information about tissue structure, allowing pathologists to clearly observe the internal structure of tissues and cells, thus improving the efficiency and accuracy of disease diagnosis. Changes in the morphology and structure of cell nuclei are important indicators of many diseases, and precise classification of cell nuclei helps pathologists determine tumor regions and types. Therefore, accurate segmentation and classification of cell nuclei in pathology images are of great significance for advancing precision medicine.
[0003] Because different types of cell nuclei may share similar morphological and structural features, and even nuclei of the same species can exhibit significant differences at different stages, classifying cell nuclei remains a major challenge. Traditional cell nucleus classification methods primarily extract morphological, textural, and topological features from within the cell nucleus and then utilize machine learning algorithms (Support Vector Machines, Random Forests, etc.) to predict cell nucleus classification. However, these methods rely excessively on the design of handcrafted features. If the features are not representative enough or the number of handcrafted features is insufficient to support the classifier's judgment, the accuracy of cell nucleus classification will be reduced, rendering it without clinical value. With the development and maturation of deep learning technology, convolutional neural network models, which have shown good classification performance in natural images, have gradually been applied to cell nucleus classification. HoVer-Net effectively solves the problem of overlapping cell nucleus boundaries by utilizing the horizontal and vertical distances from the kernel pixels to the centroid. It is designed with three branches: first, it separates kernel pixels from the background; then, it separates overlapping cell nuclei; and finally, it clusters the cell kernel pixels to predict the type. This is the first deep neural network capable of simultaneously segmenting and classifying cell nuclei in histopathological images. However, this method requires a large amount of data with cell nucleus segmentation and classification annotations for training, which is time-consuming and costly in practice. Moreover, the model's generalization ability needs further validation for cell nucleus classification tasks in unknown domains. In summary, existing cell nucleus classification work still has significant room for improvement in terms of data utilization and methodological innovation. Summary of the Invention
[0004] To address the problems existing in current technologies, the main objective of this invention is to propose a weakly supervised method for cell nucleus classification based on domain adaptation. The ultimate goal is to reduce the data discrepancy between the source and target domains, achieving the ability to simultaneously classify multiple cell nuclei from both domains. First, two widely used labeled pathological image cell nucleus datasets are selected as the source and target domains, respectively, in the domain adaptation method. Individual cells are cropped based on the label values of the two datasets. Then, a convolutional neural network is used to extract features from the individual cell images, and a fully connected network is used as a classifier to classify cell nuclei from the training data. Simultaneously, a weakly supervised learning method is employed to supervise the entire training process, thereby reducing the data discrepancy between the source and target domains. Compared with existing technologies, this invention has significant advancements and advantages, effectively solving many problems in current cell nucleus classification techniques, such as insufficient labeled data and large discrepancies between datasets, while simultaneously achieving more accurate and comprehensive cell nucleus classification results.
[0005] To achieve the above objectives, this invention provides a weakly supervised cell nucleus classification method based on domain adaptation, comprising the following steps:
[0006] Step 1) Perform cropping and preprocessing on the two cell nucleus datasets and construct corresponding new datasets;
[0007] Step 2) Extract features using a convolutional neural network based on the source domain dataset and perform high-dimensional mapping;
[0008] Step 3) Use a pre-trained network based on the source domain data to extract features and perform high-dimensional mapping on the target domain data;
[0009] Step 4) Align the target domain data with the source domain data and transfer labels in the high-dimensional mapping space using a Gaussian mixture model;
[0010] Step 5) Iterate through the process in Step 4) until all target domain data has been aligned and labels have been transferred;
[0011] Step 6) Use a classifier model to predict cell nucleus classification in the target domain data after data integration and label transfer;
[0012] Furthermore, the cropping and preprocessing of the cell nucleus dataset in step 1) includes:
[0013] The mask value is determined by the label for the input RGB cell nucleus image data;
[0014] The boundary contour of the cell nucleus is determined by the mask of a single cell nucleus, and the minimum bounding rectangle of the cell nucleus and the geometric moments of the contour are calculated accordingly.
[0015] Calculate the centroid coordinates of a single cell nucleus using moments;
[0016] The cell nucleus image data of the source and target domains are cropped using the centroid coordinates of the cell nucleus and the minimum bounding rectangle.
[0017] Furthermore, step 2) involves feature extraction and high-dimensional mapping based on source domain data, including:
[0018] Image preprocessing based on source domain data augmentation (translation, rotation, scaling);
[0019] The residual neural network ResNet is used to extract features from cell nucleus image data with labeled values in the source domain.
[0020] The generated feature maps are subjected to global average pooling, which reduces the spatial dimension of each feature map to a single value, forming a high-dimensional feature vector.
[0021] The classifier is used to make initial class predictions for source domain data based on cell features, to determine the discriminative power of cell features, and to evaluate the reliability of the cell classification model based on the classification loss.
[0022] Furthermore, step 3) involves feature extraction and high-dimensional mapping of the target domain data, including:
[0023] Image preprocessing based on data augmentation (translation, rotation, scaling) of target domain data;
[0024] Feature extraction of cell nucleus image data without label values in the target domain data is performed based on the pre-trained ResNet network.
[0025] The feature vectors of the high-dimensional mapping of the feature map are obtained by calculating the feature extraction.
[0026] Furthermore, step 4) includes data alignment and label transfer, which includes:
[0027] In a high-dimensional space, the known cell category space with labels in the source domain is moved away, while the target domain data space without labels is gathered together;
[0028] A Gaussian mixture model is established for source and target domain data to help find outlier samples in the target domain and improve the robustness of the system.
[0029] Compare the cell characteristics of known cell types in the source domain with the features calculated from the target domain;
[0030] Align the cell classification features of the source and target domains;
[0031] Pass the aligned source domain label to the target domain.
[0032] Furthermore, step 5) specifically includes the following steps:
[0033] The statistics do not yet include target domain data with cell labels;
[0034] Iterate through the above data pairs and the steps for transferring labels from the source domain to the target domain;
[0035] Gradually reduce the data difference between the source and target domains until all unknown types have cell labels.
[0036] Furthermore, step 6) specifically includes the following steps:
[0037] A fully connected layer that performs cell classification on source domain data is used as a classifier;
[0038] The pre-trained classifier is used to perform the final cell nucleus classification on the target domain data.
[0039] The design concept of this invention is as follows: First, two widely used labeled pathological image cell nucleus datasets are selected as the source and target domain data in the domain adaptation method, respectively. Individual cell nuclei are then cropped based on their label values to construct a new dataset containing individual cell nucleus images and corresponding cell type label values. Next, feature extraction is performed on the cell nucleus images in both the source and target domains. A convolutional neural network can be used as the encoder in this feature extraction process to establish a Gaussian mixture model for the extracted features, enabling data interaction and label transfer. During this process, multiple iterations are used to transfer source domain labels to target domain data, thereby reducing the differences between the source and target domain data. Finally, a classifier is used to classify cell nuclei across all data in the target domain, addressing the main difficulties in current cell nucleus classification tasks, such as scarce cell nucleus annotations and large data variability.
[0040] Compared with existing technologies, the beneficial effects of this invention are as follows: This invention's domain-adaptation-based weakly supervised cell nucleus classification method reduces data discrepancies between the source and target domains through domain adaptation, improving data integration. Weakly supervised learning reduces dependence on cell annotation labels, improving the accuracy of cell nucleus classification and better meeting clinical needs. Compared to other cell nucleus classification methods, it offers optimizations in hardware requirements, data label dependence, generalization ability, and overall classification task performance. Attached Figure Description
[0041] Figure 1 This is a flowchart illustrating the operation of the present invention.
[0042] Figure 2 This is a model diagram of the algorithm proposed in this invention;
[0043] Figure 3 This is a diagram illustrating data integration and label transfer. Detailed Implementation
[0044] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0045] This invention designs a weakly supervised cell nucleus classification method based on domain adaptation. First, individual cell nuclei are extracted, and feature extraction and high-dimensional vector mapping are performed using a convolutional neural network. Then, data alignment and label transfer are performed in the high-dimensional space to ensure that all data in the target domain have labeled values. Finally, a classifier is used to predict cell nucleus classification. The innovations of this invention include: 1) data integration based on domain adaptation, reducing the data difference between the source and target domains; 2) the use of weakly supervised learning to supervise the training process, reducing dependence on labeled data; and 3) constructing a new dataset by pruning individual cells, simplifying the model training process. Compared with existing cell nucleus classification methods, this method reduces computational complexity and the workload of doctors manually labeling cells, making it highly practical in clinical practice.
[0046] Implementation process:
[0047] A weakly supervised cell nucleus classification method based on domain adaptation, the specific steps of which are as follows:
[0048] S1: Perform cropping and preprocessing on the two cell nucleus datasets and construct corresponding new datasets;
[0049] In the input RGB cell nucleus image data, the mask value is determined by the label. The boundary contour of the cell nucleus is determined by the mask of a single cell nucleus. The minimum bounding rectangle of the cell nucleus and the geometric moments of the contour are calculated based on these. The centroid coordinates of a single cell nucleus are calculated using the moments. The cell nucleus image data of the source and target domains are cropped using the centroid coordinates and the minimum bounding rectangle of the cell nucleus.
[0050] S2: Extract features using a convolutional neural network based on the source domain dataset and perform high-dimensional mapping;
[0051] Image preprocessing (translation, rotation, scaling) is performed on the source domain data to augment the data. ResNet residual neural network is used to extract features from the cell nucleus image data with labeled values in the source domain. The generated feature maps are then subjected to average global pooling, reducing the spatial dimension of each feature map to a single value to form a high-dimensional feature vector. A classifier is used to predict the initial category of the source domain data based on cell features to determine the discriminative power of cell features. The reliability of the cell classification model is evaluated based on the classification loss.
[0052] S3: Feature extraction and high-dimensional mapping of target domain data based on a pre-trained feature extraction network;
[0053] Image preprocessing for data augmentation (translation, rotation, scaling) of target domain data involves using a pre-trained ResNet network in the source domain data to extract features from cell nucleus image data without label values in the target domain data, and then calculating the feature vector of the high-dimensional mapping of the feature map.
[0054] S4: Based on the Gaussian mixture model in the high-dimensional mapping space, the alignment and label transfer between the target domain and source domain data are realized;
[0055] In a high-dimensional space, the known cell categories with labels in the source domain are separated from each other, while the unknown categories in the target domain without labels are clustered together. A Gaussian mixture model is established for the source and target domain data to help find outlier samples in the target domain and improve the robustness of the system. The cell classification features of the known cell categories in the source domain are compared with the features calculated in the target domain to align the cell features of the source and target domains, and the aligned source domain labels are then passed to the target domain.
[0056] S5: The iterative process of data alignment and label transfer from the source domain to the target domain based on a high-dimensional mapping space;
[0057] For target domain data that does not yet have cell labels, iterate through the above data pairs and source domain to target domain label transfer steps to gradually reduce the data difference between the source domain and the target domain until all unknown types have cell labels.
[0058] S6: Use a classifier model to predict cell nucleus classification in the target domain data after data integration and label transfer;
[0059] After statistical analysis of the target domain data, including data alignment and label transfer, a classifier pre-trained in the source domain data is used to perform final cell classification on the target domain data.
[0060] Example:
[0061] S1: Data Preprocessing
[0062] PanNuke and CoNSeP were selected from common cell nuclear image datasets as the source and target domain datasets in this embodiment. In other use cases, they can be replaced with other cell nuclear image datasets according to actual needs.
[0063] The mask and cell type are determined based on the label information. In this embodiment, cell types include tumor cells, immune cells, epithelial cells, and inflammatory cells, which can be replaced with other cell types depending on the specific scenario. The boundary contour of the cell nucleus is determined based on the mask, and the geometric moments M of the cell nucleus contour are calculated based on these moments. The centroid coordinates of a single cell nucleus are then calculated using these moments. The calculation method is as follows:
[0064]
[0065] Among them, M 00 M is the zeroth moment, representing the sum of all pixel values in the image. 10 M is the first moment, representing the sum of the products of the x-coordinate and the pixel value for each pixel. 01 It is a first-order moment, representing the sum of the products of the y-coordinate and the pixel value of each pixel.
[0066] Based on the centroid coordinates and cell outline, individual cell nuclei are cropped according to the minimum bounding rectangle to construct a new dataset D containing individual cell nuclei and their corresponding label information. S and D T .
[0067] S2: Source Domain Data Feature Extraction and High-Dimensional Mapping
[0068] For source domain dataset D S Cell image I S Image enhancement (translation, rotation, scaling) is represented as I. enhanced In this embodiment, ResNet-50 is used to target I. enhanced Feature extraction yields feature F. In other scenarios, this can be replaced by neural networks such as DenseNet and EfficientNet for feature extraction.
[0069] F = ResNet(I) enhanced )
[0070] The generated features are reduced to a single value using average pooling, forming a high-dimensional feature vector. A fully connected layer is then used as a classifier to classify cells in the source domain data.
[0071] f = GAP(F)
[0072] Y S =FC(f)
[0073] Where GAP represents global average pooling, FC represents cell classification using a fully connected layer, and Y... S This indicates the results of cell classification.
[0074] S3: Feature extraction and high-dimensional mapping of target domain data
[0075] For the target domain dataset D T Cell image I T Image enhancement (translation, rotation, scaling) is represented as I. e ' nhancedThe ResNet-50 pre-trained network on the source domain data is used to extract features and map high-dimensional vectors to the target domain data, resulting in a high-dimensional vector f'.
[0076] S4: Data alignment and label transfer between source and destination domains
[0077] Use X s X represents the tagged cellular features in the source domain. T This represents cell features that are not labeled in the target domain, where (N s (Number of samples in the source domain) (N T (where the number of samples in the target domain is ), and a Gaussian mixture model is used to fit the features of the source domain and the features of the target domain:
[0078]
[0079] Where θ represents the parameters of the Gaussian mixture model, π represents the weights of the Gaussian distribution, and K... s Let μ represent the number of Gaussian distributions, and let Σ represent the mean and covariance matrices, respectively. This represents the probability density function.
[0080] Align the features of the source and target domains, and then pass the source domain labels to the target domain based on the aligned features. This is achieved through the following method:
[0081]
[0082] Where A and B are alignment matrices, used to map the features of the source and target domains to a common space, X s and X T Representing the characteristics of the source and target domains, This indicates the tag of the target domain after the tag is passed.
[0083] S5: Iteration process of S4
[0084] For target domain data that does not yet have cell labels, iterate through the above data pairs and source domain to target domain label transfer steps to gradually reduce the data difference between the source domain and the target domain until all unknown types have cell labels.
[0085] S6: Cell nuclear classification prediction
[0086] Calculate the feature vector of the target domain after all data alignment and label propagation are completed. The final classification of cell nuclei in the target domain is performed using a classifier trained on the source domain data.
[0087]
[0088] in This represents the final prediction for the target domain cell.
[0089] The above embodiments are merely preferred embodiments of the present invention and are not intended to limit the technical solutions of the present invention. Any technical solution that can be implemented based on the above embodiments without creative effort should be considered to fall within the scope of protection of the patent of the present invention.
Claims
1. A weakly supervised cell nucleus classification method based on domain adaptation, characterized in that, Includes the following steps: Step 1) Select two cell nucleus datasets from the cell nucleus image dataset, as the source domain dataset and the target domain dataset, respectively; crop and preprocess the two cell nucleus datasets, and construct a new source domain dataset. and target domain dataset ; The cropping and preprocessing of the cell nucleus dataset in step 1) is as follows: The mask value is determined by the label for the input RGB cell nucleus image data; The boundary contour of the cell nucleus is determined by the mask of a single cell nucleus, and the minimum bounding rectangle and geometric moments of the contour of the cell nucleus are calculated accordingly. Calculate the centroid coordinates of a single cell nucleus using moments; By cropping cell nucleus image data from the source and target domains using the centroid coordinates and minimum bounding rectangle of the cell nucleus, a new dataset containing individual cell nuclei and their corresponding label information is constructed. and ; Step 2) Based on the source domain dataset Convolutional neural networks are used to extract features from source domain data and perform high-dimensional mapping. The specific process of step 2) is as follows: Image preprocessing based on source domain data augmentation, where data augmentation includes translation, rotation, and scaling; The source dataset was processed using a ResNet residual neural network. Feature extraction was performed on cell nucleus image data containing labeled values; The generated feature maps are subjected to global average pooling, which reduces the spatial dimension of each feature map to a single value, forming a high-dimensional feature vector. The classifier is used to make initial class predictions for source domain data based on cell features, to determine the discriminative power of cell features, and to evaluate the reliability of cell classification based on classification loss. Step 3) Based on the target domain dataset Use a pre-trained network based on source domain data to perform feature extraction and high-dimensional mapping on target domain data; The specific process of step 3) is as follows: Image preprocessing based on data augmentation of target domain data, where data augmentation includes translation, rotation, and scaling; The ResNet pre-trained network based on source domain data is used on the target domain dataset. Feature extraction was performed on cell nucleus image data without label values. Calculate the feature vectors of the high-dimensional mapping of the feature map obtained by feature extraction; Step 4) Align the target domain data with the source domain data and transfer labels in the high-dimensional mapping space using a Gaussian mixture model; The specific process of step 4) is as follows: In a high-dimensional space, the known cell category space with labels in the source domain data is moved away, while the target domain data space without labels is clustered together; Gaussian mixture models are established for source and target domain data to remove cell nucleus image data in the target domain that differ from those in the source domain, thereby improving the robustness of the system. Compare the cell characteristics of known cell categories in the source domain data with the features calculated in the target domain data; Align the cell classification features of the source and target domains; The source domain label is passed to the target domain based on the aligned features; Step 5) Iterate through the process in Step 4) until all target domain data has been aligned and labels have been transferred; Step 6) Use a classifier model to predict cell nucleus classification on the target domain data after data alignment and label transfer.
2. The weakly supervised cell nucleus classification method based on domain adaptation according to claim 1, characterized in that, Step 5) includes the following specific steps: The statistics do not yet include target domain data with cell labels; Iterate through the above data pairs and the steps for transferring labels from the source domain to the target domain; Gradually reduce the data difference between the source and target domains until all unknown types have cell labels.
3. The weakly supervised cell nucleus classification method based on domain adaptation according to claim 1, characterized in that, Step 6) includes the following specific steps: A fully connected layer that performs cell classification on source domain data is used as a classifier; The pre-trained classifier is used to perform the final cell nucleus classification on the target domain data.
Citation Information
Patent Citations
Pseudo label loss unsupervised adversarial domain adaptive picture classification method based on Gaussian uniform mixture model
CN114492574A
Bone marrow cell classification method, system and equipment based on unsupervised domain adaptation and medium
CN118552951A