Histopathological Image Classification Method Based on a Quadruple Cascade Domain Adaptation Mechanism

Through the histopathological image classification method of the quadruple-cascade domain adaptation mechanism, the characteristics of the convolutional neural network are integrated to solve the problem of overfitting caused by the small number of labeled images, and the high accuracy and stability of histopathological image classification of breast cancer is achieved, which is suitable for clinical diagnosis.

CN115700794BActive Publication Date: 2025-07-25CHONGQING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211437766.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-15
Publication Date
2025-07-25
Estimated Expiration
2042-11-15

AI Technical Summary

Technical Problem

The prior art has problems of overfitting and poor generalization ability in histopathological image classification, especially because the number of labeled images is small and expensive, resulting in insufficient classification accuracy and robustness of convolutional neural networks, making it difficult to effectively utilize cellular structure information.

Method used

The quadruple cascade domain adaptation mechanism is adopted to construct a convolutional neural network, integrate features at different depths, and perform feature migration and cascade domain adaptation, including feature fusion, cluster envelope alignment and manifold fusion alignment, build an octa heterogeneous sample space, and use a small amount of labeled data to train the model.

Benefits of technology

It significantly improves the classification performance of breast cancer histopathological images, has high accuracy and stability, can process images of different formats and resolutions, has strong robustness and anti-overfit performance, and is suitable for clinical diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115700794B_ABST
    Figure CN115700794B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of intelligent diagnosis technology in biomedical information processing, and specifically discloses a method for classifying histopathological images based on a quadruple cascaded domain adaptation mechanism. By building a convolutional neural network for feature migration, integrating features extracted at different depths in the convolutional neural network, constructing an eight-fold heterogeneous sample space, and performing quadruple cascaded domain adaptation on the features in different sample spaces (including two clustering envelope alignments, manifold fusion alignment, and manifold clustering envelope domain adaptation). This method can significantly improve the classification performance by training the model with only a small amount of labeled data, meet the classification requirements of breast cancer histopathological images, and has strong robustness, self-adaptability, and anti-overfitting performance. This method can allow the input of images in different formats and resolutions, has high accuracy and stability, and shows great potential in clinical diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent diagnosis in biomedical information processing, and particularly to a method for classifying histopathological images based on a quadruple cascaded domain adaptation mechanism. Background Art

[0002] Breast cancer is currently one of the most serious diseases endangering women's health, with a prevalence rate exceeding 8%, ranking first among female malignancies. Since the exact cause of the disease is still unclear, early detection and diagnosis of breast cancer have become very important. Early diagnosis can not only reduce the treatment cost and duration of the disease, but also effectively improve the survival rate of patients. At present, the diagnosis of breast cancer mainly relies on the analysis of histopathological images by specialized pathologists. However, manual annotation of histopathological pictures requires pathologists to have considerable experience, and there are defects such as long time consumption and easy misdiagnosis. Therefore, computer-aided diagnosis (CAD) has become an effective tool for pathologists to shorten the diagnosis time and improve the diagnostic sensitivity and specificity.

[0003] Compared with traditional machine learning methods, convolutional neural networks (CNNs) can mine and learn representative and discriminative information from raw images. It has been widely used in the field of medical image diagnosis, especially in the field of histopathological images. However, the inherent characteristics of CNNs determine that their training requires a large number of labeled images. A low number of labeled samples will lead to a series of problems such as overfitting and poor generalization ability. Since the labeling of histopathological images is expensive and difficult to obtain, a transfer learning method that transfers information from the source domain to the target domain using limited labeled images has been proposed. Among them, the domain adaptation method can consider the similarity between the two domains and complete the classification task under different sample feature distributions.

[0004] Although recent research has made progress in cell classification algorithms, accurate classification is still a challenge due to problems such as irregularity, overlap, and uneven staining of cells in histopathological images. In addition, the performance of various algorithms is also limited by feature design and selection, and most only consider high-level features in the network, wasting a large amount of cell structure information and performing poorly. Summary of the Invention

[0005] The present invention provides a method for classifying histopathological images based on a quadruple cascaded domain adaptation mechanism, and the technical problem to be solved is: how to comprehensively utilize features of different depths in a convolutional neural network to solve the problem of histopathological image classification.

[0006] To solve the above technical problems, the present invention provides a method for classifying histopathological images based on a quadruple cascaded domain adaptation mechanism, including the steps:

[0007] S1. Construct a sample database consisting of normal and pathological histopathological images, and divide the sample database into source domain training samples and target domain training samples for training, as well as test samples for testing. All the source domain training samples are labeled, and a small part of the target domain training samples are labeled;

[0008] S2. Build a convolutional neural network and pre-train it using a publicly available image dataset to meet the feature extraction requirements, obtaining a pre-trained model;

[0009] S3. Transfer the pre-trained model as a feature extraction layer and reconstruct it with a quadruple cascaded domain adaptation mechanism and a new fully connected classifier into a classification model;

[0010] S4. Use the source domain training samples and the target domain training samples to train the classification model;

[0011] S5. Use the test samples to test the trained classification model to obtain prediction results.

[0012] Further, during the training process, the quadruple cascaded domain adaptation mechanism in step S3 is described as the following steps:

[0013] S31. Input the source domain training samples and the target domain training samples into the pre-trained model to extract low-order, middle-order, and high-order features and perform feature fusion to obtain a source domain fusion sample F S and a target domain fusion sample F T ;

[0014] S32. Perform fusion feature clustering on the source domain fusion sample F S to obtain a source domain fusion clustering envelope sample μ S , and align the source domain fusion clustering envelope sample μ S with the source domain fusion sample F S to obtain a source domain fusion alignment sample F′ S ;

[0015] S33. Perform manifold fusion alignment on the source domain fusion alignment sample F′ S and the target domain fusion sample F T , and then perform manifold projection to obtain a source domain fusion projection sample and a target domain fusion projection sample

[0016] S34. Perform fusion projection feature clustering on the source domain fusion projection sample to obtain a source domain fusion clustering projection sample and align the source domain fusion clustering projection sample Fuse and project the samples with the source domain Perform secondary clustering envelope alignment to obtain source domain fusion projection alignment samples

[0017] S35. Align the source domain fusion projection alignment samples With the target domain fusion projection samples Perform manifold clustering envelope alignment.

[0018] Furthermore, during the training process, the loss function of the classification model is expressed as:

[0019]

[0020]

[0021] Among them, L C (F S , μ) represents the loss of fusing and clustering the features of the source domain fusion sample F S in step S32, and μ represents the cluster center of this clustering; represents the loss of fusing and projecting the features of the source domain fusion projection sample in step S34, represents the cluster center of this clustering;

[0022] represents the loss of the quadruple cascaded domain adaptation mechanism, and L CEA (F S , μ S ) represents the loss of aligning the clustering envelope of the source domain fusion clustering envelope sample μ S with the source domain fusion sample F S in step 32, and L MCFA (F′ S , F T ) represents the loss of performing manifold fusion alignment on the source domain fusion alignment sample F′ S and the target domain fusion sample F T in step S33, represents the loss of performing secondary clustering envelope alignment on the source domain fusion clustering projection sample and the source domain fusion projection sample , represents the loss of performing manifold clustering envelope alignment on the source domain fusion projection alignment sample and the target domain fusion projection sample in step S35, and α represents the balance coefficient of the quadruple cascaded domain adaptation loss function;

[0023] Denotes the loss of manifold projection in manifold clustering envelope alignment in step S35;

[0024] L(X S ,Y S ,X Tl ,Y Tl ) = L(X S ,Y S ) + L(X Tl ,Y Tl ) represents the sum of the cross - entropy loss functions of the source domain and the target domain. Among them, the cross - entropy loss of the source domain The cross - entropy loss of the target domain c represents the number of picture categories, and respectively represent the i - th sample of the m - th category in the source domain and the target domain, and represent the true labels of the i - th sample of the m - th category in the source domain and the target domain, represents the probability that the source - domain sample is predicted as true, represents the probability that the target - domain sample is predicted as true.

[0025] Furthermore, in step S2, the K - means algorithm is used for fused - feature clustering. First, k sample points are selected as the initial central points of each cluster {μ1, μ2,..., μ k}, that is, the source - domain fused - clustering envelope samples, and then the following two steps are iteratively repeated:

[0026] 1) Calculate the distances between all sample points and the central points of each cluster, and then assign the sample points to the nearest cluster c i :

[0027]

[0028] where, F i represents the i - th feature sample point in the generated fused - sample feature F, and μ j represents the j - th initial central point;

[0029] 2) Recalculate the cluster center μ j :

[0030]

[0031] where, is the number of samples in cluster c i ;

[0032] When the maximum number of iterations is reached or there is no change in sample assignment, the clustering ends;

[0033] Finally, cluster the source domain fusion samples into k clusters \(c_1, c_2, \cdots, c_k\), k while ensuring that the clustering loss function \(L\) C is minimized:

[0034]

[0035] where \(\|\cdot\|\) represents the norm.

[0036] Furthermore, in the step S32, the loss \(L\) CEA \((F S , \mu S )\) is expressed as:

[0037]

[0038] where and represent a pair of random sample features in the source domain fusion sample \(F\) S , represents a pair of random sample features in the source domain fusion clustering envelope sample \(\mu\) S , \(n\) S represents the number of the source domain fusion samples \(F\) S , \(n\) μ represents the number of the clustering centers of the source domain fusion clustering envelope samples \(\mu\) S , and \(\Theta(\cdot)\) represents the Gaussian kernel function, which is expressed as:

[0039]

[0040] where \(x\) and \(y\) represent two samples, and \(\sigma\) represents the local action range parameter.

[0041] Furthermore, in the step S33, the loss \(L\) MCFA \((F' S , F T )\) is expressed as:

[0042]

[0043] where and represent a pair of random sample features in the source domain fusion alignment sample \(F'\) S , and represent a pair of random sample features in the target domain fusion sample \(F\) T , \(n\) T represents the number of the target domain fusion samples \(F\) T .

[0044] Furthermore, in the step S34, the loss Expressed as:

[0045]

[0046] The source domain fusion clustering projection samples and the source domain fusion projection samples are subjected to secondary clustering envelope alignment, specifically expressed as:

[0047]

[0048] Among them, represents a pair of random sample features in the source domain fusion projection samples ; represents a pair of random sample features in the source domain fusion clustering projection samples ; represents the number of clustering centers of the source domain fusion clustering projection samples .

[0049] Furthermore, in the step S35, it is expressed as:

[0050]

[0051] Among them, and represent a pair of sample features in the source domain fusion projection alignment samples ; represents a pair of adjacent sample features in the target domain fusion projection samples ; W MP represents the manifold projection matrix, is the diagonal matrix of W, and L = D - W is defined as the difference matrix between D and W, represents the square of the second norm, Tr(·) represents the trace of the matrix, and W represents the neighbor matrix between samples in the original space, expressed as:

[0052] F i and F j are two fusion sample features, σ is the local action range parameter, and n represents the total number of features.

[0053] Furthermore, in the step S35, it is expressed as:

[0054]

[0055] Among them, represents a pair of random sample features in the source domain fusion projection alignment samples .

[0056] Further, in the step S1, the histopathological images are breast cancer histopathological images in different formats and resolutions;

[0057] In the step S2, the pre-trained model uses the VGG model, and the publicly available image dataset uses the ImageNet dataset;

[0058] In the step S3, the fully connected classifier includes a fully connected layer, a batch normalization layer, and a dropout layer; the batch normalization layer is used to implement batch normalization so that its output has a mean of 0 and a variance of 1; the dropout layer is used to reduce the number of neurons to prevent overfitting;

[0059] In the step S5, the classification results output by the classification model are either normal cell image labels or abnormal condition image labels.

[0060] The histopathological image classification method based on the quadruple cascaded domain adaptation mechanism provided by the present invention builds a convolutional neural network for feature migration, synthesizes features extracted at different depths in the convolutional neural network, constructs an eight-fold heterogeneous sample space (including a source domain fusion sample space and a target domain fusion sample space, a source domain fusion clustering envelope sample space, a source domain fusion projection sample space and a target domain fusion projection sample space, a source domain fusion projection clustering envelope sample space, a source domain fusion projection alignment sample space, and a fusion projection alignment sample space), and performs quadruple cascaded domain adaptation on the features in different sample spaces (including two clustering envelope alignments, manifold fusion alignment, and manifold clustering envelope domain adaptation). This method can significantly improve the classification performance by training the model with only a small amount of labeled data, meet the classification requirements of breast cancer histopathological images, and has strong robustness, self-adaptability, and anti-overfitting performance. This method can allow images of different formats and resolutions to be input, has high accuracy and stability, and shows great potential in clinical diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 is a flowchart of the histopathological image classification method based on the quadruple cascaded domain adaptation mechanism provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0062] The following specifically illustrates the embodiments of the present invention in conjunction with the drawings. The given embodiments are only for illustrative purposes and should not be construed as limiting the present invention. The drawings are only for reference and illustration and do not constitute a limitation on the scope of the patent protection of the present invention, because many changes can be made to the present invention without departing from the spirit and scope of the present invention.

[0063] The histopathological image classification method based on a quadruple cascaded domain adaptation mechanism provided by an embodiment of the present invention, as Figure 1 shown, in this embodiment, includes the steps:

[0064] S1. Construct a sample database composed of normal and pathological histopathological images, and divide the sample database into source domain training samples and target domain training samples for training, and test samples for testing. All the source domain training samples are labeled, and a small part of the target domain training samples are labeled;

[0065] S2. Build a convolutional neural network and pre-train it using a publicly available image dataset to meet the feature extraction requirements, and obtain a pre-trained model;

[0066] S3. Transfer the pre-trained model as a feature extraction layer, and reconstruct it with a quadruple cascaded domain adaptation mechanism and a new fully connected classifier into a classification model;

[0067] S4. Use the source domain training samples and the target domain training samples to train the classification model;

[0068] S5. Use the test samples to test the trained classification model to obtain prediction results.

[0069] In step S1:

[0070] The source domain and target domain samples for training and testing in the sample database are normal and cancer breast histopathological images, which are divided into sampling blocks and subjected to a series of image preprocessing operations;

[0071] In the sample database, all the source domain training samples are labeled, and a small part of the target domain training samples are labeled;

[0072] The histopathological images in the sample database can accept RGB color images in JPEG, PNG or TIF formats with different resolutions.

[0073] In step S2:

[0074] The pre-trained model adopts a convolutional neural network model, which can include various convolutional neural models such as the common VGG model, GoogleNet model, ResNet, etc. The pre-trained model in this embodiment is defined as BreNet, which is composed of five convolutional modules and one fully connected module. Each convolutional module contains two convolutional layers and one pooling layer. The fully connected module contains three fully connected layers. Each output feature map of the convolutional layer is obtained by convolving multiple input feature maps and kernels. The pooling layer performs pooling on the output of the previous layer through a kernel function. The fully connected layer connects all neurons through weights.

[0075] In this embodiment, the constructed neural network model is pre-trained using the ImageNet dataset to meet the feature extraction requirements, and is reconstructed with a quadruple cascaded domain adaptation mechanism and a new fully connected classifier into a complete classification model for breast cancer histopathological image classification tasks.

[0076] In step S3, the fully connected classifier includes a fully connected layer, a batch normalization layer, and a dropout layer; the batch normalization layer is used to implement batch normalization so that its output has a mean of 0 and a variance of 1; the dropout layer is used to reduce the number of neurons to prevent overfitting.

[0077] In step S3, the quadruple cascaded domain adaptation mechanism is described as the following steps:

[0078] S31. Input the source domain training samples and the target domain training samples into the pre-trained model to extract low-order, middle-order, and high-order features and perform feature fusion to obtain the source domain fusion sample F S and the target domain fusion sample F T ;

[0079] S32. Perform fusion feature clustering on the source domain fusion sample F S to obtain the source domain fusion clustering envelope sample μ S , and align the source domain fusion clustering envelope sample μ S with the source domain fusion sample F S to obtain the source domain fusion alignment sample F′ S ;

[0080] S33. Perform manifold fusion alignment on the source domain fusion alignment sample F′ S and the target domain fusion sample F T , and then perform manifold projection to obtain the source domain fusion projection sample and the target domain fusion projection sample

[0081] S34. Perform fusion projection feature clustering on the source domain fusion projection sample to obtain the source domain fusion clustering projection sample and align the source domain fusion clustering projection sample with the source domain fusion projection sample to obtain the source domain fusion projection alignment sample

[0082] S35. Align the source domain fusion projection alignment sample with the target domain fusion projection sample for manifold clustering envelope alignment.

[0083] The so-called quadruple cascaded domain adaptation includes two clustering envelope alignments, manifold fusion alignment, and manifold clustering envelope domain adaptation.

[0084] In step S31, by cascading the average pooling features downsampled from different layers of BreNet, multi-layer feature fusion is used to enrich the classification information volume:

[0085]

[0086] Among them, represents the concatenation operation. GAP(·) represents the global average pooling operation. F represents the generated fused sample features. Ω represents different feature depths. N l represents the number of feature channels after the pooling operation. is the feature output of the l-th layer of the network, expressed as:

[0087]

[0088] Among them, f(·) is a non-linear activation function. is the weight matrix of the l-th layer. is the linear bias matrix.

[0089] In this step S32, the K-means algorithm is used for unsupervised fusion feature clustering. First, k sample points are selected as the initial central points of each cluster {μ1, μ2,..., μ k}, that is, the source domain fusion clustering envelope samples, and then the following two steps are iteratively repeated:

[0090] 1) Calculate the distances between all sample points and the central points of each cluster, and then assign the sample points to the nearest cluster c i as follows:

[0091]

[0092] Among them, F i represents the i-th feature sample point in the generated fused sample features F, and μ j represents the j-th initial central point;

[0093] 2) Recalculate the cluster center μ j as follows:

[0094]

[0095] Among them, is the number of samples in cluster c i ;

[0096] When the maximum number of iterations is reached or there is no change in sample assignment, the clustering ends;

[0097] Finally, the source domain fusion samples are clustered into k clusters c1, c2,..., c k , while ensuring that the clustering loss function L CMinimization:

[0098]

[0099] where ||·|| represents the norm.

[0100] In step S32, the Clustering Envelope Alignment (CEA) criterion is used to measure the source domain fused clustering envelope sample space and the source domain fused sample space to obtain the source domain fused alignment sample F'. S :

[0101]

[0102] where and represent a pair of random sample features in the source domain fused sample F S in the source domain fused sample F, represents a pair of random sample features in the source domain fused clustering envelope sample μ S in the source domain fused clustering envelope sample μ, n S represents the number of the source domain fused samples F S n μ represents the number of the cluster centers of the source domain fused clustering envelope samples μ S The Θ(·) represents the Gaussian kernel function, which is expressed as:

[0103]

[0104] where x and y represent two samples, and σ represents the local action range parameter.

[0105] In step S33, the Manifold Clustering Feature Fusion Alignment (MCFA) criterion (manifold fusion alignment) is used to measure the source domain fused alignment sample space and the target domain fused sample space:

[0106]

[0107] where and represent a pair of random sample features in the source domain fused alignment sample F' S in the source domain fused alignment sample F', and represent a pair of random sample features in the target domain fused sample F T in the target domain fused sample F, n T represents the number of the target domain fused samples F T The number of.

[0108] The fused sample features are projected onto a manifold to reduce the redundancy of the fused features. The fused projection sample features in the low-dimensional space can be expressed as:

[0109]

[0110] Among them, f MP (·) is a monotonically increasing activation function, which is used to ensure that the distances between features before and after projection are proportional. W MP is the manifold projection matrix. Thus, the source domain fusion projection samples and the target domain fusion projection samples

[0111] In step S34, the source domain fusion projection samples are obtained as the source domain fusion clustering projection samples after K-means clustering. Its clustering loss function L C2 is expressed as:

[0112]

[0113] Then, the manifold reconstruction method is used to further reconstruct the common subspace to minimize the divergence of the manifold projection fusion feature distributions between the source domain and the target domain. The source domain fusion clustering projection samples and the source domain fusion projection samples are subjected to quadratic clustering envelope alignment to obtain the source domain fusion projection alignment samples

[0114]

[0115] Among them, represents a pair of random sample features in the source domain fusion projection samples , represents a pair of random sample features in the source domain fusion clustering projection samples , represents the number of clustering centers of the source domain fusion clustering projection samples .

[0116] In step S5, the source domain fusion projection alignment samples and the target domain fusion projection samples are subjected to manifold clustering envelope alignment, which is expressed as:

[0117]

[0118] Among them, represents a pair of random sample features in the source domain fusion projection alignment samples .

[0119] In step S35, the manifold projection regularization term is expressed as:

[0120]

[0121] Among them, and represent a pair of sample features in the source domain fusion projection alignment samples . represents adjacent pair of sample features in the target domain fusion projection samples W MP represents the manifold projection matrix. is the diagonal matrix of W, and L = D - W is defined as the difference matrix between D and W. represents the square of the second norm, Tr(·) represents the trace of the matrix, and W represents the neighbor matrix between samples in the original space, expressed as:

[0122] F i and F j are two fused sample features, σ is the local action range parameter, and n represents the total number of features. During the training process, the loss function of the classification model is expressed as:

[0123]

[0124]

[0125] Among them, L C (F S , μ) represents the loss of fusing feature clustering for the source domain fusion sample F S in step S32, and μ represents the cluster center of this clustering; represents the loss of fusing projection feature clustering for the source domain fusion projection sample in step S34, represents the cluster center of this clustering;

[0126]

[0127] Loss of the re - cascaded domain adaptation mechanism, L CEA (F S , μ S ) represents the loss of clustering envelope alignment between the source domain fusion clustering envelope sample μ S and the source domain fusion sample F S in step 32, L MCFA (F′ S , F T ) represents the loss of manifold fusion alignment between the source domain fusion alignment sample F′ S and the target domain fusion sample F T in step S33, represents the source domain fusion clustering projection sample and the source domain fusion projection sample The loss for secondary clustering envelope alignment represents the loss of aligning the source domain fusion projection alignment samples in step S35 with the target domain fusion projection samples The loss of performing manifold clustering envelope alignment, where α represents the balance coefficient of the quadruple cascaded domain adaptation loss function ;

[0128] represents the loss of the manifold projection in performing manifold clustering envelope alignment in step S35;

[0129] L(X S ,Y S ,X Tl ,Y Tl ) = L(X S ,Y S ) + L(X Tl ,Y Tl ) represents the sum of the cross-entropy loss functions of the source domain and the target domain. Among them, the cross-entropy loss of the source domain The cross-entropy loss of the target domain c represents the number of picture categories, and respectively represent the i-th sample of the m-th category in the source domain and the target domain, and represent the true labels of the i-th sample of the m-th category in the source domain and the target domain, represents the probability that the source domain sample is predicted to be true, represents the probability that the target domain sample is predicted to be true. Finally, the picture-level classification result is obtained by synthesizing all the classification results at the sampling block level under a breast cancer tissue pathology image. Assume that each detection sample x is cut into n sampling blocks, and the network output of each sampling block is s. Then the classification result of this image is:

[0130]

[0131] where ||s ij || represents the probability that the i-th sampling block belongs to the j-th category, represents the probability that all sampling blocks belong to the j-th category. When this value is the largest, the image will be predicted as the j-th category.

[0132] In specific implementation, three breast cancer histopathological image datasets are used: two public datasets and one private dataset. The main part of the experiment is conducted on the two public datasets. Finally, the private dataset is used to verify the robustness of the method. The public dataset BreakHis provides 7909 breast tissue sections with a resolution of 700×460. The pathological images are taken at four magnifications of 40×, 100×, 200×, and 400× and are annotated by pathologists. Another public database ICIAR-2018 contains 400 breast biopsy images with a resolution of 2048×1536. According to the main cancer type in each image, the microscopic images are labeled as normal, benign, carcinoma in situ, or invasive carcinoma. To meet the requirements of classification, the four categories are combined into two categories: benign and malignant. The private dataset contains 134 breast histopathological images with a resolution of 512×512. It includes normal cells, early-stage cancer cells, and malignant cells, which are also combined into two categories: benign and malignant. Experimental results: an F1-score of over 94% and an accuracy of over 92%.

[0133] In summary, the histopathological image classification method based on the quadruple cascaded domain adaptation mechanism provided by the embodiments of the present invention builds a convolutional neural network for feature migration, synthesizes the features extracted at different depths in the convolutional neural network, constructs an octuple heterogeneous sample space (including the source domain fusion sample space and the target domain fusion sample space, the source domain fusion clustering envelope sample space, the source domain fusion projection sample space and the target domain fusion projection sample space, the source domain fusion projection clustering envelope sample space, the source domain fusion projection alignment sample space, and the fusion projection alignment sample space), and performs quadruple cascaded domain adaptation on the features in different sample spaces (including two clustering envelope alignments, manifold fusion alignment, and manifold clustering envelope domain adaptation). This method can significantly improve the classification performance by only training the model with a small amount of labeled data, meet the classification requirements of breast cancer histopathological images, and has strong robustness, self-adaptability, and anti-overfitting performance. This method can allow the input of images in different formats and resolutions, and has high accuracy and stability. The method is tested on three breast cancer histopathological image datasets, and the experimental results (an F1-score of over 94% and an accuracy of over 92%) confirm the effectiveness and robustness of the method, showing great potential in clinical diagnosis.

[0134] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A method for classifying histopathological images based on a quadruple cascaded domain adaptation mechanism, characterized in that Including the steps: S1. Construct a sample database composed of normal and pathological histopathological images, and divide the sample database into source domain training samples and target domain training samples for training, and test samples for testing. All the source domain training samples are labeled, and a small part of the target domain training samples are labeled; S2. Build a convolutional neural network and pre-train it with a publicly available image dataset to meet the feature extraction requirements, obtaining a pre-trained model; S3. Transfer the pre-trained model as a feature extraction layer and reconstruct it with a quadruple cascaded domain adaptation mechanism and a new fully connected classifier into a classification model; S4. Use the source domain training samples and the target domain training samples to train the classification model; S5. Use the test samples to test the trained classification model to obtain prediction results; During the training process, the quadruple cascaded domain adaptation mechanism in step S3 is described as the steps: S31. Input the source domain training samples and the target domain training samples into the pre-trained model to extract low-order, medium-order, and high-order features and perform feature fusion to obtain the source domain fusion sample F S and the target domain fusion sample F T ; S32. Perform fusion feature clustering on the source domain fusion sample F S to obtain the source domain fusion clustering envelope sample μ S , and align the clustering envelope of the source domain fusion clustering envelope sample μ S with the source domain fusion sample F S to obtain the source domain fusion alignment sample F S '; S33. Perform manifold fusion alignment on the source domain fusion alignment sample F S ′ and the target domain fusion sample F T Then perform manifold projection to obtain the source domain fusion projection sample and the target domain fusion projection sample S34. Perform fusion projection on the source domain fusion projection samples to obtain source domain fusion clustering projection samples through fusion projection feature clustering and perform secondary clustering envelope alignment on the source domain fusion clustering projection samples and the source domain fusion projection samples to obtain source domain fusion projection alignment samples S35. Align the source domain fusion projection alignment samples with the target domain fusion projection samples to perform manifold clustering envelope alignment.

2. The histopathological image classification method based on a quadruple cascaded domain adaptation mechanism according to claim 1, wherein During the training process, the loss function of the classification model is expressed as: Among them, L C (F S , μ) represents the loss of fusing feature clustering for the source domain fusion sample F S in step S32, and μ represents the cluster center of this clustering; represents the loss of fusing projection feature clustering for the source domain fusion projection sample in step S34, and represents the cluster center of this clustering; Denote the loss of the quadruple cascaded domain adaptation mechanism as \(L\) CEA (F S , \(\mu\) S ) denotes the loss of clustering envelope alignment between the source domain fusion clustering envelope sample \(\mu\) S and the source domain fusion sample \(F\) S in step 32, \(L\) MCFA (F S ', \(F\) T ) denotes the loss of manifold fusion alignment between the source domain fusion alignment sample \(F\) S ' and the target domain fusion sample \(F\) T in step S33, Denote the loss of secondary clustering envelope alignment between the source domain fusion clustering projection sample and the source domain fusion projection sample as Denote the loss of manifold clustering envelope alignment between the source domain fusion projection alignment sample and the target domain fusion projection sample in step S35, \(\alpha\) denotes the balance coefficient of the quadruple cascaded domain adaptation loss function ; Represents the loss of the manifold projection in the manifold clustering envelope alignment in step S35; Denote the sum of the cross - entropy loss functions of the source domain and the target domain, where the cross - entropy loss of the source domain The cross - entropy loss of the target domain $c$ represents the number of image categories, and $x_{s,m}^i$ and $x_{t,m}^i$ respectively represent the $i$-th sample of the $m$-th category in the source domain and the target domain, and $y_{s,m}^i$ and $y_{t,m}^i$ represent the true labels of the $i$-th sample of the $m$-th category in the source domain and the target domain, $p(x_{s,m}^i)$ represents the probability that the source domain sample is predicted as true, $p(x_{t,m}^i)$ represents the probability that the target domain sample is predicted as true.

3. The histopathological image classification method based on a quadruple cascaded domain adaptation mechanism according to claim 2, wherein In step S2, the K-means algorithm is used for fused feature clustering. First, k sample points are selected as the initial center points of each cluster {μ1, μ2,..., μ k}, that is, the source domain fused clustering envelope samples, and then the following two steps are iteratively repeated: 1) Calculate the distances between all sample points and each cluster center, and then assign the sample points to the nearest cluster c i where: Among them, F i represents the i-th feature sample point in the generated fused sample feature F, and μ j represents the j-th initial center point; 2) Recalculate the cluster center μ j : Among them, is the number of samples of cluster c i ; The clustering ends when the maximum number of iterations is reached or there is no change in the sample assignment; Finally, cluster the source domain fusion samples into k clusters c1, c2,..., c k , while ensuring that the clustering loss function is minimized: Where, ||·|| represents the norm.

4. The histopathological image classification method based on a quadruple cascaded domain adaptation mechanism according to claim 3, characterized in that In the step S32, the loss L CEA (F S , μ S ) is expressed as: Among them, and represent a pair of random sample features in the source domain fusion sample F S ; represents a pair of random sample features in the source domain fusion clustering envelope sample μ S ; n S represents the number of the source domain fusion samples F S ; n μ represents the number of the clustering centers of the source domain fusion clustering envelope samples μ S , and Θ(·) represents the Gaussian kernel function.

5. The histopathological image classification method based on the quadruple cascaded domain adaptation mechanism according to claim 4, wherein In the step S33, the loss L MCFA (F S ′,F T ) is expressed as: Among them, and represent a pair of random sample features in the source domain fusion alignment sample F S ′, and represent a pair of random sample features in the target domain fusion sample F T , and n T represents the number of the target domain fusion samples F T .

6. The histopathological image classification method based on a quadruple cascaded domain adaptation mechanism according to claim 5, wherein In the step S34, the loss is expressed as: Fuse and cluster the projection samples of the source domain Fuse the projection samples of the source domain Perform secondary clustering envelope alignment, specifically expressed as: Among them, represents a pair of random sample features in the source domain fusion projection sample ; represents a pair of random sample features in the source domain fusion clustering projection sample ; represents the number of cluster centers of the source domain fusion clustering projection sample .

7. The histopathological image classification method based on a quadruple cascaded domain adaptation mechanism according to claim 5, characterized in that, In the step S35, It is expressed as: Among them, and represent a pair of sample features in the source domain fusion projection alignment samples ; represents a pair of adjacent sample features in the target domain fusion projection samples ; W MP represents the manifold projection matrix, and W represents the neighbor matrix between samples in the original space. is the diagonal matrix of W, and L = D - W is defined as the difference matrix between D and W. represents the square of the second norm, and Tr(·) represents the trace of the matrix.

8. The histopathological image classification method based on a quadruple cascaded domain adaptation mechanism according to claim 7, wherein In the step S35, It is expressed as: Among them, represents a pair of random sample features in the source domain fusion projection alignment sample .

9. The histopathological image classification method based on the quadruple cascaded domain adaptation mechanism according to any one of claims 1 to 8, characterized in that: In step S1, the histopathological images are breast cancer histopathological images in different formats and resolutions; In step S2, the pre-trained model uses the VGG model, and the publicly available image dataset uses the ImageNet dataset; In step S3, the fully connected classifier includes a fully connected layer, a batch normalization layer and a dropout layer; the batch normalization layer is used to implement batch normalization so that its output has a mean of 0 and a variance of 1; the dropout layer is used to reduce the number of neurons to prevent overfitting; In step S5, the classification results output by the classification model are two cases: a normal cell image label or an abnormal image label.