Breast cancer axillary lymph node metastasis pathological image classification method

CN122737601APending Publication Date: 2026-09-11SHANGHAI FIRST PEOPLES HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610883915.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-18
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

[0009]有鉴于现有技术的上述缺陷,本发明所要解决的技术问题是现有技术依然存在因特征提取尺度单一、难以同时捕捉这些多尺度的形态学特征,容易造成关键信息丢失,全局信息的利用不够充分而导致淋巴结中的微转移及孤立肿瘤细胞簇等微小病灶的识别能力有限,影响分类的准确性等问题

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122737601A_ABST
    Figure CN122737601A_ABST
Patent Text Reader

Abstract

This invention discloses a method for classifying pathological images of axillary lymph node metastasis in breast cancer, comprising the following steps: collecting whole-slice pathological images (WSI) of axillary lymph nodes in breast cancer, and preprocessing all images; inputting the preprocessed image patches into a feature extraction network for feature extraction to obtain feature vectors; constructing an improved classification network, including constructing several stacked dual attention modules; training and optimizing the improved classification network using training set data, and saving the best model. This invention provides a method for classifying pathological images of axillary lymph node metastasis in breast cancer, effectively extracting tissue and cell features at different scales in pathological images to enhance the model's ability to represent morphological diversity; utilizing global contextual information extracted at different levels using a Transformer network to improve the sensitivity to identify small lesions; and achieving a method for high-accuracy prediction of axillary lymph node metastasis in breast cancer using only image-level labels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical image processing technology, and in particular to a method for classifying pathological images of axillary lymph node metastasis in breast cancer. Background Technology

[0002] Breast cancer is one of the most common malignant tumors in women. Axillary lymph node metastasis is a key indicator for determining clinical stage, developing treatment plans (such as whether axillary lymph node dissection is necessary), and assessing patient prognosis. Pathological examination is the gold standard for determining lymph node metastasis; however, traditional pathological diagnosis relies on pathologists observing stained sections under a microscope, which has the following characteristics:

[0003] 1. Low efficiency and high subjectivity: Manual image reading is time-consuming and laborious, and the diagnostic results are affected by the doctor's experience and subjective state. Micrometastases (0.2mm-2mm) and isolated tumor cell clusters (<0.2mm) are easily missed or misdiagnosed.

[0004] 2. Huge data scale: Whole slice pathological images (WSI) have extremely high resolution (usually up to 105×105 pixels), which cannot be directly input into existing computer vision models for processing.

[0005] Currently, deep learning-based WSI classification primarily employs the Multiple Instance Learning (MIL) framework, treating a WSI as a "bag" and its segmented image patches as "instances" within the bag. Typical implementation schemes include:

[0006] 1. CNN-based feature extraction + pooling method: First, a Convolutional Neural Network (CNN) (such as ResNet) is used to extract features from each image patch. Then, max pooling or mean pooling is used to aggregate the features of all image patches into a single feature vector. Finally, this feature vector is used for classification. This method is simple, but it assumes that all image patches are independent and identically distributed, ignoring the spatial and contextual relationships between image patches.

[0007] 2. DSMIL (Dual-stream Multiple Instance Learning): This method uses a two-stream structure. One stream is used to find key instances (the most representative image patches), and the other stream is used to calculate the correlation between key instances and all other instances. Finally, the results of the two streams are fused, which takes into account the interaction between instances to some extent.

[0008] 3. TransMIL (Transformer based Multiple Instance Learning): This method is the first to introduce Transformer into the MIL framework. It models the long-range dependencies and spatial relationships between different image patches through a self-attention mechanism, effectively solving the problem of traditional methods ignoring contextual information. However, this method extracts only one scale and has difficulty capturing morphological features at multiple scales simultaneously, which can easily lead to the loss of key information. The standard Vision Transformer model usually only uses the class token output from the last layer for the final prediction, ignoring the rich texture and structural information at different levels of abstraction contained in the shallow layers of the Transformer network. This results in insufficient utilization of global information, which limits the ability to identify micro-metastases in lymph nodes and isolated tumor cell clusters, thus affecting the accuracy of classification. Summary of the Invention

[0009] In view of the aforementioned shortcomings of the prior art, the technical problem to be solved by the present invention is that the existing technology still suffers from problems such as limited ability to identify micro-metastases and isolated tumor cell clusters in lymph nodes due to the single feature extraction scale, difficulty in simultaneously capturing these multi-scale morphological features, easy loss of key information, and insufficient utilization of global information, thus affecting the accuracy of classification. The present invention provides a pathological image classification method for axillary lymph node metastasis in breast cancer, which effectively extracts tissue and cell features at different scales in pathological images to enhance the model's ability to represent morphological diversity; makes fuller use of global contextual information extracted by the Transformer network at different levels to improve the sensitivity to the identification of micro-lesions; and realizes a method for high-accuracy prediction of axillary lymph node metastasis in breast cancer without pixel-level annotation, using only image-level labels.

[0010] To achieve the above objectives, the present invention provides a method for classifying pathological images of axillary lymph node metastasis in breast cancer, comprising the following steps:

[0011] We collected whole-section pathological images of axillary lymph nodes from breast cancer using Western blotting (WSI). All images were preprocessed, divided into n image blocks, and then divided into training and testing sets.

[0012] The preprocessed image patch is input into a feature extraction network for feature extraction to obtain feature vectors;

[0013] An improved classification network is constructed, which includes building several stacked dual attention modules, linearly projecting the feature vectors in the dual attention modules, then performing depthwise separable convolutions on each module and then adding and fusing them. The fused features are then input into a standard Transformer encoding layer for self-attention calculation to obtain a class token. The class tokens are then concatenated and flattened into a one-dimensional vector, which is then input into the classification head of a fully connected network to obtain the probability that the WSI belongs to the transition / non-transition category.

[0014] The improved classification network is trained and optimized using the data from the training set, the best model is saved, and the best model is evaluated on the test set.

[0015] Furthermore, the collected whole-section pathological images of axillary lymph nodes from breast cancer were only annotated at the image level, without any pixel-level annotation information.

[0016] Furthermore, the collected pathological images of whole sections of axillary lymph nodes from breast cancer were divided into training and testing sets in a 7:3 ratio.

[0017] Furthermore, preprocessing of all images includes dividing the TIF format WSI into n 224×224 pixel JPG image patches and zero-padding boundary areas that are less than 224 pixels; and eliminating a large number of blank and useless noise areas in the WSI, including background areas and tissue fluid areas.

[0018] Furthermore, the preprocessed image patches are input into a feature extraction network for feature extraction to obtain feature vectors. Specifically, this involves constructing a SimCLR model based on ResNet18, performing self-supervised pre-training on all valid image patches to obtain a feature extractor. Then, all valid image patches of each WSI are input into the feature extractor to obtain an n×512 feature matrix.

[0019] Furthermore, the dual attention module includes multi-scale convolutional layers and self-attention layers. The multi-scale convolutional layers use K different sizes of convolutional kernels in parallel on the input feature map F, and each convolutional kernel outputs a feature map M. F And perform multi-scale feature fusion;

[0020] The self-attention layer inputs the multi-scale fused features output from the first layer into the self-attention mechanism of the Transformer, capturing the global contextual dependencies between them, thereby obtaining deep features.

[0021] Furthermore, the improved classification network comprises L stacked TransConv modules, with a learnable class token vector concatenated into the input sequence of each TransConv module;

[0022] All L class tokens are concatenated along the channel dimension to obtain a multi-level fusion vector, which is then linearly reduced in dimensionality.

[0023] Furthermore, the improved classification network is trained and optimized using the data from the training set, including training the improved classification network using a training set with image-level labels, optimizing the network parameters through the backpropagation algorithm, and minimizing the loss function between the predicted results and the true labels.

[0024] Once training is complete, the optimal model can be used to predict the axillary lymph node metastasis status of newly input WSI.

[0025] Furthermore, the loss function is set to cross-entropy loss.

[0026] Furthermore, when training and optimizing the improved classification network using the training set data, the learning rate is 1e-4, the batch size is 1, and the optimizer is AdamW.

[0027] Technical effect

[0028] This invention provides a pathological image classification method for axillary lymph node metastases in breast cancer. Through multi-scale convolutional layers in the Transformer module (TransConv), it can simultaneously capture morphological information at different scales in pathological images (from tiny cell nuclei to large tissue structures), resulting in stronger feature representation capabilities and greater sensitivity to the identification of micrometastases. By using Multi-Level Transformer (MLTT), it aggregates the classification tokens output from each layer of the network, effectively fusing multi-level global contextual information from low-level details to high-level semantics, avoiding information loss and improving the overall performance of the model.

[0029] Experimental results on the CAMELYON16 dataset show that the proposed method achieves a classification accuracy (ACC) of 88% and an AUC of 0.94 using only image-level annotations, which is significantly better than the comparison scheme that uses only a single-layer class token and a single attention, demonstrating its superior performance.

[0030] The following will further explain the concept, specific structure, and technical effects of the present invention in conjunction with the accompanying drawings, so as to fully understand the purpose, features, and effects of the present invention. Attached Figure Description

[0031] Figure 1 This is a flowchart illustrating a preferred embodiment of the present invention for classifying pathological images of axillary lymph node metastasis in breast cancer.

[0032] Figure 2This is a schematic diagram of a preprocessing method for classifying pathological images of axillary lymph node metastasis in breast cancer according to a preferred embodiment of the present invention.

[0033] Figure 3 This is a schematic diagram of the structure of a convolutional-transformer module (TransConv) in a pathological image classification method for axillary lymph node metastasis of breast cancer according to a preferred embodiment of the present invention;

[0034] Figure 4 This is a schematic diagram of a multi-level token transformer (MLTT) method for classifying pathological images of axillary lymph node metastasis in breast cancer according to a preferred embodiment of the present invention. Detailed Implementation

[0035] To make the technical problems to be solved, the technical solutions, and the beneficial effects of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the present invention and are not intended to limit the present invention.

[0036] In the following description, specific details, such as particular internal procedures and techniques, are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of the invention. However, those skilled in the art will appreciate that the invention may be practiced in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of the invention with unnecessary detail.

[0037] like Figure 1 As shown, this invention provides a method for classifying pathological images of axillary lymph node metastasis in breast cancer. Its core lies in constructing a deep learning model using a convolutional-transformer module (TransConv) and a multi-level token transformer (MLTT). The specific steps are as follows:

[0038] Step 1: Image preprocessing and segmentation

[0039] Step 101: The whole-section pathological images (WSI) of axillary lymph nodes used in this experiment were all from the publicly available CAMELYON16 dataset. This dataset contains 398 WSIs, including 270 training images and 128 test images. The training set contains 159 normal WSIs and 111 WSIs containing tumor regions. Due to the extremely high image resolution and an average storage size exceeding 1GB, this dataset only contains image-level annotations and no pixel-level annotation information.

[0040] After preprocessing and block partitioning, and before inputting the value classification model, the original training set of 270 WSI images is divided into training data and test data in a 7:3 ratio.

[0041] Step 102: Preprocess the WSI. For example... Figure 2 As shown, since the WSI is very large, it is impossible to directly input the image into the network for model computation. Therefore, a TIF format WSI is divided into n 224×224 pixel JPG image patches, and zero-padding is applied to boundary regions that are less than 224 pixels. Based on the idea of ​​multiple instance learning, if all these patches are negative and all labels are 0, it means that this WSI is negative and the label is 0; but if at least one patch is positive and the label is 1, then this WSI is positive and the label is 1.

[0042]

[0043] Among them, when all patches If it is 0, then the WSI label =0; there is at least one patch. If the value is 1, then the WSI label The value is 1.

[0044] For each WSI image, the OTSU algorithm is used for foreground-background segmentation to eliminate a large number of blank and useless noise areas, including background areas and tissue fluid areas. The OTSU algorithm first converts the segmented patch into a grayscale image, setting the grayscale range to [0, 255], and then calculates the pixel values ​​of the grayscale image. The sum is compared with a predefined threshold. If the sum is less than the threshold, it means the patch contains a large number of blank areas and needs to be discarded; otherwise, it needs to be retained.

[0045]

[0046] in, Representing the pixel value of each point, first through Calculate the sum of all pixel values, then divide by the image area to obtain the average pixel value of the patch. .

[0047] Step 2: Feature Extraction

[0048] The n 224*224 patches retained in step 1 are input into the feature extraction network. Since the dataset only contains image-level annotations and lacks detailed pixel-level annotations, this embodiment of the invention uses the self-supervised contrastive learning framework SimCLR. Its core idea is to maximize the representational consistency between two different augmented views of the same image while minimizing the representational consistency between different image augmented views, thereby learning high-quality visual representations without labeled data.

[0049] In this step, the encoder network chosen is ResNet18. Since a single WSI image generates a large number of patches, compared to the ResNet50 used in existing technologies, ResNet18, selected in this embodiment, can process images with higher computational efficiency. Furthermore, the 224*224 input has good practicality in WSI, and ResNet18's receptive field can effectively capture key morphological features such as cell nuclei and glands. ResNet18 consists of 17 convolutional layers and 1 fully connected layer, including residual connection structures, with an output dimension of 512. The input image in (n, 224, 224) format is processed to generate an (n, 512) feature vector.

[0050] Step 3: Build and train the improved classification network

[0051] Because the WSI resolution is too large to be loaded into memory all at once, it is typically stored in a pyramid structure, containing several layers, usually downsampling upwards from the highest resolution. The bottom layer has the highest resolution and the richest local details; the top layer has the lowest resolution and the richest global information; the middle layers take into account both local and global information. Therefore, in this experiment, the n×512 patch-level feature map output from step 2 and the downsampled feature map are simultaneously input into the model for training. Specifically:

[0052] Patch-level feature maps: Step 2 extracts n image patches of size 224×224 pixels from each WSI. Each patch is processed by a pre-trained feature extractor to obtain a 512-dimensional feature vector, forming a feature matrix X ∈ R. {n ×512} To leverage the local spatial correlation of convolution operations, X is rearranged into a spatial feature map F ∈ R based on the two-dimensional grid coordinates (row and column indices) of each patch in the original WSI. {H×W×512} , where H×W = n (H and W are the number of rows and columns of the grid, respectively).

[0053] Downsampled feature map: Extracting the global feature map F from the top layer (lowest resolution) of the WSI pyramid. global ∈ R {h ×w×C_global}(In this experiment, h and w are much smaller than H and W, and C) global =256), this figure retains the overall organizational structure information.

[0054] Feature alignment: Bilinear interpolation is used to align F global The input feature map F is obtained by concatenating it along the channel dimension with F. in ∈ R {H×W×(512+256)} = R {H×W×768} .

[0055] The core of the classification network designed in this invention contains two innovative modules:

[0056] (1) Convolution-Transformer module (TransConv)

[0057] The TransConv module consists of multi-scale convolutional layers and self-attention layers, with its input being the feature map F. in .

[0058] First layer: Multi-scale convolutional layer on the input feature map F ∈ R {H×W×512} K different sizes of convolution kernels are used in parallel (in this design, K=3, with sizes of 3×3, 5×5, and 7×7, stride of 1, and padding of (kernel) size-1 / / 2 to ensure the output size remains constant). Each convolutional kernel outputs a feature map M. k ∈ R {H×W×C} Where C=256, then dimensionality is reduced to R by 1×1 convolution. {H×W×D} (D=512). The multi-scale fusion features are:

[0059]

[0060] The second layer is a self-attention layer, which inputs the multi-scale fused features output from the first layer into the Transformer's self-attention mechanism. This layer captures the global contextual dependencies between all image patches by calculating the association weights between them, thereby obtaining more discriminative deep features. This overcomes the deficiency of traditional MIL methods in ignoring interactions between image patches.

[0061]

[0062] The input vector undergoes a linear transformation to generate a query matrix Q, a key matrix K, and a value matrix V. After multi-head attention concatenation, residual connections, layer normalization, and a feedforward network (FFN), the self-attention layer output F is obtained. att ∈ R {T×512} Then reshape back to the spatial dimension F out ∈ R{H×W×512} This is the output of the current TransConv module. This self-attention layer captures the global contextual dependencies between all image patches by calculating the association weights, thereby obtaining more discriminative deep features and overcoming the deficiency of traditional Multiple Instance Learning (MIL) methods in ignoring interactions between image patches. Figure 3 As shown, Cov Layer1, Cov Layer2, and Cov Layer3 represent convolutional layers with kernels of 3×3, 5×5, and 7×7, respectively, and Transformer Layer is a self-attention layer.

[0063] (2) Multi-Level Token Transformer (MLTT)

[0064] The classification network in this embodiment of the invention consists of L stacked TransConv modules. A learnable class token vector is concatenated into the input sequence of each TransConv module. All L class tokens are concatenated along the channel dimension to obtain a multi-level fusion vector, which is then linearly reduced to decrease the risk of overfitting.

[0065]

[0066] Typically, standard Transformers only use the class token of the last layer, losing shallow layer details. In WSI, the determination of metastatic lesions often relies on joint evidence of local cellular atypia (shallow layers) and overall structural disruption (deep layers). Multi-level fusion can improve discriminative ability, preserve the independent channels of features from each layer, and allow the classification layer to adaptively learn the importance of different layers. For example... Figure 4 As shown, obtain the CLS Token for each layer in Transformer Layer 1 (M layer) and Transformer Layer 2 (N layer).

[0067] Concatenate all L class tokens along the channel dimension to obtain a multi-level fusion vector C ∈ R. {1 ×(L×512)} To reduce the risk of overfitting, a linear layer is used to reduce its dimensionality:

[0068]

[0069]

[0070] final The two-dimensional vector is transformed by the softmax function into the probability that the WSI belongs to negative (category 0) and positive (category 1):

[0071]

[0072] The predicted category is the one with the higher probability. This multi-level fusion strategy preserves the independent channels of features at each layer, allowing the classification layer to adaptively learn the importance of different layers. It is particularly suitable for metastatic lesion determination in WSI, which usually relies on joint evidence of local cellular atypia (shallow features) and global structural disruption (deep features).

[0073] Step 4: Model Training and Optimization

[0074] The end-to-end network described above was trained using a training set with image-level labels. The network parameters were optimized using backpropagation to minimize the loss function (e.g., cross-entropy loss) between the predicted results and the true labels. After training, the model can be used to predict the axillary lymph node metastasis status of newly input WSIs. The model training used the BCEWithLogitsLoss loss function. The formula for calculating BCEWithLogitsLoss is as follows: :

[0075]

[0076] in, For the true label of the sample, The Sigmoid function ultimately yields ℓ n Sample loss value.

[0077] This experiment uses accuracy and area under the curve (AUC) as quantitative evaluation metrics to assess the classification results of the model. Accuracy is the proportion of correctly predicted samples out of the total number of samples, ranging from [0,1]. A higher value indicates that the predicted result is closer to the true result, and the better the classification performance. The calculation formula is as follows:

[0078] In this embodiment, accuracy and area under the curve (AUC) are used as quantitative evaluation metrics to assess the classification results of the model. It should be noted that during model training, the original training set is divided into training and validation data in a 7:3 ratio according to step 101 for model parameter tuning and preliminary evaluation; while the final performance evaluation uses the independent test set (128 WSI images included with the CAMELYON16 dataset, which does not participate in any training or validation process).

[0079] Accuracy is the proportion of correctly predicted samples out of the total number of samples in the test set, ranging from [0,1]. A higher value indicates that the predicted result is closer to the true result, and the better the classification performance. The calculation formula is as follows:

[0080]

[0081] in:

[0082] TP (True Positive): The number of samples correctly predicted as positive.

[0083] TN (True Negative): The number of samples correctly predicted as negative.

[0084] FP (False Positive): The number of samples incorrectly predicted as positive (false positives);

[0085] FN (False Negative): The number of samples that were incorrectly predicted as negative (false negatives).

[0086] However, accuracy becomes ineffective when the number of positive and negative samples is significantly different. The ROC curve is plotted with the false positive rate (FPR) on the horizontal axis and the true positive rate (TPR, i.e., recall) on the vertical axis, by varying the decision threshold of the classifier. AUC is the area under the ROC curve, ranging from [0.5, 1], with a higher value indicating better classification performance. Its physical meaning can be understood as the probability that the model will rank the positive sample before the negative sample when randomly selecting one positive sample and one negative sample.

[0087] After training, the model achieved a classification accuracy (ACC) of 88% and an AUC of 0.94 on the test set. This indicates that the model has effectively learned the discriminative features of WSI. In practical applications, users only need to input the WSI to be predicted into the model according to the same preprocessing procedure (blocking, OTSU foreground filtering), and the model can automatically output the classification result of the WSI as positive or negative, without any manual annotation or additional intervention.

[0088] The following will use a specific example to illustrate the implementation process of the present invention:

[0089] 1. Data Preparation: 270 training WSIs and 129 test WSIs were obtained from the CAMELYON16 challenge dataset. All images were image-level labeled (i.e., only labeled whether the lymph node has metastasis, without pixel-level labeling).

[0090] 2. Preprocessing and Block Segmentation: For each WSI image, the OTSU algorithm is used for foreground and background segmentation. At a magnification of 20x, the tissue region is divided into 224x224 pixel image blocks with a stride of 224. The average gray value of each image block is calculated; if it is less than a threshold T (e.g., T = 0.8 × 255), it is considered a blank area and discarded.

[0091] 3. Feature Extraction: A SimCLR model based on ResNet18 is constructed and pre-trained in a self-supervised manner on all valid image patches to obtain a feature extractor. Then, all valid image patches from each WSI are input into this feature extractor to obtain an n × 512 feature matrix.

[0092] 4. Construct a classification network:

[0093] ① Construct a stack of 4 convolutional-transformer modules (TransConv).

[0094] ② In each TransConv, the input n×512 features are first linearly projected to obtain n×512 features with unchanged dimensions. Then, the three outputs are fused by depthwise separable convolutions of 3x3, 5x5, and 7x7 (to reduce the number of parameters). Next, the fused features are input into a standard Transformer encoding layer for self-attention calculation, outputting new n×512 features and a 1×512 class token for that layer.

[0095] ③ Collect the class tokens output by all 4 TransConv modules, concatenate them into a 4×512 matrix, and then flatten it into a (4*512) one-dimensional vector.

[0096] ④ Input the vector into a classification head containing a two-layer fully connected network (e.g., the first layer is 4096-dimensional and the second layer is 2-dimensional), and finally obtain the probability of WSI belonging to transition / non-transition through the Softmax function.

[0097] 5. Training and Testing: Training was performed using an NVIDIA GPU within the PyTorch framework. The learning rate was set to 1e-4, the batch size to 1 (each WSI is one batch), and the optimizer to AdamW. During training, the AUC value was monitored on the validation set, and the best model was saved. Finally, the model's accuracy and AUC were evaluated on the test set to obtain the results.

[0098] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.

Claims

1. A breast cancer axillary lymph node metastasis pathological image classification method, characterized in that, Includes the following steps: We collected whole-section pathological images of axillary lymph nodes from breast cancer using Western blotting (WSI). All images were preprocessed, divided into n image blocks, and then divided into training and testing sets. The preprocessed image patch is input into a feature extraction network for feature extraction to obtain feature vectors; An improved classification network is constructed, which includes building several stacked dual attention modules, linearly projecting the feature vectors in the dual attention modules, then performing depthwise separable convolutions on each module and then adding and fusing them. The fused features are then input into a standard Transformer encoding layer for self-attention calculation to obtain class tokens. The class tokens are then concatenated and flattened into a one-dimensional vector, which is then input into the classification head of a fully connected network to obtain the probability that WSI belongs to transition / non-transition. The improved classification network is trained and optimized using the data from the training set, the best model is saved, and the best model is evaluated on the test set.

2. The breast cancer axillary lymph node metastasis pathological image classification method of claim 1, wherein, The collected whole-section pathological images of axillary lymph nodes from breast cancer were only annotated at the image level, without any pixel-level annotation information.

3. The breast cancer axillary lymph node metastasis pathological image classification method of claim 2, wherein, The collected pathological images of whole sections of axillary lymph nodes from breast cancer were divided into training and testing sets in a 7:3 ratio.

4. The breast cancer axillary lymph node metastasis pathological image classification method of claim 1, wherein, Preprocessing of all images includes dividing the TIF format WSI into n 224×224 pixel JPG image patches and zero-padding boundary areas that are less than 224 pixels; and eliminating a large number of blank and useless noise areas in the WSI, including background areas and tissue fluid areas.

5. The breast cancer axillary lymph node metastasis pathological image classification method of claim 1, wherein, The preprocessed image patches are input into a feature extraction network to extract features and obtain feature vectors. Specifically, a SimCLR model based on ResNet18 is constructed and self-supervised pre-trained on all valid image patches to obtain a feature extractor. Then, all valid image patches of each WSI are input into the feature extractor to obtain an n×512 feature matrix.

6. The breast cancer axillary lymph node metastasis pathological image classification method of claim 1, wherein, The double-attention module comprises a multi-scale convolution layer and a self-attention layer, the multi-scale convolution layer uses K different sizes of convolution kernels in parallel for an input feature map F, each convolution kernel outputs a feature map M F , and performs multi-scale fusion features The self-attention layer inputs the multi-scale fused features output from the first layer into the self-attention mechanism of the Transformer, capturing the global contextual dependencies between them, thereby obtaining deep features.

7. The breast cancer axillary lymph node metastasis pathological image classification method of claim 6, wherein, The improved classification network comprises L stacked TransConv modules, with a learnable class token vector concatenated into the input sequence of each TransConv module. All L class tokens are concatenated along the channel dimension to obtain a multi-level fusion vector, which is then linearly reduced in dimensionality.

8. The breast cancer axillary lymph node metastasis pathological image classification method of claim 1, wherein, The improved classification network is trained and optimized using the training set data, including training the improved classification network using a training set with image-level labels, optimizing the network parameters through the backpropagation algorithm, and minimizing the loss function between the predicted results and the true labels. Once training is complete, the optimal model can be used to predict the axillary lymph node metastasis status of newly input WSI.

9. The breast cancer axillary lymph node metastasis pathological image classification method of claim 8, wherein, The loss function is set to cross-entropy loss.

10. The breast cancer axillary lymph node metastasis pathological image classification method of claim 8, wherein, When training and optimizing the improved classification network using the training set data, the learning rate is 1e-4, the batch size is 1, and the optimizer is AdamW.