Breast cancer pathology image scoring method based on deep learning
By combining the neural network model and attention mechanism of CNN and BiLSTM, the long diagnosis time and inaccurate results in pre-screening of early pathological images of breast cancer are solved, and efficient and accurate breast cancer pathological images are achieved.
Patent Information
- Application Number
- CN202510237376.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-02
- Publication Date
- 2025-07-18
AI Technical Summary
The existing pre-screening methods for early breast cancer pathological images have a long diagnosis time and inaccurate results.
A neural network model of encoder and decoder structure combined with convolutional neural network (CNN) and long and short-term memory network (BiLSTM) is used, and an attention mechanism is introduced to construct a breast cancer pathological image scoring method, and scored through tubules, mitotic counting and nuclear pleomorphic characteristics.
It improves the recognition accuracy of breast cancer pathological images, shortens recognition time and reduces recognition cost.
Smart Images

Figure CN120339166A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for pre-screening breast cancer pathological images, and particularly to a method for pre-screening breast cancer pathological images based on deep learning. Background Art
[0002] Breast cancer, as one of the most lethal diseases, brings many physical and psychological troubles to women patients. The breast is composed of three structures: connective tissue, lobules, and ducts. Under normal circumstances, the ducts and lobules are exactly the areas where cancer cells are most likely to grow. Breast cancer patients will present various symptoms. Swelling, color change in the infected area, and specific inflammations that appear in the breast and armpit are some external manifestations of the disease. Early detection and diagnosis of breast cancer are of great significance. It can not only save many lives but also help patients recover faster and avoid extremely painful surgeries or chemotherapy.
[0003] In the field of machine learning, deep learning technology currently plays a crucial role in the early detection of major diseases such as cancer. Taking breast cancer as an example, based on pathological image data and existing scoring rules (as shown in Figure 7), breast cancer can be scored from 3 to 9 points according to the characteristics of tubule formation, mitotic count, and nuclear pleomorphism, and further divided into 3 grades. In recent years, a large number of experiments on histopathological images have achieved positive results. These tiny images can be collected and integrated for constructing a computer-aided detection system.
[0004] Since most cells are often in an irregular, unpredictable, and arbitrary perspective state, manual recognition is not only time-consuming and laborious but also extremely error-prone, making the diagnosis process of breast cancer time-consuming and costly. In view of this, the present invention innovatively proposes a breast cancer pathological image scoring method that combines a convolutional neural network (CNN) with a bidirectional long short-term memory network (BiLSTM) and introduces an attention mechanism. Summary of the Invention
[0005] The purpose of the present invention is to solve the problems of long diagnosis time and inaccurate diagnosis results existing in the existing method for pre-screening breast cancer through early breast cancer pathological images, and to propose a breast cancer pathological image scoring method based on deep learning.
[0006] The above purpose is achieved through the following technical solutions: A breast cancer pathological image scoring method based on deep learning, the method is implemented through the following steps: Step 1: Obtain breast cancer pathological images and make a dataset of breast pathological images containing the characteristics of tubule formation, mitotic count, and nuclear pleomorphism, and label the corresponding grades of the three characteristics of the pathological images; Step 2. Perform data preprocessing on the dataset, and the preprocessing includes: Uniformly adjust the images in the dataset to grayscale images with a size of 244×244×1, and perform data augmentation through elastic transformation, morphological transformation, and noise addition methods; then, divide the dataset into a training set, a validation set, and a test set according to the ratio of 7:2:1; Step 3. Construct a neural network model with an encoder and a decoder structure that combines CNN and BiLSTM, and use the dataset preprocessed in Step 2 to train the neural network model; Step 4. Use the trained neural network model to identify the image to be recognized and obtain the corresponding grades of the image to be recognized and the three characteristics according to the characteristics of tubule formation, mitosis count, and nuclear pleomorphism; For the above trained model, the maximum values of the output values of the three branches are taken as the scoring results of each branch, and the scoring results are added to obtain the final score Score, that is, Score = Max(S1)+Max(S2)+Max(S3); Among them, Max(S1) is the scoring result after the maximum value operation of the first branch, Max(S2) is the scoring result after the maximum value operation of the second branch, and Max(S3) is the scoring result after the maximum value operation of the third branch.
[0007] Further, the process of constructing a neural network model with an encoder and a decoder structure that combines CNN and BiLSTM in Step 3 and using the dataset preprocessed in Step 2 to train the neural network model is specifically as follows: Step 3-1. Construct a neural network model with an encoder and a decoder structure that combines CNN and BiLSTM. The neural network model includes a CNN encoder and a BiLSTM decoder. The neural network model includes: Conv, Mish, Conv, Mish, InstanceNormal2d, Conv, TripleAttention; the CNN encoder includes 6 CNN composite modules, denoted as CNN_TripleAttention; the BiLSTM decoder includes 2 bidirectional BiLSTM modules; Among them, Mish is an activation function, and its expression is: f ( x )= x * tanh ( ln (1+ e x )); TripleAttention is an attention mechanism; The CNN encoder structure described above is as follows: The size of the feature map corresponding to the model input is 1×244×244, the size of the feature map corresponding to CNN_TripleAttention1 is 16×32×32, the size of the feature map corresponding to CNN_TripleAttention2 is 32×16×16, the size of the feature map corresponding to CNN_TripleAttention3 is 32×8×8, the size of the feature map corresponding to CNN_TripleAttention4 is 64×4×4, the size of the feature map corresponding to CNN_TripleAttention5 is 64×2×2, and the size of the feature map corresponding to CNN_TripleAttention6 is 128×1×1; Step 3.2: Feed the training set into the neural network model with the combined encoder and decoder structure of CNN and BiLSTM according to the input size of batchsize = 4 for training; The network model has three branches, which respectively score three characteristics: tubule formation, mitosis count, and nuclear pleomorphism. The loss function Loss of the entire network is the sum of the scoring results of the three branches, that is, Loss = Loss1 + Loss2 + Loss3; where Loss1 is the result of scoring the tubule formation characteristic through one branch of the network model, Loss2 is the result of scoring the mitosis count characteristic through one branch of the network model, and Loss3 is the result of scoring the nuclear pleomorphism characteristic through one branch of the network model.
[0008] The beneficial effects of the present invention are as follows: CNN is good at extracting implicit spatial features, while the long short-term memory network (BiLSTM) is good at extracting sequence or temporal data. Therefore, the present invention proposes to combine the convolutional neural network (CNN) and the long short-term memory network (BiLSTM), thereby constructing a neural network model with a combined encoder and decoder structure of CNN and BiLSTM, and introducing an attention mechanism for breast cancer pathological image scoring method. Using the three branches of the network model corresponding to three characteristics of tubule formation, mitosis count, and nuclear pleomorphism respectively for scoring, accurate scoring results are obtained. And the method of the present invention shortens the time required for the recognition process and reduces the recognition cost.
[0009] The method of the present invention can be used to assist doctors in diagnosing breast cancer by identifying breast cancer pathological images. Description of the Drawings
[0010] Figure 1 is the flowchart of the method of the present invention; Figure 2 is the neural network model with a combined encoder and decoder structure of CNN and BiLSTM involved in the present invention; Figure 3 It is a branch of the breast cancer pathological image scoring model that combines CNN and BiLSTM involved in the present invention; Figure 4 It is a schematic diagram of the TripleAttention attention mechanism network structure involved in the present invention; Figure 5 Training curve of the pre-screening model for breast cancer images based on deep learning; Figure 6 It is a breast cancer pathological diagnosis image involved in the background technology part of the present invention; Figure 7 It is a diagram showing invasive ductal carcinoma of the breast and the Nottingham histological grading system involved in the background technology part of the present invention. Detailed implementation manners
[0011] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. Detailed implementation manner one:
[0013] A method for scoring breast cancer pathological images based on deep learning in this embodiment, as Figure 1 shown, the method is implemented through the following steps: Step 1, obtain breast cancer pathological images and make a data set of breast pathological images containing characteristics of tubule formation, mitotic count, and nuclear pleomorphism, and label the corresponding grades of the three characteristics of the pathological images; Step 2, perform data preprocessing on the data set, and the preprocessing includes: Uniformly adjust the images in the data set to grayscale images of size 244×244×1, and perform data augmentation by methods such as elastic transformation, morphological transformation, and adding noise; then, divide the data set into a training set, a validation set, and a test set according to a ratio of 7:2:1; Step 3, construct a neural network model with an encoder and a decoder structure that combines CNN and BiLSTM, and train the neural network model using the data set preprocessed in Step 2. The specific process is as follows: Step 3-1, construct a neural network model with an encoder and a decoder structure that combines CNN and BiLSTM, as Figure 2As shown, the neural network model includes a CNN encoder and a BiLSTM decoder, and the neural network model includes: Conv, Mish, Conv, Mish, InstanceNormal2d, Conv, TripleAttention; the CNN encoder includes 6 CNN composite modules, denoted as CNN_TripleAttention; the BiLSTM decoder includes 2 bidirectional BiLSTM modules; Among them, Mish is an activation function, and its expression is: f ( x ) = x * tanh ( ln (1 + e x )); TripleAttention is an attention mechanism, and its network structure is as Figure 4 shown.
[0014] The structure of the CNN feature encoder includes: the size of the feature map corresponding to the model input is 1×244×244, the size of the feature map corresponding to CNN_TripleAttention1 is 16×32×32, the size of the feature map corresponding to CNN_TripleAttention2 is 32×16×16, the size of the feature map corresponding to CNN_TripleAttention3 is 32×8×8, the size of the feature map corresponding to CNN_TripleAttention4 is 64×4×4, the size of the feature map corresponding to CNN_TripleAttention5 is 64×2×2, and the size of the feature map corresponding to CNN_TripleAttention6 is 128×1×1, that is: Step 3-2: Feed the training set into the neural network model of the encoder and decoder structure combined with CNN and BiLSTM with an input size of batchsize = 4 for training; the network model has three branches, which respectively score three characteristics of tubule formation, mitotic count, and nuclear pleomorphism. The loss function Loss of the entire network is the sum of the scoring results of the three branches, that is, Loss = Loss1 + Loss2 + Loss3; among them, Loss1 is the result of scoring the tubule formation characteristic through a branch of the network model, Loss2 is the result of scoring the mitotic count characteristic through a branch of the network model, and Loss3 is the result of scoring the nuclear pleomorphism characteristic through a branch of the network model; Step 4: Use the trained neural network model to recognize the image to be recognized and obtain the corresponding levels of the image to be recognized and the three characteristics according to the characteristics of tubule formation, mitotic count, and nuclear pleomorphism; For the above-mentioned trained model, the maximum values of the output values of the three branches are taken as the scoring results of each branch, and the scoring results are added to obtain the final score Score, that is, Score = Max(S1)+Max(S2)+Max(S3); Among them, Max(S1) is the scoring result after the maximum value operation of the first branch, Max(S2) is the scoring result after the maximum value operation of the second branch, and Max(S3) is the scoring result after the maximum value operation of the third branch.
[0015] Finally, display and save the classification results of the scores.
[0016] The embodiments disclosed in the present invention are preferred embodiments, but not limited thereto. Those of ordinary skill in the art can easily understand the spirit of the present invention based on the above embodiments and make different extensions and changes. However, as long as they do not depart from the spirit of the present invention, they are within the protection scope of the present invention.
Claims
1. A breast cancer pathological image scoring method based on deep learning, characterized in that: The method is implemented through the following steps: Step 1: Obtain breast cancer pathological images and make a dataset of breast pathological images containing characteristics such as tubule formation, mitotic count, and nuclear pleomorphism, and label the corresponding grades of the three characteristics of the pathological images; Step 2: Perform data preprocessing on the dataset, and the preprocessing includes: Uniformly adjust the images in the dataset to grayscale images with a size of 244×244×1, and perform data augmentation through elastic transformation, morphological transformation, and noise addition methods; then, divide the dataset into a training set, a validation set, and a test set according to the ratio of 7:2:1; Step 3: Construct a neural network model with an encoder and a decoder structure combining CNN and BiLSTM, and train the neural network model using the dataset preprocessed in Step 2; Step 4: Use the trained neural network model to identify the image to be recognized and obtain the corresponding grades of the image to be recognized and the three characteristics according to the characteristics of tubule formation, mitotic count, and nuclear pleomorphism; For the trained above model, the maximum values of the output values of the three branches are taken as the scoring results of each branch, and the scoring results are added to obtain the final score Score, that is, Score = Max(S1)+Max(S2)+Max(S3); Among them, Max(S1) is the scoring result after the maximum value operation of the first branch, Max(S2) is the scoring result after the maximum value operation of the second branch, and Max(S3) is the scoring result after the maximum value operation of the third branch.
2. The method for scoring breast cancer pathological images based on deep learning according to claim 1, wherein: The process of constructing a neural network model with an encoder and a decoder structure combining CNN and BiLSTM in Step 3 and training the neural network model using the dataset preprocessed in Step 2 is specifically: Step 3-1: Construct a neural network model with an encoder and a decoder structure combining CNN and BiLSTM. The neural network model includes a CNN encoder and a BiLSTM decoder. The neural network model includes: Conv, Mish, Conv, Mish, InstanceNormal2d, Conv, TripleAttention; The CNN encoder includes 6 CNN composite modules, denoted as CNN_TripleAttention; The BiLSTM decoder includes 2 bidirectional BiLSTM modules; Among them, The Mish activation function has the following expression: f ( x )= x * tanh ( ln (1 + e x )); TripleAttention is an attention mechanism; The CNN encoder structure described is as follows: The size of the feature map corresponding to the model input is 1×244×244, the size of the feature map corresponding to CNN_TripleAttention1 is 16×32×32, the size of the feature map corresponding to CNN_TripleAttention2 is 32×16×16, the size of the feature map corresponding to CNN_TripleAttention3 is 32×8×8, the size of the feature map corresponding to CNN_TripleAttention4 is 64×4×4, the size of the feature map corresponding to CNN_TripleAttention5 is 64×2×2, and the size of the feature map corresponding to CNN_TripleAttention6 is 128×1×1; Step 3-2: Feed the training set into the neural network model of the combined encoder and decoder structure of CNN and BiLSTM with an input size of batchsize = 4 for training; the network model has three branches, which respectively score three characteristics of tubule formation, mitotic count, and nuclear pleomorphism. The loss function Loss of the entire network is the sum of the scoring results of the three branches, that is, Loss = Loss1 + Loss2 + Loss3; where Loss1 is the result of scoring the tubule formation characteristic through one branch of the network model, Loss2 is the result of scoring the mitotic count characteristic through one branch of the network model, and Loss3 is the result of scoring the nuclear pleomorphism characteristic through one branch of the network model.