Breast cancer pathological image classification method and device, electronic equipment and medium
By extracting and enhancing the information of the nucleus edge and using an enhanced vision transformer embedded in the wavelet position for classification, the problem of inaccurate classification of breast cancer pathological images under low resolution conditions is solved, and higher classification accuracy is achieved.
Patent Information
- Application Number
- CN202510158775.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-06-03
AI Technical Summary
The prior art has inaccurate classification of breast cancer pathological images under low resolution conditions, making it difficult to effectively extract key features of the cell nucleus.
By acquiring the pathological images to be processed, the nuclear region is extracted using the segmentation model, the nuclear contour is extracted and the original image is fused to obtain an enhanced image, and the enhanced image is classified and processed using an enhanced vision transformer embedded in the wavelet position.
The classification model's ability to identify morphological characteristics of the nucleus is improved, and the key features of the nucleus can be effectively extracted at low resolution, improving the accuracy of classification results.
Smart Images

Figure CN120088780A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to a breast cancer pathological image classification method, apparatus, electronic device, and medium. Background Art
[0002] Breast cancer is an epithelial malignant tumor originating from the terminal duct lobular unit of the breast. Its histological morphology is extremely complex and can be roughly divided into two types: non-invasive carcinoma and invasive carcinoma. Among them, invasive ductal carcinoma is the most common type of breast cancer, with an incidence rate as high as 70%.
[0003] In the study of breast cancer pathological images, there is a close connection between morphological features and potential molecular features. This observation result is of great significance for the application of artificial intelligence technology in breast cancer research. Specifically, potential molecular changes and biological characteristics can be inferred from image features, and the success of this inference process depends to a large extent on the quality of the image, especially the resolution. Theoretically, with high-resolution pathological images, the microscopic structure and features of breast cancer tissues, including the shape of the cell nucleus, cell arrangement, and blood vessel distribution, can be clearly shown, which is crucial for the accurate analysis and inference of artificial intelligence models.
[0004] However, in practical applications, due to various factors such as technical limitations, sample processing problems, or storage conditions, it is difficult to obtain images with sufficient resolution under the current technical conditions, or the implementation cost is extremely high. In this case, the importance of low-resolution images becomes particularly obvious. Although low-resolution images may not provide the same details as high-resolution images, they still contain valuable information. Extracting key features, such as specific textures, colors, shapes, etc., from low-resolution images is equally crucial for the diagnosis and classification of breast cancer. However, some common classification methods currently have relatively low accuracy in classifying low-resolution pathological images. Summary of the Invention
[0005] In view of the above defects of the prior art, the present invention provides a breast cancer pathological image classification method, apparatus, electronic device, and medium to solve the technical problem of inaccurate classification of breast cancer pathological images under low-resolution images.
[0006] To achieve the above object and other related objects, the present invention provides a breast cancer pathological image classification method, including: obtaining a pathological image to be processed; using a segmentation model to extract the cell nucleus region in the pathological image to obtain a segmentation image; extracting the cell nucleus contour from the segmentation image and fusing the cell nucleus contour with the pathological image to obtain an enhanced image; using a classification model to perform classification processing on the enhanced image to obtain a classification result, where the classification model is an enhanced vision transformer embedded with wavelet positions.
[0007] In one embodiment of the present invention, a segmentation model is used to extract the nucleus region in the pathological image to obtain a segmentation image, including: using a Unet model or a ResUnet model or a UTnet model to extract the nucleus region in the pathological image to obtain a segmentation image.
[0008] In one embodiment of the present invention, the classification model includes a tokenizer and a classifier; using the classification model to perform classification processing on the enhanced image to obtain a classification result, including: performing normalization processing on the enhanced image to convert the enhanced image into a tensor; using the tokenizer to convert the tensor into a low-dimensional feature; using the classifier to process the low-dimensional feature to obtain the classification result.
[0009] In one embodiment of the present invention, the tokenizer is formed by cascading a plurality of convolutional blocks, and each convolutional block is formed by cascading a first convolutional layer, an activation layer, and a max pooling layer.
[0010] In one embodiment of the present invention, the classifier includes a cascaded wavelet position embedding module, a transformer encoder, a sequence pooling layer, and a first fully connected layer; using the classifier to process the low-dimensional feature to obtain the classification result, including: using the wavelet position embedding module to decompose the low-dimensional feature into components of different frequencies and embed position information to enhance the feature representation to obtain an input sequence; using the transformer encoder to process the input sequence to capture long-range dependencies and context information in the input sequence; using the sequence pooling layer to summarize the output of the transformer encoder into a fixed-length vector; using the first fully connected layer to convert the fixed-length vector into the classification result.
[0011] In one embodiment of the present invention, the wavelet position embedding module includes a cascaded transposed layer, a second fully connected layer, a second convolutional layer, a trigonometric connection layer, and a third fully connected layer; using the wavelet position embedding module to decompose the low-dimensional feature into components of different frequencies, including: using the cascaded transposed layer, the second fully connected layer, the second convolutional layer, the trigonometric connection layer, and the third fully connected layer to sequentially process the low-dimensional feature, and adding the processed feature to the low-dimensional feature and then outputting to decompose the low-dimensional feature into components of different frequencies.
[0012] In an embodiment of the present invention, the transformer encoder includes a normalization layer, a multi-head self-attention module, and a multi-layer perceptron; processing the input sequence using the transformer encoder includes: processing the input sequence using a cascaded normalization layer and a multi-head self-attention module to obtain a first feature; adding the input sequence and the first feature to obtain a second feature; processing the second feature using a cascaded normalization layer and a multi-layer perceptron to obtain a third feature; adding the second feature and the third feature to obtain a fourth feature; processing the fourth feature as a new input sequence according to the above steps again to obtain the output feature of the transformer encoder.
[0013] To achieve the above object and other related objects, the present invention also provides a breast cancer pathological image classification device, including: a data acquisition unit for acquiring a pathological image to be processed; a first processing unit for extracting the nucleus region in the pathological image using a segmentation model to obtain a segmentation image; a second processing unit for extracting the nucleus contour from the segmentation image and fusing the nucleus contour with the pathological image to obtain an enhanced image; a third processing unit for classifying the enhanced image using a classification model to obtain a classification result, wherein the classification model is an enhanced vision transformer embedded with wavelet positions.
[0014] To achieve the above object and other related objects, the present invention also provides an electronic device, including a processor, a memory, and a communication bus; the communication bus is used to connect the processor and the memory; the processor is used to execute a computer program stored in the memory to implement the method provided in any one of the above embodiments.
[0015] To achieve the above object and other related objects, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and the computer program is used to make a computer execute the method provided in any one of the above embodiments.
[0016] Advantages of the present invention: A breast cancer pathological image classification method, device, electronic device, and medium proposed by the present invention. This method improves the recognition ability of the classification model for the morphological features of the nucleus by extracting and enhancing the nucleus edge information. Even in the case of low resolution, it can effectively extract the key features of the nucleus; at the same time, using an enhanced vision transformer embedded with wavelet positions as the classification model can not only well represent the smooth and low-frequency components in the feature map, but also accurately capture fine details and high-frequency parts, thereby further improving the accuracy of the classification result. Description of the Drawings
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0018] Figure 1 Flow chart of the classification method provided by an embodiment of the present invention;
[0019] Figure 2 Structural diagram of the classification model provided by an embodiment of the present invention;
[0020] Figure 3 Processing flow chart of the classification model provided by an embodiment of the present invention;
[0021] Figure 4 Processing flow chart of the classifier provided by an embodiment of the present invention;
[0022] Figure 5 Structural diagram of the wavelet position embedding module provided by an embodiment of the present invention;
[0023] Figure 6 Structural diagram of the transformer encoder provided by an embodiment of the present invention;
[0024] Figure 7 Confusion matrix comparison diagram of the performance of the test data of the EVT model using kernel information enhancement and wavelet position transformation provided by an embodiment of the present invention;
[0025] Figure 8 ROC curve and AUC value of the classification performance of various EVT model combinations on the test data provided by an embodiment of the present invention;
[0026] Figure 9 Schematic diagram of the classification device provided by an embodiment of the present invention;
[0027] Figure 10 Schematic structural diagram of an electronic device provided by an embodiment of the present invention.
[0028] Explanation of reference numerals: 100, tagger; 200, classifier; 201, wavelet position embedding module; 202, transformer encoder; 203, sequence pooling layer; 204, first fully connected layer; 301, data acquisition unit; 302, first processing unit; 303, second processing unit; 304, third processing unit; 401, processor; 402, memory. Detailed implementation manners
[0029] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. In addition to the specific methods, devices, and materials used in the embodiments, according to the knowledge of those skilled in the art in the technical field and the description of the present invention, any methods, devices, and materials similar to or equivalent to those described in the embodiments of the present invention can also be used to implement the present invention.
[0030] It should be understood that the terms used in the embodiments of the present invention are for the purpose of describing specific embodiments, rather than limiting the protection scope of the present invention. Unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those skilled in the technical field of the present invention.
[0031] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In some of these embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.
[0032] The flowcharts and block diagrams in the accompanying drawings illustrate the architectures, functions, and operations that the methods and computer program products according to various embodiments disclosed in the present invention may achieve. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and this module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0033] Please refer to Figure 1 , Figure 1 A breast cancer pathological image classification method provided in an embodiment of the present invention, including steps S100 to S400.
[0034] Step S100: Obtain the pathological image to be processed. In the actual acquisition process of breast cancer pathological images, due to various factors, it is inevitable to obtain low-resolution images. However, in the existing technology, when dealing with such low-resolution images, the classification accuracy is often not satisfactory, and a large amount of potential effective information fails to be fully mined and utilized. The classification method adopted subsequently in the present invention is tailored for low-resolution scenarios. Therefore, in this step, the obtained pathological image can be a common low-resolution image.
[0035] Step S200: Use the segmentation model to extract the nucleus region in the pathological image to obtain the segmentation image. In this step, the segmentation model is pre-trained and specifically used to process pathological images to segment the nucleus region from the pathological image. In the obtained segmentation image, the nucleus region is different from the non-cell region. For example, the nucleus region is black and other regions are white.
[0036] In a specific embodiment of the present invention, using the segmentation model to extract the nucleus region in the pathological image to obtain the segmentation image includes: using the Unet model or the ResUnet model or the UTnet model to extract the nucleus region in the pathological image to obtain the segmentation image.
[0037] The Unet model is a classic convolutional neural network architecture mainly used for medical image segmentation tasks. It adopts an encoder-decoder structure, extracts features through downsampling (such as max pooling), and then restores the spatial resolution of the image through upsampling (such as transposed convolution).
[0038] The ResUnet model is a fusion of the Unet and ResNet architectures. It introduces the residual blocks in ResNet on the basis of Unet, solves the gradient vanishing problem in deep networks through skip connections, and better captures local and global features at the same time.
[0039] The UTnet model is a hybrid Transformer architecture that combines the advantages of convolutional layers and self-attention mechanisms. On the basis of Unet, it replaces some convolutional layers with efficient self-attention modules, can capture long-range correlation features, and is suitable for medical image segmentation.
[0040] In addition to the above three models, other models that can perform nucleus segmentation can also be used.
[0041] It can be understood that in practical applications, the pathological image needs to be preprocessed such as normalized and resized, and finally converted into a tensor form before it can be input into the segmentation model for training or inference. And the outputs of these segmentation models are not in a direct image format but tensors, and these tensors can be further processed or converted into an image format to obtain the segmentation image.
[0042] Step S300: Extract the nucleus contour from the segmented image and fuse the nucleus contour with the pathological image to obtain an enhanced image. In this step, the original pathological image is enhanced based on the segmented image. In the enhanced image, the pixel region corresponding to the nucleus contour will be prominently displayed, while other regions remain unchanged. This step not only highlights the nuclear contour in the image but also effectively retains the information of the surrounding tissues, which is of great significance for in-depth analysis of the nucleus and its surrounding environment.
[0043] In a specific embodiment of the present invention, for example, the Canny edge detection operator can be used to extract the nuclear contour from the segmented image. After obtaining the nuclear contour, when fusing it with the original pathological image, for example, the pixel points corresponding to the nuclear contour can be directly set to black, or the pixel values of the pixel points corresponding to the nuclear contour can be weighted with black, or other fusion methods can be used, as long as the nuclear contour can be highlighted.
[0044] Step S400: Use the classification model to perform classification processing on the enhanced image to obtain a classification result, where the classification model is an enhanced vision transformer embedded with wavelet positions. In this step, using the enhanced vision transformer embedded with wavelet positions as the classification model can effectively reduce the aliasing effect in the downsampling process, accurately extract key features, and achieve efficient feature extraction and accurate classification.
[0045] Please refer to Figure 2 , in a specific embodiment of the present invention, the classification model includes a tokenizer 100 and a classifier 200; using the classification model to perform classification processing on the enhanced image to obtain a classification result includes steps S410 - S430, and its flowchart is as Figure 3 shown.
[0046] Step S410: Normalize the enhanced image to convert it into a tensor. Similar to the above segmentation model, for the classification model, the enhanced image also needs to be preprocessed to facilitate the processing of the classification model. The normalization parameters of the enhanced image can be set to [[0.485, 0.456, 0.406], [0.229, 0.224, 0.225]] for example, so as to convert the enhanced image into a tensor with three channels for classification model inference.
[0047] It can be understood that the classification model is also pre-trained. When training or validating the classification model, the images used also need to be preprocessed.
[0048] Step S420: Use the tokenizer 100 to convert the tensor into low-dimensional features. In this step, the tokenizer 100 provides a basis for subsequent feature extraction and classification tasks, ensuring that the subsequent classifier 200 can effectively perform classification processing.
[0049] Please refer to Figure 2 that in a specific embodiment of the present invention, the tagger 100 is formed by cascading a plurality of convolutional blocks. For example, it can be formed by cascading 3 convolutional blocks. Each convolutional block is formed by cascading a first convolutional layer, an activation layer, and a max pooling layer.
[0050] Step S430: Process the low-dimensional features using the classifier 200 to obtain a classification result.
[0051] Please refer to Figure 2 that in a specific embodiment of the present invention, the classifier 200 includes a cascaded wavelet position embedding module 201, a transformer encoder 202, a sequence pooling layer 203, and a first fully connected layer 204. Step S430 specifically includes steps S431 to S434, and its flowchart is as Figure 4 shown.
[0052] Step S431: Use the wavelet position embedding module 201 to decompose the low-dimensional features into components of different frequencies, and embed position information to enhance the feature representation, obtaining an input sequence. The wavelet transform technology decomposes the sequence into amplitude and phase components, providing a novel frequency domain method that can accurately identify and extract key features for the classification of histopathological images. This technology can effectively utilize image features and exhibits excellent performance in the classification of pathological images.
[0053] Please refer to Figure 5 that in a specific embodiment of the present invention, the wavelet position embedding module 201 includes a cascaded transposed layer, a second fully connected layer, a second convolutional layer, a trigonometric function connection layer, and a third fully connected layer. Using the wavelet position embedding module 201 to decompose the low-dimensional features into components of different frequencies specifically includes: sequentially processing the low-dimensional features using the cascaded transposed layer, the second fully connected layer, the second convolutional layer, the trigonometric function connection layer, and the third fully connected layer, and adding the processed features to the low-dimensional features and then outputting them to decompose the low-dimensional features into components of different frequencies.
[0054] Specifically, first, use the transposed layer to perform a transpose operation on the input tensor to adjust its dimensions to ensure the adaptability of subsequent calculations; secondly, perform a linear transformation through the second fully connected layer to generate a preliminary feature representation; then use the second convolutional layer to extract spatial features from these features; then combine cosine and sine functions (i.e., the trigonometric function connection layer) to enhance the expression ability of the features; then apply the third fully connected layer to transform the fused features to improve the nonlinear fitting ability of the model and enhance the complexity and expression ability of the features; finally, add the processed features to the original input through a residual connection to effectively transfer information and promote the training stability and convergence of the model in the deep structure.
[0055] Step S432: Process the input sequence using the Transformer encoder 202 to capture long-range dependencies and context information in the input sequence. The Transformer encoder 202, namely the Transformer encoder, can also be referred to as the transformer encoder. The Transformer encoder has become a cutting-edge technology in the field of deep learning and can achieve robust results in various computer vision tasks. In the field of histopathological imaging, the transformer method has been widely applied to multiple aspects such as image segmentation, classification, detection, representation, cross-modal retrieval, image generation, survival analysis, and survival prediction.
[0056] Please refer to Figure 6 , in a specific embodiment of the present invention, the Transformer encoder 202 includes a normalization layer, a multi-head self-attention module, and a multi-layer perceptron. Processing the input sequence using the Transformer encoder 202 specifically includes the following steps: (1) Process the input sequence using the cascaded normalization layer and the multi-head self-attention module to obtain a first feature; (2) Add the input sequence and the first feature to obtain a second feature; (3) Process the second feature using the cascaded normalization layer and the multi-layer perceptron to obtain a third feature; (4) Add the second feature and the third feature to obtain a fourth feature; (5) Use the fourth feature as the new input sequence and process it again according to the above steps to obtain the output feature of the Transformer encoder 202. The i-th feature here is only for facilitating the description of the processing process of the Transformer encoder 202. In fact, the Transformer encoder 202 can be regarded as a whole, which processes the input sequence and obtains the output feature.
[0057] Step S433: Use the sequence pooling layer 203 to summarize the output of the Transformer encoder 202 into a fixed-length vector. The role of the sequence pooling layer 203 is to compress the information of the entire sequence into a compact representation for subsequent classification.
[0058] Step S434: Use the first fully connected layer 204 to convert the fixed-length vector into a classification result. The fully connected layer is a common layer in neural networks and is used to convert the input features into output results. In this step, the classification result can be healthy, unhealthy, or other classification results such as healthy, suspected, highly suspected, etc.
[0059] As mentioned above, both the segmentation model and the classification model need to be trained. The following introduces the training process in combination with the actual application scenario for reference.
[0060] In specific applications, three different breast cancer pathological image datasets were used for the training and validation of the segmentation model and the classification model, namely the SHOW dataset, the TNBC dataset, and the BreaKHis dataset. Among them, the SHOW dataset is a large-scale synthetic pathological image dataset, combined with nuclear semantic segmentation annotations, called the Synthetic Nucleus and Annotation Wizard. It has a sample size of 20,000 images and is mainly used for the training of the segmentation model; the TNBC dataset comes from 11 triple-negative breast cancer patients, representing the differences among patients of the same cancer type. Many cells are detailedly recorded in the dataset, and the sample size is 4,022 images, which is mainly used for the validation of the segmentation model; the BreaKHis dataset comes from 82 patients, including benign and malignant images, and is mainly used for the training and validation of the classification model. Specifically, in the BreaKHis dataset, 248 benign images and 531 malignant images were designated as test samples and randomly divided. The remaining 2,231 benign images and 4,773 malignant images were randomly divided into a training set and a validation set.
[0061] During the training process of the segmentation model, for example, the Adam optimizer can be combined with the StepLR scheduler to adjust the learning rate according to the iteration step. Specifically, the learning rate is multiplied by 0.1 every 30 iterations. In addition, the random seeds are set to 21, 42, 84 (to ensure reproducibility), the number of epochs is set to 100, and a loss function that combines binary cross-entropy loss (BCE) and Dice score in a 1:1 ratio is adopted. Selecting different random seeds will result in different segmentation effects of the segmentation model. Table 1 shows the Dice (Dice coefficient, used to evaluate the accuracy of the segmentation result), Jaccard (Jaccard index, used to evaluate the segmentation precision), and Hausdorff (Hausdorff distance, used to evaluate the boundary precision of the segmentation result) corresponding to different random seeds (seed). In a specific embodiment of the present invention, when the random seed is set to 42, the comprehensive performance of the segmentation model is the best.
[0062] Table 1: Segmentation model effects corresponding to different random seeds
[0063] Seed Dice Jaccard Hausdorff 21 0.78649 0.65848 7.30393 42 0.79294 0.66642 7.30142 84 0.77985 0.65219 7.17857
[0064] For the entire classification method, the nuclear segmentation and image enhancement in steps S200 and S300 can be denoted as nie, the wavelet position embedding module 201 can be denoted as wpe, and the classification model without the introduction of the wavelet position embedding module 201 can be denoted as EVT. Therefore, the classification method in the present invention is actually EVT + wpe + nie. To evaluate the performance impact of the nie module and the wpe module on this classification method, we specifically compared the classification effects when the nie module and / or the wpe module were missing, as shown in the following table, where EVT means that the wavelet position embedding module 201 was not introduced and image enhancement was not performed, EVT + wpe means that the wavelet position embedding module 201 was introduced but image enhancement was not performed, and EVT + nie means that the wavelet position embedding module 201 was not introduced but image enhancement was performed.
[0065] Table 2: Comparison of the impact of the nie and wpe modules on the performance of the EVT model
[0066]
[0067] As can be seen from Table 2, the nie module focuses on enhancing the nuclear information features in the image, enabling the model to more accurately identify and distinguish the key nuclear structures in the image. The enhancement effect of EVT + nie is relatively small. When the nie module and the wpe module are jointly applied to the EVT model, the performance of the EVT + wpe + nie model is further improved, and the accuracy rate reaches 0.9487. The experimental results of the confusion matrix of the experimental records are shown as Figure 7 shown, and the kappa value is calculated. The kappa value of the EVT + wpe + nie model has increased significantly, indicating a slight improvement in the model consistency. The ROC curve ( Figure 8 ) was further plotted. The best classification chi-square value of the blue curve is as high as 0.99, which is significantly higher than other methods. This result shows that the synergistic effect of the two modules can more comprehensively mine the image features, enabling the EVT model to have stronger classification ability when processing breast cancer pathological data with BreaKHis as an example.
[0068] It should be noted that the step division of the above various methods is only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as the same logical relationship is included, they are all within the protection scope of this application; adding insignificant modifications to the algorithm or process or introducing insignificant designs, but not changing the core design of its algorithm and process, are all within the protection scope of this patent.
[0069] Please refer to Figure 9 , Figure 9A breast cancer pathological image classification device provided by an embodiment of the present invention includes: a data acquisition unit 301, a first processing unit 302, a second processing unit 303, and a third processing unit 304; wherein, the data acquisition unit 301 is used to acquire the pathological image to be processed; the first processing unit 302 is used to extract the cell nucleus region in the pathological image by using a segmentation model to obtain a segmented image; the second processing unit 303 is used to extract the cell nucleus contour from the segmented image and fuse the cell nucleus contour with the pathological image to obtain an enhanced image; the third processing unit 304 is used to perform classification processing on the enhanced image by using a classification model to obtain a classification result, wherein the classification model is an enhanced vision transformer embedded with wavelet positions.
[0070] It should be noted that the classification device in this embodiment is a device corresponding to the above classification method, and the functional modules in the classification device respectively correspond to the corresponding steps in the classification method. The classification device in this embodiment can be implemented in cooperation with the classification method. That is, without conflict, the relevant technical details mentioned in the classification method of the above embodiment can also be applied to the classification device in this embodiment.
[0071] Please refer to Figure 10 , Figure 10 An electronic device provided by an embodiment of the present invention includes a processor 401, a memory 402, and a communication bus; the communication bus is used to connect the processor 401 and the memory 402; the processor 401 is used to execute the computer program stored in the memory 402 to implement the above breast cancer pathological image classification method.
[0072] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored, and the computer program is used to make a computer execute the above breast cancer pathological image classification method.
[0073] Generally speaking, in the present invention, an enhanced vision transformer embedded with wavelet positions is used as the classification model, which is organically combined with the cell nucleus information enhancement strategy. Through rigorous experimental design and verification, it is successfully confirmed the excellent performance and significant advantages demonstrated by this collaborative combination in processing complex pathological images, opening up a new path for the development of complex pathological image classification technology and overcoming the technical bottleneck of the collaborative application of wavelet transform and cell nucleus enhancement.
[0074] The above embodiments only illustrate the principles and effects of the present invention, rather than limiting the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.
Claims
1. A breast cancer pathology image classification method, characterized in that: include: Acquiring a pathological image to be processed; Extracting the cell nucleus region in the pathological image using the segmentation model to obtain a segmented image; Extracting the cell nucleus contour from the segmented image, and fusing the cell nucleus contour with the pathological image to obtain an enhanced image; The enhanced image is classified using a classification model to obtain a classification result, wherein the classification model is an enhanced visual transformer embedded in the wavelet position.
2. The breast cancer pathology image classification method according to claim 1, characterized in that: The cell nucleus region in the pathological image is extracted using a segmentation model to obtain a segmented image, including: The cell nucleus region in the pathological image is extracted using a Unet model, a ResUnet model, or a UTnet model to obtain a segmented image.
3. The breast cancer pathology image classification method according to claim 1, characterized in that: The classification model includes a marker and a classifier; The enhanced image is classified using a classification model to obtain a classification result, including: Normalizing the enhanced image to convert the enhanced image into a tensor; Converting the tensor into low-dimensional features using the tagger; The low-dimensional features are processed using the classifier to obtain the classification result.
4. The breast cancer pathology image classification method according to claim 3, characterized in that: The marker is formed by cascading multiple convolution blocks, and each of the convolution blocks is formed by cascading a first convolution layer, an activation layer, and a maximum pooling layer.
5. The breast cancer pathology image classification method according to claim 3, characterized in that: The classifier includes a cascaded wavelet position embedding module, a transformer encoder, a sequence pooling layer, and a first fully connected layer; Processing the low-dimensional features using the classifier to obtain the classification result includes: Decomposing the low-dimensional features into components of different frequencies using the wavelet position embedding module, and embedding position information to enhance feature representation, thereby obtaining an input sequence; processing the input sequence using the transformer encoder to capture long-range dependencies and contextual information in the input sequence; aggregating the output of the transformer encoder into a vector of fixed length using the sequence pooling layer; The fixed-length vector is converted into the classification result using the first fully connected layer.
6. The breast cancer pathology image classification method according to claim 5, characterized in that: The wavelet position embedding module includes a cascaded transposition layer, a second fully connected layer, a second convolutional layer, a trigonometric function connection layer, and a third fully connected layer; Decomposing the low-dimensional features into components of different frequencies using the wavelet position embedding module includes: The low-dimensional features are processed in sequence using the cascaded transposition layer, the second fully connected layer, the second convolutional layer, the trigonometric function connection layer, and the third fully connected layer, and the processed features are added to the low-dimensional features and then output, so as to decompose the low-dimensional features into components of different frequencies.
7. The breast cancer pathology image classification method according to claim 5, characterized in that: The transformer encoder includes a normalization layer, a multi-head self-attention module, and a multi-layer perceptron; Processing the input sequence using the transformer encoder comprises: Processing the input sequence using a cascaded normalization layer and a multi-head self-attention module to obtain a first feature; Adding the input sequence and the first feature to obtain a second feature; Processing the second feature using a cascaded normalization layer and a multi-layer perceptron to obtain a third feature; Adding the second feature and the third feature to obtain a fourth feature; The fourth feature is taken as a new input sequence and processed again according to the above steps to obtain the output feature of the transformer encoder.
8. A breast cancer pathology image classification device, characterized in that: include: A data acquisition unit, used for acquiring a pathological image to be processed; A first processing unit, configured to extract a cell nucleus region in the pathological image using a segmentation model to obtain a segmented image; A second processing unit is used to extract the cell nucleus contour from the segmented image, and fuse the cell nucleus contour with the pathological image to obtain an enhanced image; The third processing unit is used to classify the enhanced image using a classification model to obtain a classification result, wherein the classification model is an enhanced visual transformer embedded in the wavelet position.
9. An electronic device, characterized in that: The system comprises a processor, a memory and a communication bus; the communication bus is used to connect the processor and the memory; the processor is used to execute a computer program stored in the memory to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and the computer program is used to make a computer execute the method according to any one of claims 1 to 7.
Citation Information
Cited By
Quantitative index determination method and device based on hepatitis pathology image, equipment and product
CN121213551A