Analytical dichotomy pathological image quality control method based on prototype learning

Through the encoder-decoder architecture and multi-loss training based on prototype learning, the problems of low manual sampling efficiency in pathological image quality control and lack of interpretability in deep learning are solved, and high-precision and transparent pathological image quality control are achieved, improving the efficiency and accuracy of pathological diagnosis.

CN120495730APending Publication Date: 2025-08-15GUILIN UNIV OF ELECTRONIC TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510514432.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the prior art, pathological image quality control relies on manual sampling to have problems such as low efficiency, inconsistent results, high cost and inability to fully evaluate. In addition, deep learning methods lack interpretability in pathological image quality control, affecting the accuracy and transparency of computer-assisted diagnosis.

Method used

Using an encoder-decoder architecture based on prototype learning, image features are extracted through convolutional layers, bottleneck blocks and residual blocks, combined with the prototype update layer to dynamically optimize category prototypes, and using multiple loss functions to train the model to achieve high-precision and interpretable pathological image quality control.

Benefits of technology

It improves the efficiency and accuracy of pathological image quality control, enhances the interpretability of the model, enables doctors to solve the problem of the strategy process, and significantly improves the efficiency and classification robustness of pathological diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495730A_ABST
    Figure CN120495730A_ABST
Patent Text Reader

Abstract

The invention discloses an interpretable dichotomy pathological image quality control method based on prototype learning, and the method comprises the steps: firstly carrying out the preprocessing of a full-width scanning pathological image, and extracting an effective region; an encoder extracts multi-level features through a convolution layer, a bottleneck block and a residual block, a prototype updating layer is embedded in a potential space, prototype vectors of focusing / out-of-focus categories are dynamically optimized, similarity vectors are generated by calculating the Euclidean distance between sample features and the prototypes, and the similarity vectors are used for calculating the sample features and the prototypes; inputting a linear classification layer output category probability; the decoder reconstructs the image through transposition convolution and jump connection, and complements details in combination with an optimization prototype; model training is combined with coding and decoding loss, classification loss and prototype loss, and finally high-precision classification is achieved. According to the method, image features are extracted through an encoder-decoder architecture, a category prototype is dynamically optimized in combination with a prototype learning mechanism, and high-precision classification and interpretability are achieved through multi-loss joint training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of digital pathology and deep learning, and in particular to an interpretable binary classification pathology image quality control method based on prototype learning. Background Art

[0002] Digital pathology, a key branch of modern medicine, digitizes traditional pathology slides into full-width scanned images, enabling remote diagnosis, long-term storage, and intelligent analysis of pathology images. However, the pathology slide preparation and digitization process is prone to quality issues, such as out-of-focus scans or blurred areas caused by tissue wrinkles, varying tissue thickness, bubbles, pen marks, dust, scratches, and other factors. These issues directly impact the reliability of subsequent computer-aided diagnosis systems.

[0003] In this context, pathology image quality control has become a key link in ensuring diagnostic accuracy. The traditional pathology image quality control process mainly relies on manual sampling, but this method faces severe challenges in the digital age. At present, pathology diagnosis and treatment centers and hospital pathology departments mainly rely on manual sampling to control the quality of pathology sections. Although this method is simple and effective and plays an important role in improving the level of preparation and the quality of sections, it also has several shortcomings: (1) Quality control personnel are required to make careful observations under a microscope and work for long hours every day, which can easily lead to missed detections and false detections due to visual fatigue; (2) For the same section, different quality control personnel may have different judgments, resulting in inconsistent inspection results and a lack of clear standardization and quantitative evaluation methods. This subjectivity may lead to inconsistent results and make it difficult to establish a unified quality control standard. At the same time, due to human resource limitations, it is impossible to conduct a comprehensive evaluation and analysis of a large number of pathology sections; (3) The sampling method cannot cover all pathology sections, and partial sample sampling cannot truly reflect the quality of all sections. Based on the above reasons, the traditional manual sampling method can no longer meet the high requirements of modern digital pathology development in terms of efficiency, accuracy, cost, etc.

[0004] With the advancement of digital scanning technology, pathology slides can be rapidly digitized, making digital pathology images easier to display and store. This provides the foundation and prerequisite for computer-assisted pathology image analysis. Compared to manual analysis, computers are highly efficient, free from fatigue, and more objective and impartial. In the field of digital pathology, artificial intelligence (AI) technology is being used to assist physicians with pathological diagnosis, pathology image analysis, and lesion localization. As digital pathology research continues to deepen, scholars both domestically and internationally are beginning to apply AI technology to the quality control of pathology slides. Rigorously validated AI quality control methods assist technicians in quality control of pathology slides, including screening out substandard slides and detecting out-of-focus images.

[0005] With the widespread integration of deep learning techniques in medical and healthcare applications, particularly in the field of digital pathology diagnosis, diagnostic and prognostic capabilities have been significantly enhanced. However, these methods are computationally complex, and their effectiveness relies heavily on large annotated datasets. The process of acquiring and labeling large numbers of medical images is not only time-consuming but also costly. As machine learning algorithms become increasingly important for important societal issues, interpretability (transparency) has become a key issue in whether we can trust the predictions from these models. In some cases, the model's decision-making process is not transparent.

[0006] Therefore, there is an urgent need to introduce interpretable pathology image quality control methods based on artificial intelligence and deep learning technologies to improve the efficiency of pathology slide quality control. By using a model with simple algorithms, it is possible to automate the analysis of pathology images and quickly identify in-focus and out-of-focus images. This not only improves detection accuracy but also ensures transparency of the results, allowing doctors to understand the model's decision-making process and significantly enhance the efficiency of pathology diagnosis. Summary of the Invention

[0007] The main purpose of the present invention is to overcome the shortcomings and deficiencies of the existing technology and provide an interpretable binary classification pathology image quality control method based on prototype learning. Image features are extracted through an encoder-decoder architecture, and the category prototypes are dynamically optimized in combination with the prototype learning mechanism. Multi-loss joint training is used to achieve high-precision classification and interpretability.

[0008] In order to achieve the above object, the present invention adopts the following technical solutions:

[0009] In a first aspect, the present invention provides an interpretable binary classification pathology image quality control method based on prototype learning, comprising the following steps:

[0010] Acquire the full-width scanned pathological image and perform preprocessing to obtain the effective area of the digital pathological image;

[0011] An interpretable binary classification model was constructed, specifically by inputting the valid region of the digital pathology image into an encoder to extract features. The encoder comprises convolutional layers, bottleneck blocks, residual blocks, and transition layers, and generates latent space features through multi-level feature fusion. A prototype update layer is embedded in the encoder's latent space to dynamically optimize the prototype vectors of the focused and out-of-focus categories. A similarity vector is generated by calculating the Euclidean distance between the sample features and the prototypes. This similarity vector is then input into a linear classification layer, which outputs the probability of the focused or out-of-focus category. Simultaneously, the latent space features are reconstructed into the original input image through transposed convolutions and skip connections in the decoder, and the image details are completed using the optimized prototypes.

[0012] The model is trained by combining the sum of encoder-decoder loss, classification loss and prototype loss, and the trained model is used to perform binary classification on the pathological images to be processed.

[0013] As a preferred technical solution, the preprocessing is to adjust the pathological image to a fixed size, then convert the adjusted image from the PIL image format to the tensor format, and finally normalize the image in the tensor format.

[0014] As a preferred technical solution, the encoder is specifically:

[0015] The convolutional layer consists of two convolution blocks, which process the input RGB image respectively. Each convolution block is followed by batch normalization and ReLU activation function.

[0016] The residual block includes multiple bottleneck blocks connected in series, each bottleneck block is composed of a 1×1 convolutional layer, a 3×3 convolutional layer and a 1×1 convolutional layer stacked in sequence, and is used to perform deep extraction of input features;

[0017] The convolution operation with a stride of 2 reduces the spatial size of the input feature map to half, while adjusting the number of channels to optimize the feature expression capability.

[0018] As a preferred technical solution, the prototype update layer is specifically:

[0019] Prototypes are randomly generated during model initialization and are learnable parameters. Their initial values will be updated through backpropagation during training. In each training cycle, the model first extracts features from the input sample through the encoder, and then obtains the global feature representation of each sample through a pooling operation. By calculating the Euclidean distance between each sample feature and the prototype, the model can understand the similarity between the sample and the prototype of each category. After multiple iterations, the model will continuously update the prototype based on the sample features and loss feedback, so that it gradually converges to the optimal position that can effectively distinguish different categories.

[0020] As a preferred technical solution, the decoder consists of transposed convolution and skip connection. Each layer gradually upsamples the potential features output by the encoder through preset input channel number, output channel number, kernel size, stride and padding parameters, and uses cross-layer feature fusion technology to splice the feature maps of different stages of the encoder with the upsampling results until the spatial size of the reconstructed image is consistent with the original input image.

[0021] After each convolution transposition operation, the ReLU activation function is applied, and the final output layer uses the Sigmoid activation function to normalize the pixel values of the reconstructed image to a preset range;

[0022] The hierarchical parameter configuration of the decoder satisfies the following requirements: the number of output channels of the first-layer convolution transpose operation is less than the number of input channels, and the stride is greater than 1, so as to achieve size expansion of the feature map and adjustment of the number of channels.

[0023] As a preferred technical solution, the decoder loss is a standard mean square error loss, that is, the square distance L2 between the original input and the reconstructed input is used to penalize the reconstruction error of the codec;

[0024] The classification loss is used to directly optimize the model's focus / out-of-focus binary classification performance for pathological images;

[0025] The prototype loss includes aggregation loss and dispersion loss. The aggregation loss represents the loss component of the distance between the aggregation sample features of the same type and the prototype, and the dispersion loss represents the loss component of the distance between the dispersion heterogeneous prototypes.

[0026] As a preferred technical solution, the prototype loss is specifically:

[0027]

[0028] Where d is the number of prototypes retained for each category, N is the number of categories, and x i is the i-th sample, y i is the true label p of the i-th sample k is the prototype of category j, P j is the prototype set of category j.

[0029] In a second aspect, the present invention provides an interpretable two-class pathology image quality control system based on prototype learning, which is applied to the above-mentioned interpretable two-class pathology image quality control method based on prototype learning, including a preprocessing module, a model building module and a model training module;

[0030] The preprocessing module is used to obtain the full-scan pathology image and perform preprocessing to obtain the effective area of the digital pathology image;

[0031] The model construction module is used to construct an interpretable binary classification model. Specifically, the encoder is input into the valid area of the digital pathology image to extract features. The encoder includes a convolutional layer, a bottleneck block, a residual block, and a transition layer, and generates latent space features through multi-level feature fusion. The prototype update layer is embedded in the latent space of the encoder to dynamically optimize the prototype vectors of the focused and out-of-focus categories. The similarity vector is generated by calculating the Euclidean distance between the sample features and the prototype. The similarity vector is input into the linear classification layer, and the probability of the focused or out-of-focus category is output. At the same time, the latent space features are reconstructed into the original input image through the transposed convolution and skip connection of the decoder, and the image details are completed by combining the optimized prototype.

[0032] The model training module is used to perform model training by combining the sum of codec loss, classification loss and prototype loss, and to perform binary classification on the pathological image to be processed using the trained model.

[0033] In a third aspect, the present invention provides an electronic device, comprising:

[0034] at least one processor; and,

[0035] a memory communicatively connected to the at least one processor; wherein,

[0036] The memory stores computer program instructions that can be executed by the at least one processor, and the computer program instructions are executed by the at least one processor to enable the at least one processor to perform the prototype learning-based interpretable binary classification pathology image quality control method.

[0037] In a fourth aspect, the present invention provides a computer-readable storage medium storing a program, which, when executed by a processor, implements the prototype learning-based interpretable binary classification pathology image quality control method.

[0038] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0039] (1) This paper proposes a robust classification mechanism based on dynamic prototype learning. It introduces a prototype update layer and dynamically optimizes the category prototypes through aggregation loss and dispersion loss, forcing similar samples to cluster closely in the latent space and heterogeneous prototypes to stay away from each other. This mechanism significantly improves the classification robustness of the model in scenarios with class imbalance and small sample sizes. At the same time, it intuitively displays the classification basis through prototype visualization, enhancing credibility.

[0040] (2) This invention achieves image reconstruction through encoder-decoder joint training and prototype association analysis. This design enables doctors to verify the model decision logic and solves the "black box" problem of traditional deep learning models;

[0041] (3) This paper proposes a lightweight multi-level feature fusion technology: the encoder uses the Bottleneck Block module (1x1→3x3→1x1 convolution) and the Transition Layer to reduce the number of parameters while fusing shallow detail features with deep semantic features to enhance sensitivity to fuzzy areas. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0043] Figure 1 Schematic diagram of common quality problems of pathological sections (such as tissue wrinkles, bubbles, and scratches) in an embodiment of the present invention;

[0044] Figure 2 This is a flow chart of an interpretable binary classification pathology image quality control method based on prototype learning according to an embodiment of the present invention;

[0045] Figure 3 This is a diagram showing the overall architecture of an interpretable binary classification pathology image quality control method based on prototype learning according to an embodiment of the present invention;

[0046] Figure 4 Schematic diagram of the performance of various methods according to the embodiments of the present invention on 10%, 25%, 50% and 100% of the original data set;

[0047] Figure 5 Schematic diagram of the structure of an interpretable two-class pathology image quality control system based on prototype learning according to an embodiment of the present invention;

[0048] Figure 6 Schematic diagram of the electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0049] In order to enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0050] References to "embodiments" in this application mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described in this application may be combined with other embodiments.

[0051] With the widespread application of digital technology in the medical field, digital pathology diagnosis technology has been gradually promoted in clinical practice. However, the preparation of pathology slices and the acquisition of whole slide images (WSI) often have defects such as tissue wrinkles, bubbles, and scratches. Figure 1 As shown in parts (a) to (h) of the figure, these defects can cause out-of-focus or blurred areas in the image, which in turn affects the accuracy of automated analysis in computational pathology. To address this problem, the present invention proposes an interpretable two-classification quality detection method based on prototype learning. This method optimizes the model's recognition performance for blurred pathology images through a prototype learning mechanism, while innovatively introducing a decoder design to enhance the interpretability of the model's decision-making process. Experimental results show that this method achieves a classification accuracy of 97.21% on the FocusPath dataset, which is a significant improvement over existing technologies and provides a more effective solution for pathology image quality control.

[0052] like Figure 2 、 Figure 3 As shown, in one embodiment of the present application, a method for interpretable binary classification pathology image quality control based on prototype learning is provided, comprising the following steps:

[0053] S1 acquires the WSI image and performs preprocessing to obtain the effective area of the digital pathology image;

[0054] Furthermore, the preprocessing involved resizing the image to a fixed size of 228x228, converting the image from PIL image format to tensor format, and finally normalizing the tensor to scale the pixel values to [0, 1]. Given the plentiful data available for the experiment, no data augmentation was performed.

[0055] S2. Prepare training and test sets. Specifically, the FocusPath public dataset used contains 8,640 annotated in-focus and out-of-focus pathology images. In the dataset, if the absolute value of the label score of an image is less than or equal to 2, the image is considered in-focus; otherwise, it is considered out-of-focus. The dataset is divided into training and test sets in a ratio of 4:1.

[0056] S3. Construct an interpretable binary classification model. This model uses an encoder-decoder architecture to extract image features and dynamically optimizes category prototypes in combination with a prototype learning mechanism. Specifically, the effective area of the digital pathology image is input into the encoder to extract features. The encoder includes a convolutional layer, a bottleneck block, a residual block, and a transition layer, and generates latent space features through multi-level feature fusion. In particular, a prototype update layer is embedded in the encoder's latent space to dynamically optimize the prototype vectors of the focused and out-of-focus categories. A similarity vector is generated by calculating the Euclidean distance between the sample features and the prototype, and the similarity vector is input into the linear classification layer to output the category probability of focused or out-of-focus. The latent space features are reconstructed into the original input image through the transposed convolution and jump connection of the decoder, and the image details are completed in combination with the optimized prototype.

[0057] Furthermore, the model first transforms the input sample into a compressed latent space through an encoder, while simultaneously extracting feature representations of the input sample. Specifically, the encoder is a novel image feature extraction network consisting of a convolutional layer, a bottleneck block, a residual block (Layer 1), and a transition layer. The initial portion of the network consists of two convolutional blocks, each processing the input RGB image, gradually reducing the spatial dimensionality and enhancing feature representation. Each convolutional block is followed by batch normalization and a ReLU activation function to ensure the model's nonlinear capabilities and training convergence. The network utilizes a Bottleneck Block module, which combines 1x1 convolution, 3x3 convolution, and another 1x1 convolution to effectively reduce the number of parameters and computational complexity, enabling the extraction of deeper features with fewer parameters, thereby improving the network's expressiveness. The combination of Layer 1 and the Transition Layer modules enables the encoder network to efficiently transfer information between feature layers of different resolutions. The introduction of the transition layer allows for more flexible adjustment of the number of channels, helping to improve feature representation while preserving detail.

[0058] Furthermore, the prototype update layer, specifically, the prototype update layer is embedded in the latent space of the encoder, and its functions include prototype initialization, dynamic optimization and similarity measurement. Prototype initialization refers to the allocation of d prototypes (d is a hyperparameter) to each category (focus / out of focus), and the prototype vector is randomly selected from the training sample features of the corresponding category. Similarity measurement refers to calculating the Euclidean distance between the input feature vector and all prototypes to generate a similarity vector. The dynamic optimization mechanism refers to the dynamic optimization of category prototypes through aggregation loss and dispersion loss during training, forcing similar samples to be closely clustered in the latent space and heterogeneous prototypes to stay away from each other. The prototype is updated through backpropagation, and the optimization goal is to minimize the distance between similar samples and the prototype, while maximizing the distance between heterogeneous prototypes. This mechanism significantly improves the classification robustness of the model in class imbalance and small sample scenarios.

[0059] Furthermore, the linear classification layer converts the m-dimensional vector generated by the prototype update layer into a probability distribution over N categories. This layer consists of a fully connected layer and a softmax layer to ensure that the sum of the output probabilities is 1. The entire network, including the prototype update layer and the linear classification layer, is trained jointly.

[0060] S4. Designing the loss function to constrain the model training process

[0061] The model is trained from scratch on the training set. Since the network consists of three parts, the loss function used to optimize the network also consists of three parts, namely encoder-decoder loss, classification loss, and prototype loss. The encoder-decoder loss (AE-Loss) is the standard mean squared error loss. AE-loss can be expressed as:

[0062]

[0063] In order to ensure the consistency between the prototype and the training data, the present invention designs a prototype loss. The first component of the prototype loss is the aggregation loss, which is used to measure the similarity between samples of the same category and their prototypes. The goal is to cluster similar data points near the same prototype. The second component is the dispersion loss, which is used to ensure that each prototype can be optimized around the cluster center of its category samples, thereby reflecting the overall characteristics of the category. It helps the model maintain a close relationship between the prototype and its category samples while maintaining an appropriate distance from other category prototypes. However, for unbalanced data sets (i.e., the number of samples in some categories is significantly greater than that in other categories), the present invention emphasizes optimizing the prototype at the class level during training rather than optimizing on the entire data set. Specifically, the optimization process focuses on the distance between the prototype of each class and the samples of that class. The prototype loss is expressed as:

[0064]

[0065] Finally, in order to ensure the best classification performance, the standard Cross-Entropy loss is used to optimize the model. The classification loss is expressed as:

[0066]

[0067] The complete loss function is expressed as:

[0068] loss=loss class +λ1*loss AE +λ2*loss p

[0069] Comparative experiment:

[0070] Experimentally, we compared our method with several existing methods on the publicly available FocusPath dataset, using ROC-AUC, PR-AUC, Accuracy, Precision, Recall, and F1-score as our evaluation metrics. Our method outperformed all of the compared methods, as shown in Table 1. Our method performed exceptionally well across all metrics, achieving a top accuracy of 97.21%, demonstrating its robust performance on the FocusPath dataset.

[0071] Table 1 Performance of all networks on the FocusPath dataset

[0072]

[0073] In order to verify that the method still performs well in the case of sparse data, the present invention compares the performance of various methods on 10%, 25%, 50% and 100% of the data set. The results show that the classification accuracy of this method is superior to other methods in the case of sparse data. Figure 4 shown.

[0074] As shown in Table 2, the present invention also comprehensively evaluates the performance of the model by showing the differences between the input image, the reconstructed image, and the prototype distance distribution. The input image presents the real sample features, and the reconstructed image is obtained by inverse processing the features through the decoder. By comparing the distances of each category, it can be seen that the prototype distance of the input image and its predicted category (such as Class 1) is closer, indicating that the model has a higher confidence in this category. The distance to another category (such as Class 0) is significantly larger, proving that the image is unlikely to belong to this category. This result quantitatively supports the effectiveness of the model's classification decision.

[0075] Table 2 Input image, reconstruction difference and prototype distance distribution

[0076]

[0077] In summary, this paper provides an interpretable binary pathology image quality control method based on prototype learning. Through dynamic prototype learning and interpretable design, it achieves excellent results in the classification of out-of-focus and in-focus quality in digital pathology images. Furthermore, even in the case of data scarcity, this method demonstrates superior classification accuracy. Furthermore, by demonstrating the differences between input and reconstructed images and the prototype distance distribution, we comprehensively evaluate the model's performance, demonstrating its strong practical value.

[0078] It should be noted that, for the sake of convenience, the aforementioned method embodiments are all expressed as a series of action combinations, but those skilled in the art should know that the present invention is not limited to the described order of actions, because according to the present invention, certain steps can be performed in other orders or simultaneously.

[0079] Based on the same concept as the prototype learning-based interpretable two-classification pathology image quality control method in the above-mentioned embodiment, the present invention also provides a prototype learning-based interpretable two-classification pathology image quality control system, which can be used to execute the above-mentioned prototype learning-based interpretable two-classification pathology image quality control method. For ease of explanation, the structural diagram of the embodiment of the prototype learning-based interpretable two-classification pathology image quality control system only shows the parts related to the embodiment of the present invention. Those skilled in the art will understand that the illustrated structure does not constitute a limitation of the device, and it may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0080] See also Figure 5 In another embodiment of the present application, an interpretable binary classification pathology image quality control system 100 based on prototype learning is provided, the system comprising a pre-processing module 101, a model building module 102 and a model training module 103;

[0081] The preprocessing module is used to obtain the full-scan pathology image and perform preprocessing to obtain the effective area of the digital pathology image;

[0082] The model construction module is used to construct an interpretable binary classification model. Specifically, the encoder is input into the valid area of the digital pathology image to extract features. The encoder includes a convolutional layer, a bottleneck block, a residual block, and a transition layer, and generates latent space features through multi-level feature fusion. The prototype update layer is embedded in the latent space of the encoder to dynamically optimize the prototype vectors of the focused and out-of-focus categories. The similarity vector is generated by calculating the Euclidean distance between the sample features and the prototype. The similarity vector is input into the linear classification layer, and the probability of the focused or out-of-focus category is output. At the same time, the latent space features are reconstructed into the original input image through the transposed convolution and skip connection of the decoder, and the image details are completed by combining the optimized prototype.

[0083] The model training module is used to perform model training by combining the sum of codec loss, classification loss and prototype loss, and to perform binary classification on the pathological image to be processed using the trained model.

[0084] It should be noted that the interpretable two-class pathology image quality control system based on prototype learning of the present invention corresponds one-to-one to the interpretable two-class pathology image quality control method based on prototype learning of the present invention. The technical features and beneficial effects described in the above-mentioned embodiment of the interpretable two-class pathology image quality control method based on prototype learning are applicable to the embodiment of the interpretable two-class pathology image quality control based on prototype learning. For specific contents, please refer to the description in the embodiment of the method of the present invention. No further details will be given here. This is hereby declared.

[0085] In addition, in the implementation of the interpretable two-class pathology image quality control system based on prototype learning in the above-mentioned embodiment, the logical division of each program module is only an example. In actual application, the above-mentioned functions can be assigned to different program modules as needed, for example, for the convenience of corresponding hardware configuration requirements or software implementation. That is, the internal structure of the interpretable two-class pathology image quality control system based on prototype learning is divided into different program modules to complete all or part of the functions described above.

[0086] See also Figure 6 In one embodiment, an electronic device for implementing an interpretable binary classification pathology image quality control method based on prototype learning is provided. The electronic device 200 may include a first processor 201, a first memory 202 and a bus, and may also include a computer program stored in the first memory 202 and executable on the first processor 201, such as an interpretable binary classification pathology image quality control program 203 based on prototype learning.

[0087] The first memory 202 includes at least one type of readable storage medium, including flash memory, a mobile hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the first memory 202 can be an internal storage unit of the electronic device 200, such as a mobile hard disk of the electronic device 200. In other embodiments, the first memory 202 can also be an external storage device of the electronic device 200, such as a plug-in mobile hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 200. Furthermore, the first memory 202 can also include both an internal storage unit of the electronic device 200 and an external storage device. The first memory 202 can not only be used to store application software and various types of data installed in the electronic device 200, such as the code of the prototype learning-based interpretable binary classification pathology image quality control program 203, but can also be used to temporarily store data that has been output or is about to be output.

[0088] In some embodiments, the first processor 201 may be composed of an integrated circuit, for example, a single packaged integrated circuit, or a plurality of packaged integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The first processor 201 is the control core (Control Unit) of the electronic device, connecting the various components of the entire electronic device using various interfaces and lines, and executing or executing programs or modules stored in the first memory 202, as well as calling data stored in the first memory 202, to perform various functions of the electronic device 200 and process data.

[0089] Figure 6 Only the electronic device with components is shown, and it can be understood by those skilled in the art that Figure 6 The structure shown does not constitute a limitation on the electronic device 200 , and the electronic device 200 may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.

[0090] The prototype learning-based interpretable binary classification pathology image quality control program 203 stored in the first memory 202 of the electronic device 200 is a combination of multiple instructions. When executed in the first processor 201, the following can be achieved:

[0091] Acquire the full-width scanned pathological image and perform preprocessing to obtain the effective area of the digital pathological image;

[0092] An interpretable binary classification model was constructed, specifically by inputting the valid region of the digital pathology image into an encoder to extract features. The encoder comprises convolutional layers, bottleneck blocks, residual blocks, and transition layers, and generates latent space features through multi-level feature fusion. A prototype update layer is embedded in the encoder's latent space to dynamically optimize the prototype vectors of the focused and out-of-focus categories. A similarity vector is generated by calculating the Euclidean distance between the sample features and the prototypes. This similarity vector is then input into a linear classification layer, which outputs the probability of the focused or out-of-focus category. Simultaneously, the latent space features are reconstructed into the original input image through transposed convolutions and skip connections in the decoder, and the image details are completed using the optimized prototypes.

[0093] The model is trained by combining the sum of encoder-decoder loss, classification loss and prototype loss, and the trained model is used to perform binary classification on the pathological images to be processed.

[0094] Furthermore, if the modules / units integrated in the electronic device 200 are implemented as software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium. The computer-readable medium may include any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard drive, a magnetic disk, an optical disk, a computer memory, or a read-only memory (ROM).

[0095] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0096] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0097] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be considered as equivalent replacement methods and are included in the scope of protection of the present invention.

Claims

1. An interpretable binary classification pathology image quality control method based on prototype learning, characterized by: The steps include: Acquire the full-width scanned pathological image and perform preprocessing to obtain the effective area of the digital pathological image; An interpretable binary classification model was constructed, specifically by inputting the valid region of the digital pathology image into an encoder to extract features. The encoder comprises convolutional layers, bottleneck blocks, residual blocks, and transition layers, and generates latent space features through multi-level feature fusion. A prototype update layer is embedded in the encoder's latent space to dynamically optimize the prototype vectors of the focused and out-of-focus categories. A similarity vector is generated by calculating the Euclidean distance between the sample features and the prototypes. This similarity vector is then input into a linear classification layer, which outputs the probability of the focused or out-of-focus category. Simultaneously, the latent space features are reconstructed into the original input image through transposed convolutions and skip connections in the decoder, and the image details are completed using the optimized prototypes. The model is trained by combining the sum of encoder-decoder loss, classification loss and prototype loss, and the trained model is used to perform binary classification on the pathological images to be processed.

2. The interpretable binary classification pathology image quality control method based on prototype learning according to claim 1 is characterized in that: The preprocessing is to adjust the pathological image to a fixed size, then convert the adjusted image from the PIL image format to the tensor format, and finally normalize the image in the tensor format.

3. The interpretable binary classification pathology image quality control method based on prototype learning according to claim 1 is characterized in that: The encoder is specifically: The convolutional layer consists of two convolution blocks, which process the input RGB image respectively. Each convolution block is followed by batch normalization and ReLU activation function. The residual block includes multiple bottleneck blocks connected in series, each bottleneck block is composed of a 1×1 convolutional layer, a 3×3 convolutional layer and a 1×1 convolutional layer stacked in sequence, and is used to perform deep extraction of input features; The convolution operation with a stride of 2 reduces the spatial size of the input feature map to half, while adjusting the number of channels to optimize the feature expression capability.

4. The interpretable binary classification pathology image quality control method based on prototype learning according to claim 1 is characterized in that: The prototype update layer is specifically: Prototypes are randomly generated during model initialization and are learnable parameters. Their initial values are updated through backpropagation during training. In each training cycle, the model first extracts features from the input sample through the encoder, and then obtains the global feature representation of each sample through pooling operations. By calculating the Euclidean distance between each sample feature and the prototype, the model can understand the similarity between the sample and the prototype of each category; after multiple iterations, the model will continuously update the prototype based on the sample features and loss feedback, so that it gradually converges to the optimal position that can effectively distinguish different categories.

5. The interpretable binary classification pathology image quality control method based on prototype learning according to claim 1 is characterized in that: The decoder is composed of transposed convolution and skip connections. Each layer gradually upsamples the latent features output by the encoder through preset input channel number, output channel number, kernel size, stride and padding parameters. The feature maps of different stages of the encoder are spliced with the upsampled results using cross-layer feature fusion technology until the spatial size of the reconstructed image is consistent with the original input image. After each convolution transposition operation, the ReLU activation function is applied, and the final output layer uses the Sigmoid activation function to normalize the pixel values of the reconstructed image to a preset range; The hierarchical parameter configuration of the decoder satisfies the following requirements: the number of output channels of the first-layer convolution transpose operation is less than the number of input channels, and the stride is greater than 1, so as to achieve size expansion of the feature map and adjustment of the number of channels.

6. The interpretable binary classification pathology image quality control method based on prototype learning according to claim 1 is characterized in that: The decoder loss is a standard mean square error loss, which uses the squared distance L2 between the original input and the reconstructed input to penalize the reconstruction error of the codec; The classification loss is used to directly optimize the model's focus / out-of-focus binary classification performance for pathological images; The prototype loss includes aggregation loss and dispersion loss. The aggregation loss represents the loss component of the distance between the aggregation sample features of the same type and the prototype, and the dispersion loss represents the loss component of the distance between the dispersion heterogeneous prototypes.

7. The method for interpretable binary classification pathological image quality control based on prototype learning according to claim 6 is characterized in that: The prototype loss is specifically: Where d is the number of prototypes retained for each category, N is the number of categories, and x i is the i-th sample, y i is the true label p of the i-th sample k is the prototype belonging to category j, P j is the prototype set of category j.

8. An interpretable binary classification pathology image quality control system based on prototype learning, characterized by: An interpretable binary classification pathology image quality control method based on prototype learning applied to any one of claims 1-7, comprising a preprocessing module, a model building module, and a model training module; The preprocessing module is used to obtain the full-scan pathology image and perform preprocessing to obtain the effective area of the digital pathology image; The model construction module is used to construct an interpretable binary classification model. Specifically, the encoder is input into the valid area of the digital pathology image to extract features. The encoder includes a convolutional layer, a bottleneck block, a residual block, and a transition layer, and generates latent space features through multi-level feature fusion. The prototype update layer is embedded in the latent space of the encoder to dynamically optimize the prototype vectors of the focused and out-of-focus categories. The similarity vector is generated by calculating the Euclidean distance between the sample features and the prototype. The similarity vector is input into the linear classification layer, and the probability of the focused or out-of-focus category is output. At the same time, the latent space features are reconstructed into the original input image through the transposed convolution and skip connection of the decoder, and the image details are completed by combining the optimized prototype. The model training module is used to perform model training by combining the sum of codec loss, classification loss and prototype loss, and to perform binary classification on the pathological image to be processed using the trained model.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores computer program instructions that can be executed by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the prototype learning-based interpretable binary classification pathology image quality control method as described in any one of claims 1-7.

10. A computer-readable storage medium storing a program, characterized in that: When the program is executed by a processor, the interpretable binary classification pathology image quality control method based on prototype learning according to any one of claims 1 to 7 is implemented.