Cervical whole slide image detection method, device, and computer program product
By cropping and extracting features from whole cervical slide images, and combining cell detectors and slide classifiers, the problem of imbalance in supervision methods in existing technologies has been solved, thereby improving the accuracy and enabling widespread application of whole cervical slide image detection.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- THE HONG KONG UNIV OF SCI & TECH
- Filing Date
- 2025-08-15
- Publication Date
- 2026-06-04
AI Technical Summary
Existing technologies cannot fully utilize cell-level and slice-level supervision in cervical whole-section image detection, resulting in insufficient detection accuracy, making it difficult to meet actual detection needs and hindering the application of this technology in a wider range of scenarios.
By acquiring whole cervical cytology slides, cropping them to obtain image patches, and using a combination of cell detectors and slide classifiers for feature extraction and prediction, the accuracy of detection is improved by combining cell-level supervision and slide-level supervision.
It achieves finer-grained cell representation and slice-level prediction, improving the accuracy of computer-aided cervical whole-slice image detection and promoting the application of this technology in a wider range of scenarios.
Smart Images

Figure CN2025115016_04062026_PF_FP_ABST
Abstract
Description
Cervical whole-section image detection methods, equipment and computer program products Technical Field
[0001] This application relates to the field of image detection technology, and in particular to a method, device and computer program product for detecting whole cervical slice images. Background Technology
[0002] With the development of artificial intelligence (AI) technology, AI-assisted image detection has been widely applied. However, in AI-assisted detection of large-scale whole-section cytology images, the balance between cell-level and section-level supervision is often insufficient. This imbalance in supervision directly limits the effectiveness of AI models in large-scale whole-section cytology image detection. For example, in the analysis of sections with diverse cell morphologies and sparse lesion distribution, the model is prone to judgment bias. Furthermore, it makes it difficult to meet the accuracy requirements of actual detection tasks, hindering the wider application of this technology. Summary of the Invention
[0003] The main objective of this application is to provide a method, device, and computer program product for detecting whole cervical slide images, aiming to improve the accuracy of computer-aided whole cervical cytology slide image detection.
[0004] To achieve the above objectives, one aspect of this application proposes a method for detecting whole-section cervical images, comprising the following steps:
[0005] A whole-section image of cervical cytology to be detected is obtained; the whole-section image of cervical cytology is cropped to obtain image blocks; the image blocks are input into a cell detector to obtain preset cell images in the image blocks output by the cell detector; features are extracted from the preset cell images through a feature extraction network in the slice classifier, and the extracted features are predicted through a multi-instance learning classifier in the slice classifier to obtain the detection result of the whole-section image of cervical cytology.
[0006] To achieve the above objectives, another aspect of this application proposes a cervical cytology whole slide image detection device. The device includes: an acquisition module for acquiring a cervical cytology whole slide image to be detected; a cropping module for cropping the cervical cytology whole slide image to obtain image blocks; a cell detection module for inputting the image blocks into a cell detector to obtain preset cell images in the image blocks output by the cell detector; and a feature extraction and prediction module for extracting features from the preset cell images through a feature extraction network in a slide classifier, and predicting the extracted features through a multi-instance learning classifier in the slide classifier to obtain a cervical cytology whole slide image detection result.
[0007] To achieve the above objectives, another aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method for detecting whole-section cervical images.
[0008] To achieve the above objectives, another aspect of the embodiments of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for detecting whole-section cervical images.
[0009] To achieve the above objectives, another aspect of this application provides a computer program product that, when run in an electronic device, causes the electronic device to execute the above-described cervical whole-slice image detection method.
[0010] The embodiments of this application include at least the following beneficial effects:
[0011] This application provides a method, device, and computer program product for detecting cervical cytology whole slide images. The scheme acquires a whole slide image of cervical cytology to be detected, crops the whole slide image to obtain image blocks, inputs the image blocks into a cell detector, obtains preset cell images in the image blocks output by the cell detector, then extracts features from the preset cell images using a feature extraction network in a slide classifier, and predicts the extracted features using a multi-instance learning classifier in the slide classifier to obtain the detection result of the cervical cytology whole slide image. This scheme, by inputting image blocks into the cell detector, can detect preset cell information from a single slide, obtaining a fine-grained cell representation. Then, the feature extraction network in the slide classifier extracts features from the preset cell images in the obtained fine-grained cell representation, obtaining slide-level preset cell features. Finally, the slide-level preset cell features are input into a multi-instance learning classifier for prediction to obtain the detection result of the cervical cytology whole slide image. This disclosure combines cell-level supervision of cell detectors (driving the model to learn fine-grained cell features) with slice-level supervision of slice classifiers (focusing on guiding the model to learn global judgment logic at the slice level), thereby making full use of cell-level and slice-level supervision, improving the accuracy of computer-aided cervical whole slice image detection, and promoting the application of this technology in a wider range of scenarios.
[0012] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0013] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:
[0014] Figure 1 is a flowchart of a cervical whole-section image detection method provided in some embodiments of this application;
[0015] Figure 2 is a flowchart of self-supervised pre-training provided in some embodiments of this application;
[0016] Figure 3 is a schematic block diagram of the principle of a cervical whole slide image detection method including preset cell detection and whole slide image classification provided by some embodiments of this application;
[0017] Figure 4 is a schematic diagram of the adaptive and evaluation process during prototype testing provided in some embodiments of this application;
[0018] Figure 5 is a schematic diagram of the cervical whole-section image detection process provided in some embodiments of this application;
[0019] Figure 6 is a visualization of the experimental verification results of the cervical whole-section image detection method provided in some embodiments of this application;
[0020] Figure 7 is a schematic block diagram of a cervical whole-section image detection device provided in some embodiments of this application;
[0021] Figure 8 is a schematic diagram of the hardware structure of an electronic device provided in some embodiments of this application. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the reference to "embodiment" herein means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments. The implementation methods described in the following exemplary embodiments do not represent all implementation methods consistent with the embodiments of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0023] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various concepts, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to a determination” as used herein may be interpreted as “when…” or “when…” or “in response to a determination.”
[0024] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.
[0025] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0026] To facilitate understanding of the inventive concept of this application, before providing a detailed description of the embodiments of this application, the English abbreviations (terms) / related concepts involved in the embodiments of this application will be explained first. The English abbreviations (terms) / related concepts involved in the embodiments of this application are subject to the following interpretation.
[0027] CCS: CCS stands for Cervical Cytology Screening, which refers to cytology screening. It uses computer algorithms to automatically analyze cervical cytology smears (such as Pap smears and liquid-based thin-layer cytology smears) to assist manual interpretation.
[0028] WSI: WSI stands for Whole Slide Image. It refers to a full-slide image, which is a high-resolution digital image that is converted from a complete cytology specimen (such as a cervical smear, sputum smear, etc.) through digital scanning technology. It contains all the cells and background information in the specimen.
[0029] Currently, cytological screening (CCS) is recommended as an early detection method for cervical disease detection applications. Cytologists examine and identify pre-selected cells using a microscope or digital slides. Reporting for CCS currently follows the Bethesda System (TBS) (Nayar and Wilbur, 2015), which provides widely accepted standardized interpretation guidelines. Although histological examination is considered the gold standard for diagnosing cervical diseases, CCS, as a non-invasive, effective, and cost-effective alternative, is more suitable for widespread cervical lesion screening. However, CCS still faces several challenges.
[0030] First, examining cytologists is both tedious and time-consuming. Each specimen typically contains 20,000 to 50,000 cells, with pre-defined cells sparsely distributed in the slide, while digital slides can reach sizes of 100,000 × 100,000 pixels.
[0031] Secondly, cell-level identification is inherently challenging due to the differences in cell morphology across different cell types (squamous cells and glandular cells), locations (surface, middle, parabasal, and basal layers), and tumor states (metaplastic cells, keratotic cells, and dyskeratotic cells).
[0032] Furthermore, patient-level screening results depend heavily on sample preparation and the experience of cytologists, which can lead to poor reliability both between and within observers.
[0033] With the widespread application of artificial intelligence-assisted image detection, the detection and analysis of whole-slice images (WSI) is helpful for cervical lesion screening and slide classification. Currently, WSI analysis is mainly used for histological detection through multi-instance learning (MIL), but its application to cytological screening still faces unique challenges.
[0034] Unlike histology, which deals with continuous regional tissues, cytology focuses on discrete cellular objects, requiring targeted methods for effective analysis. Therefore, applying WSI methods to cytology necessitates a detailed understanding of these differences to ensure the accuracy and validity of the results.
[0035] The first study applying WSI analysis followed this principle, developing a dual-path network for cell-level prediction, followed by a rule-based WSI classifier. Subsequent studies have incorporated more expertise to provide richer cytological features and guidance, thereby improving detection performance. For example, one approach designed a multi-resolution strategy, first learning low-resolution patch features and then optimizing high-resolution cell features. Other studies perform fine segmentation of the nucleus and cytoplasm to characterize cytological morphology and semantics. However, pixel-level annotation of the cytoplasm and nucleus imposes a heavy manual burden on system development; other designs employ cytological statistics or rule-based whole-slice classifiers, resulting in detection performance heavily reliant on detector performance, and weak classifiers limiting generalization ability. Furthermore, the combination and arrangement of models can increase system complexity. Some methods use multi-instance schemes for whole-slice classification, but this approach neglects cellular information in computational cytology, overemphasizing slice-level supervision. Therefore, these methods often fail to fully utilize cell-level and slice-level supervision.
[0036] Specifically, cell-level supervision requires precise annotation of the morphological and structural features of individual cells in a slice (such as marking the location and type of preset cells) to drive the model to learn fine-grained cellular features. Slice-level supervision, on the other hand, relies solely on the detection labels of the entire slice, focusing on guiding the model to learn the global judgment logic at the slice level. In practical applications, over-reliance on cell-level supervision can be difficult to adapt to the needs of large-scale detection due to the extremely high cost of fine-grained annotation (requiring professional physicians to spend a lot of time to complete). If only slice-level supervision is used, the model will struggle to capture the detailed features of key cells, resulting in insufficient recognition ability for complex samples (such as subtle variations in early cells or preset cells with significant morphological variations).
[0037] This imbalance in supervision limits the effectiveness of AI models in large-scale cytology whole-slice image detection. For example, in the analysis of slices with diverse cell morphologies and sparse lesion distribution, the model is prone to judgment bias. At the same time, it also makes it difficult to meet the accuracy requirements of actual detection tasks, hindering the application of this technology to a wider range of scenarios.
[0038] In view of this, this application proposes a method, device, and computer program product for detecting cervical cytology whole slide images. In the embodiments of this application, a cervical cytology whole slide image to be detected is acquired, and the cervical cytology whole slide image is cropped to obtain image blocks. Then, the image blocks are input into a cell detector to obtain preset cell images in the image blocks output by the cell detector. Then, the feature extraction network in the slide classifier extracts features from the preset cell images, and the extracted features are predicted by the multi-instance learning classifier in the slide classifier to obtain the detection result of the cervical cytology whole slide image. This scheme can detect the information of preset cells from a single slide by inputting the image blocks into the cell detector, and obtain a finer-grained cell representation. Then, the feature extraction network in the slide classifier extracts features from the preset cell images in the obtained fine-grained cell representation to obtain the preset cell features at the slide level. Finally, the preset cell features at the slide level are input into the multi-instance learning classifier for prediction to obtain the detection result of the cervical cytology whole slide image. This disclosure combines cell-level supervision of cell detectors (driving the model to learn fine-grained cell features) with slice-level supervision of slice classifiers (focusing on guiding the model to learn global judgment logic at the slice level), thereby making full use of cell-level and slice-level supervision, improving the accuracy of computer-aided cervical whole slice image detection, and promoting the application of this technology in a wider range of scenarios.
[0039] The cervical whole-section image detection method provided in this application relates to the field of image detection technology. It can be applied to the electronic device provided in this application. The electronic device can be a terminal or a server.
[0040] In some embodiments, the terminal may be a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smartwatch, or in-vehicle terminal, but is not limited thereto.
[0041] The server can be configured as a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network.
[0042] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0043] It should be noted that in all specific embodiments of this application, when processing data related to user characteristics, such as user biometric information and historical biometric data, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive user biometric information, separate user permission or consent is obtained through pop-ups or redirects to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.
[0044] The implementation steps of a cervical whole-section image detection method provided in this application will be described in detail below with reference to the accompanying drawings.
[0045] Please refer to Figure 1, which is a flowchart of a cervical whole-slice image detection method provided in some embodiments of this application. It should be noted that the steps shown in the flowchart can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0046] The method of the embodiments of this application includes the following steps:
[0047] Step 101: Obtain whole slide images of the cervical cytology specimen to be examined;
[0048] Step 102: Cropping the whole cervical cytology slide image to obtain image blocks;
[0049] Step 103: Input the image patch into the cell detector to obtain the preset cell image in the image patch output by the cell detector;
[0050] Step 104: Extract features from the preset cell image using the feature extraction network in the slice classifier, and predict the extracted features using the multi-instance learning classifier in the slice classifier to obtain the detection results of the whole slice image of cervical cytology.
[0051] Steps 101 to 104, as illustrated in the embodiments of this application, involve inputting image patches into a cell detector to detect information about preset cells from a single slice, obtaining a finer-grained cell representation. Then, the feature extraction network in the slice classifier extracts features from the preset cell image in the obtained fine-grained cell representation, obtaining slice-level preset cell features. Finally, the slice-level preset cell features are input into a multi-instance learning classifier for prediction, yielding the cervical cytology whole-slice image detection result. This disclosure combines cell-level supervision of the cell detector (driving the model to learn fine-grained cell features) with slice-level supervision of the slice classifier (focusing on guiding the model to learn global judgment logic at the slice level), thereby fully utilizing both cell-level and slice-level supervision, improving the accuracy of computer-aided cervical whole-slice image detection, and promoting the application of this technology in a wider range of scenarios.
[0052] It should be noted that the cervical whole-slice image detection method proposed in this application can be used for racial identification of specific populations (such as adolescents, middle-aged people or the elderly), identification of women in different physiological cycles, and assessment of the impact of the environment on the human body.
[0053] In studies targeting specific populations for racial identification, whole-section images of cervical cells contain subtle racially relevant features, such as morphological differences in cervical epithelial cells and the distribution characteristics of the extracellular matrix among different ethnic groups. This method can accurately extract these features and, by analyzing microscopic indicators such as cell size, nuclear morphology, and cytoplasmic composition, assist in constructing racial identification models, providing data support for research on physiological differences among different ethnic groups in anthropology. Simultaneously, in studies exploring the impact of the environment on the human body, the state of cervical cells can reflect the effects of environmental factors to some extent. For example, individuals exposed to polluted environments for extended periods may exhibit unique morphological changes in their cervical cells; the metabolic state of cervical cells may also be reflected in images from populations in different climatic regions. This method, through detailed analysis of whole-section images of cervical cells, captures these environment-related cellular feature changes, quantitatively assesses the impact of environmental factors on the human body, and provides a powerful analytical tool for environmental health research in the field of public health.
[0054] The specific implementation methods for each of the above steps are described below.
[0055] In step 101, a whole slide image of the cervical cytology specimen to be examined is obtained.
[0056] Optionally, a digital pathology scanner can be used to perform a full-field scan of cervical cytology smears (such as liquid-based thin-layer cytology smears, Pap smears, or centrifuged cervical cytology precipitates) to generate high-resolution digital images. It should be understood that the scanning process must follow standardized parameters (such as scan magnification and pixel precision) to ensure that the images clearly present cell morphological details (including cell nuclei, cytoplasm, and intercellular relationships) while covering the entire smear area to avoid missing potential pre-selected cells at the edges or corners.
[0057] The scanning method for cervical cytology smears can be optical microscopy or line scanning, and this application does not impose any restrictions on this.
[0058] The generated digital images can be stored in pathology-specific formats (such as SVS, NDPI, TIFF, etc.) on medical imaging servers, dedicated storage systems, or storage modules within testing systems. These digital images can be associated with basic user information (such as sample number, examination date, submitting institution) and smear preparation information (such as staining method, slide preparation technique), and data security and privacy can be ensured through encryption protocols. The storage system must support large-capacity data management and rapid retrieval functions to meet the efficient image access requirements in subsequent testing processes.
[0059] The generated digital images will enter the subsequent detection process, providing basic data support for tasks such as cell-level identification and lesion screening.
[0060] In step 102, the whole cervical cytology slide image is cropped to obtain image blocks.
[0061] For example, for a whole cervical cytology slide image, a non-overlapping sliding window of 1200×1200 pixels can be used to crop out an image patch from the whole cervical cytology slide image.
[0062] Optionally, to adapt to the model input size, reduce computational load, and focus on densely populated cell regions to improve analysis efficiency, the original image can be systematically cropped. This can be done through regular grid cropping, dividing the entire slice image into non-overlapping or partially overlapping image blocks according to a preset fixed size (e.g., 1200×1200 pixels or 512×512 pixels). Overlapping cropping (e.g., overlap rate of 20%-50%) can prevent information loss due to target cells being segmented across boundaries. Alternatively, adaptive cropping can be used, identifying cell-rich regions in the image using prior algorithms (e.g., cell density detection) (excluding invalid regions such as blank backgrounds and slide edges), and cropping only the effective regions to reduce redundant data.
[0063] Optionally, during the cropping process, the coordinate position information of each image block in the original whole slice image (such as the upper left pixel coordinates) can be recorded simultaneously, so that the block-level analysis results can be mapped back to the whole slice to realize the localization and integration of lesion areas. The cropped image blocks can also undergo basic preprocessing (such as normalizing pixel values and unifying the color space) to ensure that they meet the input requirements of subsequent models (such as cell detection and classification models), laying the foundation for accurate extraction of cell morphology features.
[0064] In step 103, the image block is input into the cell detector to obtain the preset cell image in the image block output by the cell detector.
[0065] For example, as shown in Figure 3, each image block can be input into a cell detector, and then the cell detector can output a preset cell image of that image block. It should be understood that the preset cell image can be a set type of cytology image, or it can be a normal cytology image for a certain population (such as a specific ethnic group, race, age, etc.), and this application does not limit it.
[0066] For example, an image patch can also be input into a deformable Transformer detector to obtain preset cell information in the image patch, wherein the preset cell information includes preset cell location, preset cell type and confidence level; then, a preset cell image is determined based on the preset cell information.
[0067] For a whole-section cytology image, multiple image patches can first be extracted from it using a 1200×1200 pixel non-overlapping sliding window. Then, a Deformable Detection Transformer (DDETR) is used to infer these image patches, identifying preset cell images as preset candidate objects and outputting prediction information including location, type, and confidence level. After the sliding window traverses the entire whole-section image, a "bag" can be constructed by aggregating selected cell images (e.g., the top-ranked preset confidence values, where the preset values can be flexibly set as needed). The selected cell images are cropped from the whole-section image based on the predicted preset cell location, preset cell type, and confidence level.
[0068] By inputting image patches into the cell detector, background information, normal cells, and other information irrelevant to the detection target can be filtered out, providing an accurate data foundation for subsequent slice-level detection.
[0069] Specifically, in the preset cell detection, the cytology image can first be enhanced, and then mapped into multi-scale feature maps through a backbone network. The size of the feature maps are 1 / 8, 1 / 16, 1 / 32, and 1 / 64 of the original image, respectively. These multi-scale feature maps serve as input to a Transformer structure, which contains N encoder layers and M decoder layers.
[0070] Each encoder layer consists of the following components: position encoding, residual connections, layer normalization, feedforward neural network, and the proposed multi-scale deformable attention module (MSDAttn). This module is specifically designed to capture deformable information and is a variant of the multi-head self-attention mechanism. Its implementation can be represented as follows:
[0071] Where x→y, L and K represent the number of attention heads, feature levels, and number of sampling points, respectively. q and For each query element q, x represents the content feature and the normalized coordinates of the reference point. l This is a multi-scale feature map. A mlqk and p mlqk These are attention weights and sampling offsets, respectively. W represents the scaling function. m and W′ m These are learnable weights.
[0072] The decoder includes the same components as the encoder, taking the target query as input to limit the maximum number of detections. The final prediction head is a feedforward network, serving as both a classification and regression head, used to predict probabilities and bounding box coordinates.
[0073] The aforementioned multi-scale deformable design is highly compatible with cytological screening. It is suitable for small target detection (such as naked nuclei), can adapt to the multi-scale requirements of multi-resolution cytological images, and can also cope with diverse morphological features.
[0074] Alternatively, the cell detector can also be implemented using Faster R-CNN, YOLO-v3, RetinaNet, or DETR detection models, and this application does not impose any restrictions on this.
[0075] In step 104, features are extracted from the preset cell image through the feature extraction network in the slice classifier, and the extracted features are predicted through the multi-instance learning classifier in the slice classifier to obtain the detection result of the whole slice image of cervical cytology.
[0076] For example, as shown in Figure 3, thousands of predefined cells detected from a single slice (each cell having a different degree of characteristic and confidence score) can be aggregated for slice classification. Regarding cell selection and aggregation, for each predefined category (such as: Atypical Squamous Cells of Undetermined Significance (ASC-US), Low-Grade Squamous Intraepithelial Lesion (LSIL), Atypical Squamous Cells - High-Grade Lesion Not Excluded (ASC-H), High-Grade Squamous Intraepithelial Lesion (HSIL), Squamous Cell Carcinoma (SCC), and Irregular Morphology of Atypical Glandular Cells (AGC) or Metaplastic Cells), cell images with the highest confidence scores (where the predefined scores can be flexibly selected according to actual conditions) can be selected to construct a whole-slice feature bag (WSI feature bag) to characterize the original whole-slice image. Subsequently, each cell image in this feature bag is enhanced and used as input to the classifier. It is understood that, to avoid destroying cell morphological features, the cell images maintain their original proportions and are not scaled.
[0077] In some implementations, feature extraction of a preset cell image by a feature extraction network in a slice classifier can be achieved by using a pre-trained feature extraction network in the slice classifier to extract features from the preset cell image and obtain a feature representation of the preset cell image; wherein, during the feature extraction process of the preset cell image, the weights of the feature extraction network are frozen.
[0078] For example, a pre-trained visual Transformer can be used to map cell images into 1024-dimensional features. All mapped instance features in the whole-slice feature package are concatenated into a single feature package, serving as a representation of the whole-slice image, i.e., the preset cell feature package. Subsequently, the preset cell feature package is used by a trainable classification head to obtain slice-level prediction results, i.e., the detection results of the cervical cytology whole-slice image. Optionally, the multi-instance learning classifier (classification head) can use ABMIL, MeanMIL, MaxMIL, CLAM, DSMIL, TransMIL, or S4MIL models; this application is not limited to these.
[0079] In some implementations, pre-training the feature extraction network may involve: inputting unmasked global cropped image patches into a teacher network to obtain a first cropped prototype; inputting masked global cropped image patches and locally cropped image patches into a student network to obtain a second cropped prototype; updating the parameters of the student network by aligning the second cropped prototype with the first cropped prototype based on reconstruction loss and alignment loss; and updating the parameters of the teacher network according to the exponential moving average of the student network parameters; wherein the global cropped image patches and locally cropped image patches are obtained by preprocessing cell image patches in a cytology dataset and performing global and local cropping.
[0080] Optionally, during the pre-training phase of the feature extraction network, the constructed large-scale multicenter cytology dataset CCS-100K can be fully utilized and mined using a self-supervised learning approach. This enables the effective capture and learning of inherent and general knowledge in cytology, including cell instance features (cytoplasm, nucleus, cancer cells), semantic features (morphology), and global information (distribution).
[0081] The CCS-100K dataset can be a comprehensive multicenter dataset covering multiple (e.g., 45) medical centers, containing a large number (e.g., 120,196) cervical whole-slice images. By employing a large-scale self-supervised pre-training method, the feature extraction network can learn robust features from diverse datasets, thereby improving accuracy in different scenarios.
[0082] In some implementations, the pre-training of the feature extraction network can employ the self-supervised learning framework DINOv2, which uses a student-teacher network structure to achieve knowledge distillation through the alignment of network output distributions. The DINOv2 pre-training process can be divided into several key steps, the core of which is to learn from the "teacher network" through the "student network" and ultimately enable the model to master the semantic and contextual information of the cervical slice image.
[0083] The key steps include data augmentation and cropping. For a cytological image (such as a full cervical slice), random augmentation (e.g., color adjustment, Gaussian blur, exposure adjustment) is first performed to generate multiple different versions of the image. These augmented images are then divided into "global cropped blocks" (larger regions, preserving the overall structure) and "local cropped blocks" (smaller regions, focusing on details). Some global cropped blocks are randomly "masked" (partially obscuring areas). Then, for the teacher network: the input is the unmasked global cropped block, and the output is the "feature prototype" (representing the core features of the image block). For the student network: the input is the masked global and local cropped blocks, and the output is the corresponding feature prototype. Reconstruction Loss allows the student network to predict the teacher network's output for the same region based on the pixels surrounding the masked area. The goal is to force the student network to learn "contextual relationships" (e.g., inferring the morphology of masked cells from the surrounding structure). Alignment loss aligns all feature prototypes (from the masked global and local blocks) output by the student network with the corresponding global block feature prototypes output by the teacher network (by minimizing cross-entropy loss). The goal is to ensure that the features learned by the student network are consistent with those of the teacher, guaranteeing feature robustness. The teacher network is not trained directly; instead, it is updated slowly using the exponential moving average (EMA) of the student network parameters. This makes the teacher's features more stable and avoids being affected by transient fluctuations in the student network's performance.
[0084] As shown in Figure 2, the alignment strategy in pre-training is mainly implemented through the reconstruction loss and alignment loss between the outputs of the student network and the teacher network. Specifically, for a cellular image patch, it is first randomly enhanced (including color dithering, Gaussian blur, or exposure adjustment, etc.) to generate multiple enhanced patches. Then, these enhanced patches are subjected to global cropping and multiple local cropping. The globally cropped patches are randomly masked: the unmasked global cropped patches are used as input to the teacher network, while the masked global cropped patches and local cropped patches are used as input to the student network to output the cropped prototype.
[0085] Next, the prototype alignment process updates the student network through reconstruction loss and alignment loss: Mask image modeling in the reconstruction loss uses the student network's output to predict the teacher network's output, prompting the student network to learn semantic and contextual information based on pixels surrounding the masked region; regarding the alignment loss, the aggregated cropped prototypes generated by the student network (including masked global and local cropped prototypes) are aligned with the unmasked global cropped prototypes corresponding to the teacher network by minimizing cross-entropy loss. Finally, the teacher network is updated using the exponential moving average of the student network.
[0086] This self-supervised pre-training enables the capture and accumulation of generalizable cytological knowledge from large-scale, diverse data, thereby achieving a universal understanding of cytology. After pre-training, the pre-trained model can be used as a feature extractor for downstream detection tasks. In this case, the extracted features can efficiently generalize and adapt to the entire process of cytological screening tasks (with supervised information), thus demonstrating robust and consistent performance in multi-center retrospective and prospective validation.
[0087] In AI-assisted cell detection, the good performance of deep learning models relies on an important assumption: that training and test data are independently and identically distributed from the same or similar distributions. However, for AI-assisted cell detection tasks, cross-center heterogeneity in different scenarios can become a critical problem, which may lead to a significant degrade in model performance (e.g., model bias towards specific categories).
[0088] Specifically, this heterogeneity stems from differences in sample preparation and imaging protocols among different institutions, leading to a shift in the distribution of whole-slice image (WSI) data, which may manifest as variations in resolution or image appearance. For example, differences between institutions in patient demographics, sample preparation methods, staining protocols, and slice imaging can cause inconsistencies between model training data and the test samples, resulting in severe heterogeneity. Consequently, the detection performance of this approach may significantly decline in practical applications, revealing its insufficient generalization ability and thus weakening the effectiveness of AI-driven solutions in CCS.
[0089] In view of this, before acquiring the whole cervical cytology slide image to be detected, this disclosure further adjusts the parameters of the slide classifier using knowledge distillation based on the cell sample data to be detected in the target detection scenario and the retrospective sample data with category labels, according to the test-time adaptive mechanism. This enables the cervical whole slide image detection method proposed in this disclosure to have stronger generalization ability and better overall performance under large-scale detection conditions.
[0090] Test-time adaptation refers to adapting a pre-trained task model to the scenario of the current test batch before making predictions. In this disclosure, this adaptive scheme compensates for the shortcomings of the pre-trained model by dynamically updating a small number of model weights. Specifically, it utilizes currently available test data and source supervision information to adjust specific model weights to enhance model adaptability and ultimately improve performance. The specific process is shown in Figure 4.
[0091] In some implementations, adjusting the parameters of the slice classifier using knowledge distillation based on test-time adaptive mechanisms, using test cell sample data and retrospective sample data with class labels in a target detection scenario, may include: determining class prototypes from the retrospective sample data using a pre-trained feature extraction network, wherein the class prototypes are used to characterize one or more representative features of each class of cervical cells; acquiring sample objects using test cell sample data in a target detection scenario; and performing knowledge distillation based on prototype alignment loss and consistency loss to update the parameters of the slice classifier by aligning the sample objects with the feature representations of the class prototypes.
[0092] For example, a pre-trained feature extraction network can be used to extract aggregated whole-slice image (WSI) features from retrospective data (i.e., samples seen during model training, such as the aforementioned CCS-100K dataset). The feature dimension can be n×1024 (where n is the number of instances). For each category (e.g., NILM, LSIL, HSIL, etc.), the top preset number of WSI features with the highest confidence are selected as the "prototype" (which can be understood as "typical representatives") of that category. These prototypes are aggregated to form the prototype library. The prototype library is equivalent to storing the "standard feature representations" of each category in historical data, serving as a "reference standard" for subsequent adaptation. This prototype library stores diagnostic supervision information from retrospective samples. It should be understood that the preset values can be flexibly selected according to the actual situation.
[0093] Subsequently, for WSIs from unseen centers (WSIs in object detection scenarios), pre-defined cells are inferred as candidate objects using a cell detector. Then, knowledge distillation is performed based on prototype alignment loss and consistency loss, updating the parameters of the slice classifier by aligning sample objects with the feature representations of class prototypes.
[0094] Specifically, "data augmentation" (such as slightly adjusting cell rotation angles or brightness) can be performed on multiple samples in the current test batch to generate augmented samples. The aim is to increase the diversity of the current batch of samples, allowing the model to learn more robust local features. Then, a pre-trained feature extractor (i.e., the feature extractor's parameters are frozen) is used to convert the original and augmented samples into original and augmented features (i.e., converting the image into numerical features that the computer can understand). Teacher and student networks are initialized; the teacher network receives unaugmented samples as input, and the student network receives both unaugmented and augmented samples as input. Through "knowledge distillation" (i.e., aligning the student network's output with the teacher network), the feature representations of the current test data and those in the prototype library become more consistent, and the student network's predicted classification results become more consistent with the teacher network's predicted classification results. Finally, through this alignment, the model updates the classifier's parameters (rather than the entire model), making the model's detection more suitable for the current test samples, thereby improving detection accuracy.
[0095] In Figure 4, S1 is used to generate source prototypes, i.e., to generate source prototypes from retrospective samples, resulting in a category-specific prototype library. S2 is used to obtain cell candidate objects from the test data. S3 is used to achieve test-time adaptation through the mean teacher framework, including prediction alignment and prototype alignment. The adaptive prediction result is calculated by integrating the outputs of the student network and the teacher network.
[0096] Specifically, in S1, for a whole-slice image from a source center (known center), a pre-defined cell feature package is extracted from the whole-slice image using a trained model's cell detector and feature extractor. This feature package is then stitched together to obtain slice features. These slice features are then input into a slice classifier to obtain detection and classification information, which is saved, thus constructing a category prototype library. The category prototype library stores a set of representative features from each pre-defined category of cervical cells.
[0097] In S2, the whole-slice image from the unseen center (unknown center, target detection center) is passed through a trained cell detector to identify preset cells. Cell images with the highest confidence values among the preset cells are aggregated (it should be understood that the preset values can be flexibly set according to the actual situation) to obtain cell-level results. Then, data augmentation is performed on the results to obtain preset cell candidates.
[0098] In S3, the unenhanced cell-level results are input into the teacher prediction network and the student prediction network, respectively. The enhanced cell-level results are input into the student prediction network, yielding the predicted classification results for the teacher network and the student network, respectively. Both the teacher and student prediction networks include a pre-trained feature extractor and a slice classifier with parameters to be tuned. Knowledge distillation is performed using prototype alignment loss and consistency loss, respectively. By aligning the sample object with the feature representation of the category prototype, the predicted classification results of the student network are aligned with those of the teacher network, updating the parameters of the slice classifier to achieve test-time adaptation. Prediction alignment is achieved using cross-entropy loss, and prototype alignment is achieved using contrastive loss.
[0099] In some embodiments, the slice classifier includes a multi-instance learning classifier that performs knowledge distillation based on prototype alignment loss and consistency loss. This is achieved by aligning the feature representations of sample objects with those of the class prototype, aligning the predicted classification results of the student network with those of the teacher network, and updating the parameters of the multi-instance learning classifier in the student network. The parameters of the multi-instance learning classifier in the teacher network are then updated using an exponential moving average based on the parameters of the multi-instance learning classifier in the student network. Specifically, the prototype alignment loss aligns the unenhanced and enhanced feature representations of sample objects with the class prototype through a contrastive learning mechanism. The consistency loss is determined based on the predicted classification results of the teacher network and the student network.
[0100] Specifically, in the adaptation phase, given m samples of the current batch x (i.e., the batches and samples in the preset cell candidate objects in S2), the i-th cell containing n cell instances... th Each sample will undergo data augmentation to improve the diversity of cell instances. The current batch sample x and its augmented samples... Transformed into feature f and
[0101] Subsequently, alignment knowledge distillation from the source prototype to the current test sample is performed using the mean teacher framework. Both the teacher and student networks are initialized using a cell-level screening model. Knowledge distillation is then performed using prototype alignment loss and consistency loss. The prototype alignment loss, through contrastive learning, combines the current batch features f with its enhanced features. Align with the prototype. Here, J + The positive sample pairs in the network consist of the current sample I and its corresponding augmented sample. The current sample I refers to the features obtained without augmentation, i.e., the features output by the feature extractor in the teacher network; the corresponding augmented sample refers to the features obtained after augmentation, i.e., the features output by the feature extractor in the student network.
[0102] To find the most similar sample from the category prototype library, nearest neighbor prototypes are calculated through contrastive learning. The nearest neighbor prototype is calculated using similarity: Sim(z i , z j )=z i ·z j / (||z i ||||z j ||), where z represents the projected embedding vector. The prototype alignment loss is expressed as follows:
[0103] Where τ is an adjustable temperature hyperparameter. J = J - +J + This represents the total number of positive and negative samples. Regarding consistency loss, for a given input x... i The student network and the teacher network each output the prediction results. and Consistency loss The following is used to align these predictions:
[0104] Where C represents the total number of categories.
[0105] The student network is updated by summing all the losses mentioned above. The teacher network is updated stably using an exponential moving average. The final adaptive prediction at test time consists of the integrated results of the teacher network output and the student network output.
[0106] Specifically, the predictions of the average teacher network and student network can be used as the final predictions to make the model predictions more stable during the adaptation process.
[0107] The above is an introduction to the implementation method of the cervical whole-section image detection method proposed in this disclosure.
[0108] The solutions of this disclosure embodiment will be further introduced and explained below with reference to specific application examples.
[0109] Please refer to Figure 5, which is a schematic diagram of the cervical whole-slice image detection process provided in some embodiments of this application. In Figure 5, the cervical whole-slice image detection system is used to detect the WSI to be detected. It should be understood that before detection, the WSI is stored in the storage module of the electronic device, and during detection, the WSI is input from the storage module into the detection system for detection.
[0110] Specifically, the WSI to be detected is cropped into multiple image patches. Each image patch is input into a cell detector, which identifies preset cells and outputs information such as the preset cells, confidence level, and location within that image patch. The cropped patches are then reserved for use by the classifier. The top preset value cell images based on their confidence levels are aggregated (it should be understood that the preset values can be flexibly set according to actual conditions) and input into a pre-trained ViT network with fixed weights and a multi-instance learning classifier for slice-level result detection and classification.
[0111] The aforementioned cervical whole-section image detection system is used to implement the cervical whole-section image detection method proposed in this disclosure. The cervical whole-section image detection system mainly includes three consecutive stages:
[0112] 1) Self-supervised basic model pre-training stage;
[0113] 2) Downstream task training phase;
[0114] 3) Adaptive testing phase.
[0115] All stages are closely aligned with the practical needs of actual testing scenarios.
[0116] The pre-training phase is used to perform self-supervised pre-training on large-scale multicenter cytology image patches to build a generalizable cytology feature extraction model; the training phase aims to train a whole-slice image classification model for cytology screening, which includes two components: a cell detector for pre-defined cell identification and a slice classifier for slice-level detection classification; the adaptation phase aims to adapt the model to the current samples during system deployment and evaluation, especially suitable for large-scale detection across multiple medical centers.
[0117] Specifically, in the pre-training stage of the self-supervised basic model, the large-scale multicenter cell biology dataset CCS-100K can be fully utilized and mined through self-supervised learning. This enables the effective capture and learning of inherent and general knowledge in cell biology, including cell instance features (cytoplasm, nucleus, cancer cells), semantic features (morphology), and global information (distribution).
[0118] During the adaptive testing phase, the pre-trained model is further refined and personalized to acquire the expertise required for specific downstream tasks, such as identifying predefined cells and classifying whole-slice images. In this phase, the cell detector and slice classifier can be trained using labeled data. The cell detector aims to identify predefined cells throughout the slice, and then the slice classifier aggregates the cell information to detect the final slice result.
[0119] During the test-time adaptive phase, the system can adapt to specific target scenarios through the proposed prototype-guided adaptive technique, which aligns the test-time prediction results with the prototypes from the source samples. This technique enables the cervical whole-slice image detection system to possess stronger generalization ability and better overall performance under large-scale detection conditions. The specific pre-training method and test-time adaptive method can be found in the description of the implementation method for cervical whole-slice image detection described above, and will not be repeated here.
[0120] In addition, the applicant has verified the effectiveness of the proposed cervical whole-slice image detection system and method. It should be understood that the experimental verification presented below is merely one verification method proposed by the applicant and should not limit the scope of protection of this application in any way. Furthermore, other verification methods can also be used to verify the effectiveness of the embodiments proposed in this application, and the applicant makes no limitation in this regard.
[0121] Regarding the dataset, the applicant collected the large-scale multicenter dataset CCS-100K for pre-training the self-supervised base model, training the screening model, and adaptation during testing. This dataset contains 124,118 retrospective cytology specimens from 45 medical institutions, all scanned into digital slides. After quality control, 3,922 samples were excluded, ultimately resulting in 120,196 whole-slice cytology images used for retrospective studies.
[0122] For validation, the applicant used Python (v3.9.0) and PyTorch (v2.0.0) for model training and evaluation. During pre-training, DINOv2 was used for self-supervised pre-training, with 16 80GB NVIDIA H800 GPUs. To adapt to the cell instance size, the local cropping ratio was reduced to 0.02, and the number of local crops was set to 8. In the segmentation and preprocessing of the whole-slice images, the cytology WSI was loaded using openslide-python (v1.3.1) and some known methods, and then the CLAM toolkit was used to segment the WSI into 1200×1200 image patches. The training phase experiments used 8 24GB NVIDIA GeForce RTX 3090 GPUs. For the cell detector, the MMDetection library (v3.0.0) was used to implement and compare various models. Distributed data parallelism technology was employed to accelerate cell prediction inference in the WSI.
[0123] Please refer to Figure 6, which is a visualization of whole-slice detection with interpretable cell-level and slice-level results obtained by the method and system of the embodiments proposed in this disclosure. The sample includes three typical detection results: normal (NILM), squamous epithelial abnormality (HSIL), and glandular epithelial abnormality (AGC). The cell-level visualization displays the detected preset cells and their confidence scores, providing an interpretive basis for the final slice-level prediction. The predictive information provided by the slice-level visualization includes: TBS diagnostic grade, abnormality status, and fine-grained score, along with a global whole-slice image (WSI) annotated with preset cells (represented by different labels).
[0124] In whole slice image (WSI) testing, for a cervical cytology specimen, the methods and systems of the embodiments of this disclosure can predict a positive score to report abnormalities in the test results, and can also provide fine-grained predicted categories and scores for assessing the degree of malformation, providing guidance for subsequent further examinations and biopsies.
[0125] For example, in Figure 6, samples with no intraepithelial lesion or malignancy (NILM) received a confidence score of 0.9999 for the normal category; while samples with high-grade squamous intraepithelial lesion (HSIL) had a positive score of 0.9999, with high-grade HSIL having a confidence score of 0.9915, indicating a high degree of cervical intraepithelial neoplasia and malignant potential. In cellular screening, pre-defined cells detected in each WSI were highlighted with different markers, the size of which represented the cell's confidence level. Visualization results showed that, due to uniform mixing and centrifugation during specimen preparation, the pre-defined cells were randomly distributed in each WSI. Simultaneously, the detected pre-defined cells, their confidence statistics, and characteristic morphology provided reliable clues for the final diagnosis.
[0126] In the second sample in Figure 6, a small number of low-grade squamous intraepithelial lesion (LSIL) cells (i.e., koilocytes with typical perinuclear halos) and a large number of HSIL cells (characterized by significantly enlarged nuclei, reduced cytoplasmic regions, and densely stained aggregates) were detected. These findings perfectly matched the true label of HSIL for this sample. Similarly, a large number of Atypical Glandular Cells (AGC) cells arranged in a palisade pattern were observed in the atypical glandular cell (AGC) sample, along with other atypical cells. For the NILM sample, although the slice-level prediction score was close to 1.00, some cells exhibited deep nuclear staining and enlargement, representing "indistinguishable mimic cells" or potentially malignant-related changes. Therefore, by integrating this information with an interactive interface, the system of the embodiments of this disclosure can also serve as an integrated system. The methods and systems of the embodiments of this disclosure will provide cytopathologists with interpretable AI-assisted results, output reliable detection conclusions, and guide subsequent final diagnoses, thereby optimizing the overall diagnostic process.
[0127] There is a serious problem of multicenter heterogeneity in cytology samples collected under different scenario conditions. To evaluate the performance and effectiveness of the methods and systems of the embodiments proposed in this disclosure in this regard, the applicant included 13,907 cytology specimens from 7 external centers (EC1-EC7), which did not overlap with the 38 centers in the training cohort.
[0128] Three sets of experiments were compared: a complete whole-slice image (WSI) detection system, a whole-slice image (WSI) detection system without adaptation (GSPA w / o A), and a whole-slice image (WSI) detection system without pre-training and adaptation phases (GSPA w / o PA). The results show that the whole-slice image (WSI) detection system achieved satisfactory and stable whole-slice image (WSI) classification performance at all seven external centers, with AUC values ranging from 0.844 (95% confidence interval 0.828–0.861) to 0.953 (95% confidence interval 0.943–0.963), and sensitivity ranging from 0.687 (95% confidence interval 0.657–0.717) to 0.937 (95% confidence interval 0.925–0.948). Notably, the AUC values exceeded 0.90 at four external centers, demonstrating the strong generalization ability of the whole-slice image (WSI) detection system and indicating its ability to be applied in various scenarios.
[0129] To further validate the effectiveness of the proposed system, the pre-training and adaptation phases were progressively removed, and performance changes were reported. After removing the adaptive part during testing, the evaluation results for all seven centers showed a consistent decrease in AUC values; for example, the AUC for EC2 decreased from 0.953 to 0.876, and for EC6 from 0.844 to 0.739. The decrease in sensitivity for the other centers, except for EC2 and EC6, further highlights the effectiveness of the prototype alignment strategy of the cervical whole-slice image detection system in addressing cross-center bias. This adaptive strategy allows the model to better adapt to the target test samples by updating a small number of parameters.
[0130] Furthermore, to verify the performance of the proposed system, this disclosure also trained an adaptive detection-based cervical whole-slice image detection model without a pre-trained extractor and evaluated it directly at external centers. The results showed a significant performance degradation, with an AUC value close to 0.50; worse still, the sensitivity dropped to close to 0.0 in four centers, indicating that the model is biased towards specific NILM categories and cannot identify positive samples. This demonstrates that the extracted features, lacking generalization ability, can only characterize data from the current center and are ineffective for samples from unseen cytological centers. Therefore, the cytological data generalization representation achieved through self-supervised pre-training brings significant performance improvements across all centers.
[0131] It should be understood that the above experimental results and data are only applicable to the current test samples. When the samples change, the data may be inconsistent.
[0132] The implementation of the cervical whole-section image detection device provided in this application will now be described in detail with reference to the accompanying drawings.
[0133] Regarding the cervical whole-slice image detection method provided in the above embodiments, this application embodiment also provides a cervical whole-slice image detection device for implementing the above method, as shown in FIG7. FIG7 is a block diagram of the cervical whole-slice image detection device according to an embodiment of this application. The cervical whole-slice image detection device 700 includes:
[0134] The acquisition module 701 is used to acquire whole-section images of cervical cytology to be tested;
[0135] The cropping module 702 is used to crop the whole cervical cytology slice image to obtain image blocks;
[0136] Cell detection module 703 is used to input the image block into a cell detector to obtain a preset cell image in the image block output by the cell detector;
[0137] The feature extraction and prediction module 704 is used to extract features from the preset cell image through the feature extraction network in the slice classifier, and to predict the extracted features through the multi-instance learning classifier in the slice classifier, so as to obtain the detection result of the whole slice image of cervical cytology.
[0138] It is understood that the content of the above method embodiments is applicable to the present device embodiments. The specific functions implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0139] Referring to Figure 8, this application embodiment also provides an electronic device 800, which includes a memory 801, one or more processors 802 (only one is shown in Figure 8), and a computer program stored in the memory 801 and executable on the processor 802. The memory 801 stores software programs and units, and the processor 802 executes various functional applications and data processing by running the software programs and units stored in the memory 801 to obtain resources corresponding to the aforementioned preset events. Optionally, the processor 802 implements the aforementioned cervical whole-slice image detection method by running the computer program stored in the memory 801.
[0140] Memory 801 serves as a non-transitory computer-readable medium for storing non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory 801 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory 801 may optionally include memory remotely located relative to the processor 802, and this remote memory may be connected to the processor 802 via a network.
[0141] It is understood that the content of the above method embodiments is applicable to the embodiments of this electronic device. The specific functions implemented by the embodiments of this electronic device are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0142] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described cervical whole-slice image detection method.
[0143] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0144] This application also provides a computer program product, which includes a computer program that, when executed by one or more processors, can implement the steps of the cervical whole slice image detection method described above.
[0145] It is understood that the content of the above method embodiments is applicable to this computer program product. The specific functions implemented by the embodiments of this computer program product are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0146] The cervical cytology whole slide image detection method, apparatus, electronic device, medium, and computer program product provided in this application acquire a whole slide image of cervical cytology to be detected, detects it with a cell detector to obtain a preset cell image, then extracts features from the preset cell image through a feature extraction network in the slide classifier, and predicts the extracted features through a multi-instance learning classifier in the slide classifier to obtain the cervical cytology whole slide image detection result. It combines cell-level supervision (driving the model to learn fine-grained cell features) of the cell detector with slice-level supervision of the slide classifier (focusing on guiding the model to learn global judgment logic at the slice level), thus fully utilizing cell-level and slice-level supervision and improving the accuracy of computer-aided cervical cytology whole slide image detection. Furthermore, the feature extraction network is self-supervised pre-trained on large-scale multi-center cytology image blocks, effectively capturing and learning inherent and universal knowledge in cytology. To adapt to heterogeneity issues in different scenarios, a test-time adaptive technique is used to adapt to the detection scenario. This technique aligns the prediction results during testing with the prototype from the source sample, thereby improving the generalization ability and overall performance of the method and system under large-scale screening conditions.
[0147] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0148] Although specific embodiments are described herein, those skilled in the art will recognize that many other modifications or alternative embodiments are also within the scope of this disclosure. For example, any of the functions and / or processing capabilities described in connection with a particular device or component can be performed by any other device or component. Furthermore, while various exemplary embodiments and architectures have been described according to embodiments of this disclosure, those skilled in the art will recognize that many other modifications to the exemplary embodiments and architectures described herein are also within the scope of this disclosure.
[0149] The foregoing description, with reference to block diagrams and flowcharts of systems, methods, systems, and / or computer program products according to exemplary embodiments, has described certain aspects of this disclosure. It should be understood that one or more blocks in the block diagrams and flowcharts, as well as combinations of blocks in the block diagrams and flowcharts, can be implemented by executing computer-executable program instructions, respectively. Similarly, according to some embodiments, some blocks in the block diagrams and flowcharts may not need to be executed in the order shown, or may not all need to be executed. Furthermore, additional components and / or operations beyond those shown in the blocks in the block diagrams and flowcharts may exist in some embodiments.
[0150] The embodiments of this application have been described in detail above with reference to the accompanying drawings. However, this application is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of this application.
Claims
1. A method for detecting whole-section images of the cervix, characterized in that, Includes the following steps: Obtain whole-section images of the cervical cytology specimens to be examined; The whole cervical cytology slide image is cropped to obtain image blocks; The image block is input into the cell detector to obtain a preset cell image in the image block output by the cell detector; The preset cell image is used to extract features by the feature extraction network in the slice classifier, and the extracted features are used to predict the results of the whole slice image detection of cervical cytology by the multi-instance learning classifier in the slice classifier.
2. The cervical whole-section image detection method according to claim 1, characterized in that, The step of extracting features from the preset cell image using the feature extraction network in the slice classifier includes: The preset cell image is used to extract features through a pre-trained feature extraction network in the slice classifier to obtain the feature representation of the preset cell image. During the feature extraction process of the preset cell image, the weights of the feature extraction network are frozen.
3. The method for detecting whole cervical slice images according to claim 1, characterized in that, The pre-training of the feature extraction network adopts a student-teacher network structure, and the pre-training steps of the feature extraction network include: Input the unmasked global cropped image patch into the teacher network to obtain the first cropping prototype; The masked global cropped image patch and the local cropped image patch are input into the student network to obtain the second cropping prototype; Based on reconstruction loss and alignment loss, the parameters of the student network are updated by aligning the second pruning prototype with the first pruning prototype. The teacher network parameters are updated based on the exponential moving average of the student network parameters; The globally cropped image patch and the locally cropped image patch are obtained by preprocessing cell image patches in the cytology dataset and then performing global and local cropping.
4. The method for detecting whole cervical slice images according to claim 1, characterized in that, Before acquiring the whole cervical cytology slide image to be tested, the method further includes: Based on the test-time adaptive mechanism, the parameters of the slice classifier are adjusted using knowledge distillation through the cell sample data to be tested in the target detection scenario and the retrospective sample data with category labels.
5. The cervical whole-section image detection method according to claim 4, characterized in that, The step of adjusting the parameters of the slice classifier using knowledge distillation based on the test-time adaptive mechanism, through the cell sample data to be tested and the retrospective sample data with category labels in the target detection scenario, includes: A category prototype is determined from the retrospective sample data using a pre-trained feature extraction network, wherein the category prototype is used to characterize one or more representative feature features in each class of cervical cells. The sample object is obtained by analyzing the cell sample data in the target detection scenario. Knowledge distillation is performed based on prototype alignment loss and consistency loss, by aligning the feature representations of the sample objects with the class prototypes to update the parameters of the slice classifier.
6. The method for detecting whole cervical slice images according to claim 5, characterized in that, The slice classifier includes a multi-instance learning classifier, and the step of adjusting the parameters of the slice classifier using knowledge distillation includes: Knowledge distillation is performed based on prototype alignment loss and consistency loss. The predicted classification results of the student network are aligned with the predicted classification results of the teacher network by aligning the feature representation of the sample object with the category prototype, thereby updating the parameters of the multi-instance learning classifier in the student network. The parameters of the multi-instance learning classifier in the teacher network are updated using an exponential moving average based on the parameters of the multi-instance learning classifier in the student network.
7. The method for detecting whole cervical slice images according to claim 5, characterized in that, The step of determining the category prototype from the retrospective sample data using a pre-trained feature extraction network includes: Cervical cytology whole slide image features are extracted from an aggregated preset cell sample package using a pre-trained feature extraction network, wherein the aggregated preset cell sample package is generated based on retrospective sample data. The cervical cytology whole slide image features are classified and predicted by a multi-instance learning classifier, and the top k high-confidence cervical cytology whole slide image features of each predicted category are determined as the category prototype, wherein the category prototype includes the corresponding diagnostic supervision information.
8. The method for detecting whole cervical slice images according to claim 5, characterized in that, The knowledge distillation based on prototype alignment loss and consistency loss, which updates the parameters of the slice classifier by aligning the feature representations of the sample objects with the class prototypes, includes: Inputting unenhanced sample objects into the teacher network yields the teacher network's predicted classification results; The augmented and unaugmented sample objects are input into the student network to obtain the predicted classification results of the student network. The adaptive prediction classification result during testing is determined based on the predicted classification results of the teacher network and the student network.
9. The method for detecting cervical whole-section images according to claim 5, characterized in that, The prototype alignment loss aligns the unenhanced and enhanced sample object feature representations with the category prototype through a contrastive learning mechanism.
10. The method for detecting whole cervical slice images according to claim 8, characterized in that, The consistency loss is determined based on the predicted classification results of the teacher network and the predicted classification results of the student network.
11. The method for detecting cervical whole-section images according to any one of claims 1-10, characterized in that, The step of inputting the image patch into a cell detector to obtain a preset cell image in the image patch output by the cell detector includes: The image patch is input into a deformable Transformer to obtain preset cell information in the image patch. The preset cell information includes preset cell location, preset cell type, and confidence level. The preset cell image is determined based on the preset cell information.
12. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the cervical whole-section image detection method as described in any one of claims 1 to 11.
13. A computer program product, characterized in that, When the computer program product is run in an electronic device, the electronic device performs the cervical whole slice image detection method as described in any one of claims 1 to 11.