Method and device for directly performing HPV prediction by utilizing cervical cell pathological image, electronic equipment, storage medium and program

The method addresses ASCUS diagnosis inaccuracies by using cervical cell pathology images for direct HPV prediction, enhancing diagnostic accuracy and reducing healthcare costs through a self-supervised feature extraction and multi-instance learning approach.

CN120318171APending Publication Date: 2025-07-15WUHAN LANTINGYUN MEDICAL LAB CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510383641.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

In the prior art, the accuracy of ASCUS diagnosis in cervical cancer screening is low, resulting in misdiagnosis and missed diagnosis. The HPV detection is high and invasive, increasing the economic and physical burden of patients. How to accurately judge the HPV infection status of ASCUS patients at low cost and accurately becomes a problem.

Method used

By using cervical cell pathological images, the abnormal cell regions were screened using the YOLOv7 target detection model, combined with the DINO framework and the DSMIL network, self-supervised learning and multi-instance learning, extracting and fusing image features to predict HPV infection status.

Benefits of technology

Accurate prediction of HPV infection status in ASCUS patients is achieved, reducing medical costs, reducing unnecessary HPV detection, and improving diagnostic efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318171A_ABST
    Figure CN120318171A_ABST
Patent Text Reader

Abstract

The invention provides a method and device for directly carrying out HPV prediction by utilizing a cervical cell pathology image, electronic equipment, a storage medium and a program, and relates to the technical field of medical image processing, and the method comprises the following steps: S10, data acquisition: obtaining a same-period cervical cell pathology full-slide image of a detected person which has been subjected to cervical tissue pathology analysis and a conclusion of which is evaluated as ASCUS; and S20, lesion area positioning: determining an abnormal cell area in the cervical cell pathological full-slide image by using the trained target detection model. Compared with the prior art, the method has the following beneficial effects: firstly, suspicious cells are screened out through a target detection method, and an existing method in a tissue pathology all-slide image is migrated into a cell pathology image; secondly, due to application of unsupervised learning and an attention mechanism, the model can automatically learn and emphasize the most important features for HPV detection in the image; and thirdly, the HPV infection state of the ASCUS patient is directly and accurately predicted from the cervical image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and in particular to a method, device, electronic device, storage medium and program for directly predicting HPV using cervical cell pathological images. Background Art

[0002] Cervical cancer is one of the types with relatively high incidence and mortality rates among female cancers globally, and is mainly associated with high-risk human papillomavirus (HPV) infection. HPV infection, especially HPV16 and HPV18, is an important pathogenic factor for cervical cancer. Early diagnosis of cervical cancer is crucial for improving the cure rate and survival rate. Therefore, regular cervical cancer screening is of great significance for preventing and early detecting cervical cancer. Currently, the screening methods for cervical cancer mainly include cervical cytology examination (TCT) and HPV detection. TCT detects the morphological changes of cervical cells to identify early signs of cervical cancer or precancerous lesions. However, TCT has relatively low accuracy in diagnosing samples with "atypical squamous cells of undetermined significance" (ASCUS), which is prone to misdiagnosis and missed diagnosis.

[0003] In the screening of cervical cancer, the diagnosis of ASCUS is an important and challenging problem. According to the latest TBS reporting guidelines, squamous cells are divided into NIL, ASCUS, LSIL, ASC-H, HSIL, and SCC. Among these types of cells, NIL is diagnosed as normal cells and can be observed continuously, while the others are abnormal cells. Clinically, for subjects with cytological screening results clearly indicating precancerous lesions or canceration such as LSIL, ASC-H, HSIL, and SCC, further diagnosis is required through colposcopy and histopathology. ASCUS belongs to abnormal squamous epithelial cells with unclear significance, and it is difficult to determine clinically whether it is a precancerous lesion. To further clarify whether there may be a precancerous lesion, HPV detection is often required. If it is positive, it indicates HPV virus infection and the risk of precancerous lesions, and the subsequent treatment methods are the same as those for LSIL and other results. Further colposcopy and histopathology examinations are recommended. However, the diagnosis result of ASCUS may be due to mild cell changes caused by HPV infection or other non-viral factors. That is, there are many HPV-negative patients (about 40%) among ASCUS results. If ASCUS is simply recommended for colposcopy or histopathological examination as a group of suspected precancerous lesions, it will undoubtedly cause overexamination of these HPV-negative patients. The price of HPV detection is high, which not only increases the economic burden on patients but also brings discomfort or pain to patients due to the invasiveness of histopathological examination, and there are also risks of complications. Therefore, how to accurately judge the HPV infection status of ASCUS patients at low cost has become an urgent problem to be solved.

[0004] Most hospitals currently use manual diagnosis methods for screening, but there are certain drawbacks in the existing process. First of all, histopathological diagnosis completely relies on the professional knowledge and experience of pathologists. The evaluation criteria of different doctors may vary, and there may sometimes be differences in the diagnosis results of the same patient. In addition, training an experienced pathology expert requires long-term training and high costs. Even top experts may experience a decline in diagnostic accuracy due to fatigue. Compared with the limitations of manual diagnosis, the advantages of digital pathology are gradually emerging. Through scanning and stitching technology, digital pathology laboratories can integrate cytological or histological images under the microscope into digital images, and then develop computer-aided diagnosis methods to achieve efficient, accurate, and objective pathological image analysis. At present, there is still a lack of relevant research on directly predicting HPV from cervical cell images. Summary of the Invention

[0005] Aiming at the deficiencies of the above-mentioned existing technologies, the technical problem to be solved by the present invention is to provide a method, device, electronic device, storage medium and program for directly predicting HPV using cervical cell pathological images, which can reduce medical costs by accurately predicting the HPV infection status of ASCUS patients.

[0006] To solve the above technical problems, the technical solution adopted by the present invention is: The present invention provides a method for directly predicting HPV using cervical cell pathological images, including the following steps:

[0007] S10. Data acquisition: Obtain the synchronous cervical cell pathological whole-slide images of the examinees who have undergone cervical tissue pathological analysis and whose conclusion evaluation is ASCUS;

[0008] S20. Lesion area localization: Use the trained object detection model to determine the abnormal cell area in the cervical cell pathological whole-slide image, and screen the suspicious positive cells in the abnormal cell area;

[0009] S30. Image feature extraction and fusion: Perform image feature extraction on the screened suspicious positive cells based on the self-supervised learning architecture, perform lesion prediction through the convolutional neural network, and fuse the image features through the self-attention mechanism to obtain cell-level features;

[0010] S40. HPV infection status prediction: Use the multi-instance learning network to aggregate the cell-level features to generate sample-level features, and input the classifier to determine the HPV infection status result of the sample-level features.

[0011] In a preferred solution, in step S10, a pathological slide scanner with an optical magnification of 20 to 40 times is used to scan the input cervical cell whole-slide image;

[0012] The output is a digital whole-slide image I with an image size of n×n pixels, a unified color space of RGB mode, and a brightness range of [0.8, 1.2];

[0013] Among them, the whole-slide image of cervical cells is a Papanicolaou-stained whole-slide image. The Papanicolaou-stained whole-slide image is processed by color standardization to unify the brightness and color distribution of the Papanicolaou-stained whole-slide images collected by different devices. The standardization process includes histogram equalization and color space conversion.

[0014] In the preferred solution, in step 20, the specific steps are as follows:

[0015] S21. Predict the bounding box: The object detection model is YOLOv7. The digital whole-slide image I is input into YOLOv7;

[0016] Based on the YOLOv7 object detection model, perform lightweight CNN operations on the digital whole-slide image I, and output a set of predicted bounding boxes B = {b1, b2,..., b n};

[0017] Among them, each predicted bounding box b i contains the coordinate information of x, y, w, h, the confidence E(b i ) and the class information G(b i ). i represents an index value, and b i represents the i-th predicted bounding box in B = {b1, b2,..., b n};

[0018] S22. Initial screening threshold: Set the confidence threshold T 阈值 = 0.5. According to E(b i ) ≥ T 阈值 initial screening of the predicted bounding box b i to obtain a set of candidate bounding boxes

[0019] Based on the set of candidate bounding boxes B 候选 mark the abnormal cell regions;

[0020] S23. Morphological filtering: Calculate the specific morphological features of the abnormal cell regions framed by the candidate bounding boxes ;

[0021] Compare the specific morphological features to determine the final set of suspicious positive cell bounding boxes

[0022] Among them, the specific morphological features include the area A and the circularity R;

[0023] S24. Output the result: According to the final set of suspicious positive cell bounding boxes B 最终Crop the corresponding cell images in the digitized whole-slide image I to obtain a set of suspicious positive cell images C = {c1, c2, …, c i}.

[0024] In a preferred solution, in step 30, the self-supervised learning architecture adopts the DINO framework, uses the pre-trained ResNet50 of DINO as the CNN backbone network, and the DINO framework performs contrastive learning pre-training on the unlabeled digitized whole-slide image I based on the ResNet50 backbone;

[0025] Input the suspicious positive cell image c i Extract multi-level feature maps, and the DINO framework performs global modeling on the multi-level feature maps extracted by ResNet50 through the Transformer encoder, and uses the self-attention mechanism to enhance the spatial correlation of the features;

[0026] Among them, the resolutions of the multi-level feature maps include downsamplings of 1 / 8, 1 / 16, and 1 / 32.

[0027] In a preferred solution, the specific steps in step 30 are as follows:

[0028] S31. Self-supervised feature extraction: Input the suspicious positive cell image c i Input into the DINO framework, perform data augmentation operations on the input suspicious positive cell image c i to generate a multi-view set {I1, I2, …, I n};

[0029] Extract the feature z i of each view I i based on the DINO framework, and obtain the corresponding feature vectors according to the feature z i : {z1, z2, …, z n};

[0030] Among them, i represents an index value, and I i represents the i-th view in {I1, I2, …, I n}, and z i represents the i-th feature vector in {z1, z2, …, z n};

[0031] The DINO framework minimizes the distance between the feature vectors {z1, z2, …, z n} in multiple views through the contrastive learning mechanism;

[0032] Obtain the final image feature F 自监督 extracted by DINO through multiple rounds of iterative training;

[0033] S32. Classification of positive and negative cells: Input the suspicious positive cell image c screened in step S24 i into the hierarchical multi-scale feature fusion network to obtain the positive and negative cell categories;

[0034] The multi-scale feature fusion network performs feature extraction on the input suspicious positive cell image c i to obtain a global feature vector and a local feature vector, fuse the global feature vector and the local feature vector, and output a set of generated feature maps {D1, D2,..., D n};

[0035] S33. Prediction of cell positive probability: Based on the set of feature maps {D1, D2,..., D n}, construct a channel attention module and a spatial attention module;

[0036] For each feature map D i , perform global max pooling and global average pooling respectively to obtain a channel feature vector g i ;

[0037] According to the channel feature vector g i , generate a channel attention weight vector w through a fully connected layer i ;

[0038] Multiply the channel attention weight vector w i with the feature map D i to obtain the feature map enhanced by channel attention Concatenate all the feature maps in the channel dimension and generate a spatial attention map M through a convolutional layer;

[0039] Multiply the spatial attention map M element-wise with the concatenated feature maps to obtain the feature map enhanced by spatial attention

[0040] Concatenate the self-supervised feature F 自监督 with the feature vector corresponding to the feature map obtained by the multi-scale feature fusion network to output a new feature vector F 融合 ;

[0041] Input the feature vector F into a fully connected layer 融合 , and calculate the probability P of positive cells through the Sigmoid function;

[0042] According to the probability P of positive cells and the positive cell feature F 自监督 , determine the cell level feature F 细胞 .

[0043] In a preferred solution, in step S40, the multi-instance learning network is a DSMIL network, which includes two parallel feature aggregation channels. The first stream is the key instance extraction stream, and the second stream is the instance attention aggregation stream;

[0044] According to the cell-level feature F in step S33 细胞 Obtain the cell instance feature h i (i = 1, 2, …, N);

[0045] The key instance extraction stream scores the cell instance feature h through the instance scoring network i For scoring;

[0046] Use the max pooling mechanism to select the cell instance feature h with the highest score i , and output to obtain the key instance h m ;

[0047] The selection formula is as follows:

[0048] c m (B) = max{W0h i}), i = 1, 2, …, N;

[0049] Among them, W0 is the learnable weight parameter in the instance scoring network, and h i Is the feature representation of the i-th cell instance, and N is the total number of instances;

[0050] Map each cell instance feature h i To the query vector q through linear transformation i And the information vector v i ;

[0051] Calculate the attention weight U(h i And the key instance h m Between), the formula is as follows: i ,h m ), the formula is as follows:

[0052]

[0053] Among them, Indicates the feature similarity between the i-th instance and the key instance;

[0054] According to the attention weight U(h i ,h m ) weighted aggregate the information vector of the cell instance feature h i , and output to obtain the sample feature b;

[0055] Fuse the key instance feature h m With the sample feature b, and output to obtain the final aggregated sample-level feature H样本 ;

[0056] The feature fusion formula is as follows:

[0057]

[0058] Among them, W0 and W b are trainable parameter matrices for after fusion;

[0059] Obtain cell-level statistics, and splice the cell-level statistics with the aggregated sample-level feature H 样本 to obtain an enhanced sample-level feature

[0060] Input the enhanced sample-level feature into the classifier to determine the HPV positive probability y 预测 .

[0061] In a preferred embodiment, the present invention provides a system for directly predicting HPV using cervical cell pathological images, which is characterized by including:

[0062] An image acquisition module, which is used to collect and standardize the whole-slide image of cervical cells, unify the image brightness and color space, input a Papanicolaou-stained image, and output a standardized image in the LAB color space;

[0063] A suspicious positive cell detection module, which is communicatively connected to the image acquisition module, and is used to screen positive cells through the YOLOv7 object detection model and output a predicted box of suspicious positive cells including confidence and class probability;

[0064] A feature extraction and fusion module, which is communicatively connected to the suspicious positive cell detection module, includes a DINO pre-trained model and an attention mechanism network, and outputs cell-level features after fusing multi-scale features;

[0065] A multi-instance aggregation module, which is communicatively connected to the feature extraction and fusion module, and is used to generate sample-level features by extracting key instances and aggregating attention weights in the DSMIL two-stream network for all cell-level features;

[0066] An HPV classification decision module, which is communicatively connected to the multi-instance aggregation module, and is used to calculate and output the HPV infection status result by receiving the sample-level features, including probability values and binary classification labels;

[0067] A result output module, which is communicatively connected to the classification decision module, and is used to display the prediction result and visualize the key cell regions.

[0068] In a preferred embodiment, the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of any one of the above methods for directly predicting HPV using cervical cytopathological images.

[0069] In a preferred embodiment, the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored, wherein the computer program is executed by a processor to implement the steps of any one of the above methods for directly predicting HPV using cervical cytopathological images.

[0070] In a preferred embodiment, the present invention further provides a computer program product, wherein when the computer-readable instructions are executed by one or more processors, one or more processors are caused to execute the steps of any one of the above methods for directly predicting HPV using cervical cytopathological images.

[0071] The present invention provides a method for directly predicting HPV using cervical cytopathological images, which has the following beneficial effects by comparison with the prior art:

[0072] First, reduce the computational cost: First, screen out suspicious cells through the method of object detection, and migrate the method in the existing whole-slide histopathological images to the cytopathological images, thereby reducing the computational cost of subsequent tasks and facilitating the processing of large-scale data sets or real-time image analysis tasks;

[0073] Second, enhance the feature representation ability: The application of unsupervised learning and attention mechanism enables the model to automatically learn and emphasize the features that are most important for HPV detection in the image, which helps the model capture the deep HPV features different from the lesion features, thereby improving the diagnostic effect of the model;

[0074] Third, provide reference opinions: Directly and accurately predict the HPV infection status of ASCUS patients from cervical images, achieving the functions of reducing medical costs and providing reference opinions. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] The present invention will be further described below with reference to the drawings and embodiments:

[0076] Figure 1 It is the ASCUS type cell image used in the present invention;

[0077] Figure 2 It is the work flow chart of the present invention;

[0078] Figure 3 It is the flow chart of the method steps of the present invention;

[0079] Figure 4 It is the distribution diagram of ASCUS sample data used in the present invention;

[0080] Figure 5 It is the step flow chart of the classification model of the present invention;

[0081] Figure 6 It is the present invention Figure 5 The specific step flow chart in;

[0082] Figure 7 It is the structural schematic diagram of the system of the present invention;

[0083] Figure 8 It is the structural schematic diagram of the electronic device provided by the present invention. Detailed implementation manners

[0084] Clinically, it has been found that for patients of the ASCUS type, it is difficult to clearly determine whether there is pre-cancerous lesions clinically. To further clarify whether there may be pre-cancerous lesions, it is often necessary to further perform HPV testing. At present, there is no method to directly predict the HPV status from cervical images. Therefore, to solve this problem, this implementation has developed a method to directly predict the HPV status from cervical images. The prediction results can be used as a reference for doctors during diagnosis, reducing the treatment cost.

[0085] To better understand the purpose, structure and function of the present invention, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the drawings and in combination with the embodiments.

[0086] Embodiment 1

[0087] As Figures 1 - 6 shown, a method for directly predicting HPV using cervical cytopathological images, which includes the following steps:

[0088] S10. Data acquisition: Obtain the synchronous cervical cytopathological whole-slide images of the examinees who have undergone cervical tissue pathological analysis and whose conclusion evaluation is ASCUS;

[0089] In step S10, a pathological slide scanner with an optical magnification of 20 to 40 times is used to scan the input cervical cytopathological whole-slide image;

[0090] The output is a digital whole-slide image I, with an image size of n×n pixels, a unified color space of RGB mode, and a brightness range of [0.8, 1.2];

[0091] Among them, the whole-slide image of cervical cells is a Papanicolaou-stained whole-slide image. The Papanicolaou-stained whole-slide image is processed by color standardization to unify the brightness and color distribution of the Papanicolaou-stained whole-slide images collected by different devices. The standardization process includes histogram equalization and color space conversion.

[0092] S20. Lesion area localization: Use the trained object detection model to determine the abnormal cell area in the pathological whole-slide image of cervical cells, and screen the suspicious positive cells in the abnormal cell area;

[0093] In step 20, the specific steps are as follows:

[0094] S21. Predicting bounding boxes: The object detection model is YOLOv7. Input the digital whole-slide image I into YOLOv7;

[0095] Perform lightweight CNN operations on the digital whole-slide image I based on the YOLOv7 object detection model, and output a set of predicted bounding boxes B = {b1, b2,..., b n};

[0096] Among them, each predicted bounding box b i contains the coordinate information of x, y, w, h, the confidence level E(b i ) and the class information G(b i ), i represents an index value, and b i represents the i-th predicted bounding box in B = {b1, b2,..., b n};

[0097] Specifically, in order to meet the input size requirements of the YOLOv7 model, perform a resize operation on the digital whole-slide image I, and uniformly adjust the image to the size preset by the model, that is, the image size is 640×640 pixels;

[0098] At the same time, the YOLOv7 model internally processes the digital whole-slide image I through convolution, pooling, and activation operations. The YOLOv7 model first extracts features from the input digital whole-slide image I through the backbone network. Among them, the backbone network consists of multiple convolutional layers and CSP (Cross Stage Partial) modules. The CSP module divides the feature map in the channel dimension, performs convolutional operations separately and then merges them, which not only reduces the computational amount but also effectively improves the efficiency of feature extraction. And during the feature extraction process, feature maps of different scales are generated. The feature maps contain information at different levels of the image. The shallow feature maps retain more detailed information of the image, while the deep feature maps contain more semantic information.

[0099] In addition, the confidence level E(b i) indicates the likelihood that the bounding box contains suspicious positive cells, and the value range is usually [0, 1].

[0100] S22. Initial screening threshold: Set the confidence threshold T 阈值 = 0.5. According to E(b i ) ≥ T 阈值 Initial screening prediction bounding box b i Obtain the set of candidate bounding boxes

[0101] Based on the set of candidate bounding boxes B 候选 Mark the abnormal cell regions;

[0102] Specifically, in order to retain as many real suspicious cell regions as possible during the detection process and avoid missed detections, a relatively low confidence threshold T 阈值 is set. Using this to screen the bounding boxes, from all the predicted bounding boxes B generated by YOLOv7 and the lightweight CNN, select the predicted bounding boxes B with a confidence ≥ 0.5 to form the set of candidate bounding boxes B 候选 , and determine which of the abnormal cell regions framed by the set of candidate bounding boxes B 候选 are suspicious positive cells.

[0103] It should be noted that when the YOLOv7 model is detecting, even if the confidence E(b i ) of the predicted bounding box B is relatively low, it may still correspond to real diseased cells. If the confidence threshold T 阈值 is set too high, some real positive cells may be misjudged as negative, thus affecting the accuracy of subsequent detections.

[0104] S23. Morphological filtering: According to the candidate bounding boxes Calculate the specific morphological features of the abnormal cell regions framed;

[0105] Compare the specific morphological features to determine the final set of bounding boxes of suspicious positive cells

[0106] Among them, the specific morphological features include the area A and the circularity R;

[0107] In this embodiment, for the calculation of the area: For the cell region framed by the candidate bounding box , calculate its actual area A. By counting the pixels within the candidate bounding box , the value range is [0, 1]. If the pixel value in this region is not the background value, then this pixel is included in the area statistics.

[0108] Specifically, during implementation, the candidate bounding boxes can be traversed For each pixel within, determine that the pixel value is the identifier conforming to the cell region, thereby obtaining the number of pixels in the cell region, and then converting it into the actual area;

[0109] Assume that the actual area corresponding to each pixel is s square micrometers, then A = s × the number of pixels in the cell region.

[0110] Meanwhile, calculate the area A1 of the circumscribed rectangle of the cell, and its calculation formula is:

[0111] A1 = (x max - x min ) × (y max - y min );

[0112] Wherein, (x min , y min ) and (x max , y max ) are respectively the upper left corner and the lower right corner coordinates of the candidate bounding box .

[0113] In this embodiment, calculate the circularity: calculate the circularity R through the following formula, and the formula is:

[0114]

[0115] The circularity is used to measure the degree of approximation of the shape of the cell region to a circle, and the value range is [0, 1]. The closer the value is to 1, the closer the shape of the cell region is to a circle.

[0116] During implementation, when the cell region is a standard circle, its circularity R = 1; when the cell region is a very irregular shape, the circularity R will be much less than 1.

[0117] In this embodiment, for threshold setting and filtering: preset the area threshold A min and the circularity threshold R min 、R max ;

[0118] The area threshold A min is set based on the area statistical data of normal cervical cells and diseased cells. Generally speaking, the area of diseased cells is usually larger than the minimum area of normal cells. Through a large number of sample statistical analyses, a suitable A min value is obtained. The circularity thresholds R min and R max are also set based on the research on the shape characteristics of normal cells and diseased cells. The circularity of normal cells is usually within a certain range, while the circularity of diseased cells will deviate from this range. Therefore, cells with an area less than A min or a circularity not within R min to R maxFilter the bounding boxes within the range.

[0119] Specifically, during implementation, according to the set B of candidate bounding boxes 候选 for each bounding box calculate its corresponding area A and circularity R, and determine that it satisfies A < A min or R < R min or R > R max If one or several of the conditions are met, then remove this bounding box from the set B of candidate bounding boxes 候选 .

[0120] Final bounding box set generation: After morphological filtering, obtain the final set of suspected positive cell bounding boxes From the set B of candidate bounding boxes 候选 only retain the bounding boxes whose area and circularity both meet the requirements to form the final set B of suspected positive cell bounding boxes 最终 .

[0121] S24. Output result: According to the final set B of suspected positive cell bounding boxes 最终 crop the corresponding cell images in the digital whole slide image I to obtain the set C of suspected positive cell images = {c1, c2,..., c i}}.

[0122] Specifically, based on the final set B of suspected positive cell bounding boxes 最终 , crop the corresponding cell images from the original digital whole slide image I, so as to obtain the set C of suspected positive cell images for each sample of the original digital whole slide image I = {c1, c2,..., c i}}, where c i = Crop(I, b i ), and here Crop(·) represents the operation of image cropping based on the bounding box.

[0123] S30. Image feature extraction and fusion: Perform image feature extraction on the selected suspected positive cells based on a self-supervised learning architecture, perform lesion detection through a convolutional neural network, and fuse the image features through a self-attention mechanism to obtain cell-level features;

[0124] In step 30, the self-supervised learning architecture adopts the DINO framework, uses the ResNet50 pre-trained by DINO as the CNN backbone network, and the DINO framework performs contrastive learning pre-training on the unlabeled digital whole slide image I based on the ResNet50 backbone;

[0125] In this embodiment, 6,601 samples including 5,313 ASCUS samples were used. These samples were all prepared by the thin-layer liquid-based cytology technique (TCT). After the preparation was completed, they were scanned by a pathology slide scanner magnified 20 times to produce digital whole-slide (WSI) images I. These digital whole-slide images I were then detailedly labeled by professional cytopathologists to form the original cervical cytology dataset. The model was trained on ASCUS samples and tested using all types of samples. The sample types and quantities are shown in the following table:

[0126]

[0127]

[0128] Among them, the HPV genotyping of 2,100 ASCUS samples is as Figure 4 shown;

[0129] Through the contrastive learning strategy, the DINO framework can mine multi-level features from the suspicious positive cell image ci using multi-view data augmentation without manual annotation, and learn microscopic features such as cell texture, edge, and morphology on unlabeled data;

[0130] The ViT architecture of DINO can capture global context information, while ResNet50, as a CNN backbone network, is good at extracting local detail features, such as abnormal textures in the lesion area. Combining the two can provide multi-scale information input for lesion degree classification;

[0131] At the same time, experiments of DINO in the field of cervical cytopathology analysis show that its self-supervised pre-training features are more domain-adaptive than natural image supervised training, especially in data-scarce scenarios, which can improve the performance of downstream tasks (an average increase of 1.63%-4.62%), verifying the robustness of self-supervised features to the heterogeneity of medical images.

[0132] Input the suspicious positive cell image c i Extract the multi-level feature maps. The DINO framework performs global modeling on the multi-level feature maps extracted by ResNet50 through the Transformer encoder, and uses the self-attention mechanism to enhance the spatial correlation of the features;

[0133] Among them, the resolutions of the multi-level feature maps include downsamplings of 1 / 8, 1 / 16, and 1 / 32;

[0134] In addition, the DINO framework introduces Mixed Query Selection to screen key regions from the high-resolution feature maps generated by ResNet50 as the initial bounding boxes of the decoder, improving the small object detection ability.

[0135] The specific steps in step 30 are as follows:

[0136] S31. Self-supervised feature extraction: Input the suspicious positive cell image c screened in step S24 i into the DINO framework, and perform data augmentation operations on the input suspicious positive cell image c i to generate a multi-view set {I1, I2,..., I n};

[0137] Among them, the data augmentation operations include random cropping, rotation, and color jitter;

[0138] The DINO framework performs feature learning among the generated multiple views through a contrastive learning mechanism, and extracts the similarity measure of the features of each view I i ;

[0139] Let the feature extraction network in the DINO framework be f θ , and θ be the network parameters, then the feature z i of view I i is expressed as follows:

[0140] z i = f θ (I i );

[0141] According to the feature z i obtain the corresponding feature vectors {z1, z2,..., z n};

[0142] It should be noted that i represents an index value, and I i represents the i-th view in {I1, I2,..., I n}, and z i represents the i-th feature vector in {z1, z2,..., z n};

[0143] By minimizing the contrastive loss function L, it is promoted that the feature representations of similar views are close in the feature space, and the feature representations of different views are far away, and the distance between the feature vectors {z1, z2,..., z n} in multiple views is minimized;

[0144] In this embodiment, in order to more clearly illustrate the contrastive learning mechanism, take two views I1 and I2 as an example;

[0145] The contrastive loss between the feature vectors z1 and z2 can be represented by the following InfoNCE loss formula, and the formula is:

[0146]

[0147] Among them, the DINO framework uses contrastive learning to learn the similar features of cell images under different views, that is, those features that reflect the inherent structure and patterns of cells. By minimizing the contrastive loss, the DINO framework adjusts its own parameters so that the feature vectors of different views from the same suspicious positive cell image c i are as close as possible, while being distinguished from the feature vectors of other views. z k is the feature representation of other views randomly selected from the dataset, τ is the temperature parameter, and the temperature parameter τ controls the smoothness of the contrastive loss function. A smaller τ will make the contrastive learning more strict, emphasizing the differences between feature vectors more, and is used to regulate the intensity of contrastive learning. k represents the number of negative samples, and the number of negative samples k determines the number of other views participating in the calculation in contrastive learning, affecting the comprehensiveness and accuracy of model learning.

[0148] After multiple rounds of iterative training, the self-supervised feature h of the cell is extracted from the suspicious positive cell image c i by DINO; 自监督 ;

[0149] S32. Classification of positive and negative cells: As Figure 5 shown, the suspicious positive cell image c i screened in step S24 is imported into the hierarchical multi-scale feature fusion network as input data to subdivide the positive and negative of the cells. The network design adopts a two-branch structure, and the global feature branch and local feature branch based on W-MSA respectively extract the global information and local information of the input image, and use the feature fusion network to fuse the information;

[0150] The hierarchical multi-scale feature fusion network processes the input suspicious positive cell image c i layer by layer through convolutional layers, pooling layers, activation functions, and residual connection operations. The local branch and global branch of the hierarchical multi-scale feature fusion network operate in parallel in the network, each containing four processing stages. The local branch starts with a 4×4 convolutional layer with a stride of 4, and the global branch uses a patch partition module. In the input stage of the network, the Patch Embed module is responsible for cutting the high-resolution medical image into small pieces and encoding these image patches into a vector sequence for subsequent processing. The Patch Merging module is used to merge adjacent small pieces in the processing of different stages, reducing the spatial resolution of the feature map while increasing the channel depth, which helps the model to abstract and extract features at a deeper level. The image is divided into multiple 4×4 patches and unfolded in the channel direction, doubling the feature dimension through a linear embedding layer, and then performing deep feature transformation through the global feature block;

[0151] At each stage of the network, local features and global features are processed by a spatial attention mechanism and a channel attention mechanism respectively, and are input into a feature fusion module to generate fused features.

[0152] As Figure 6 shown, in the global feature branch stage based on W-MSA, the 256-sized cell image generated by the sample image passing through the YOLOv7 network for target inspection is input into the global feature module. W-MSA divides the input feature map into multiple windows of size M×M, and performs self-attention calculations independently for each window, thus greatly reducing the consumption of computing resources. Its computational complexity can be expressed by Formula 1 and Formula 2:

[0153] Ω(MSA) = 4hwC 2 + 2(hw) 2 C(1);

[0154] Ω(W-MSA) = 4hwC 2 + 2M 2 hwC(2);

[0155] Among them, h and w represent the height and width of the feature map respectively, C represents the number of channels of the feature map, and M represents the size of each window.

[0156] During the process of extracting features in the global feature module, the image first passes through a LayerNorm layer for normalization processing, and then is passed into the W-MSA module. The LayerNorm layer is mainly used for feature regularization to accelerate the training process of the neural network and improve the stability of the model. It calculates the mean and variance of the input features within a specific window, and uses these statistical data to normalize the features. The specific expression is as Formula 3:

[0157]

[0158] Among them, x i is the input feature, μ and σ 2 are the mean and variance of the features within the window respectively, ∈ is a very small number to avoid division by zero, and γ and β are learnable parameters used to further adjust the normalized data.

[0159] The feature map processed by LayerNorm will be passed into the linear layer, which mainly performs a simple linear transformation, converting the input features into a new feature space through a weight matrix. This process can be expressed as:

[0160] Linear(x) = Wx + b(4);

[0161] Among them, W is the weight matrix and b is the bias term.

[0162] Next, the feature map passes through the GELU (Gaussian Error Linear Unit) activation function, which is a non-linear function based on the Gaussian distribution and is used to enhance the model's non-linear processing ability for input features. The mathematical expression of the GELU function is:

[0163] GELU(x) = x·Φ(x)(5);

[0164] where Φ(x) is the standard normal cumulative distribution function of the input x.

[0165] The GELU function helps the model introduce non-linear dynamics while preserving the input information, improving the model's adaptability and interpretability for complex data. After each processing module, a residual connection and a relative position bias are applied to enhance the model's expressive power. In the next module, Shift W-MSA is also introduced to adjust the window position, further improving the model's ability to capture global information. This processing process can be described as:

[0166] g i = f 1×1 (W-MSA(LN(G i-1 )))+G i-1 (6);

[0167] G i = f 1×1 (SW-MSA(LN(g i )))+g i (7);

[0168] where G i and g i represent the output features after being processed by Shift W-MSA and W-MSA respectively, f 1×1 represents a convolution operation with a 1×1 convolution kernel size, which is actually equivalent to a linear transformation, and LN represents the LayerNorm layer.

[0169] In the local feature branch, the 256-sized cell images generated by the YOLOv7 network for target inspection of the sample image are also input into the local feature module and first processed by the backbone layer.

[0170] By using a convolutional kernel with a large stride, the backbone layer can significantly reduce the width and height of the input feature map. This method uses a 4×4 convolutional kernel and a stride of S = 4.

[0171] Assume the input image size is H×W×D (height, width, depth), the convolutional kernel size is 4×4, the stride is S = 4, and the number of output channels is N. Then the size of the output feature map is:

[0172]

[0173] Reduce the spatial dimension of the input data rapidly during preliminary processing while preserving important feature information as much as possible, providing a good foundation for deep and complex feature processing.

[0174] To enhance the feature expression and integration ability while reducing the number of parameters and computational complexity, the methods of pointwise convolution in advance and depthwise separable convolution are used.

[0175] The pointwise convolution in advance first processes the input feature map using a 1×1 convolutional kernel. The main purpose is to preprocess the features, such as dimensionality reduction and feature extraction. It can reduce the computational burden of subsequent operations.

[0176] After pointwise convolution, depthwise convolution is performed. This operation uses a separate convolutional kernel for each input channel for convolution. In this way, it can focus on extracting local features within a single channel. If the input feature map has M channels, then there will be M convolutional kernels, each with a size of 3×3, and each kernel operates only on the corresponding single channel without cross-channel information processing.

[0177] After depthwise convolution, it enters the 1×1 pointwise convolution layer. This step is usually used to expand or further refine the features. The number of convolutional kernels may increase to enhance the expressive ability of the feature map.

[0178] By increasing the number of pointwise convolution layers, the depth of the model can be significantly increased without overly increasing the computational burden, because the 1×1 convolution has a lower computational complexity compared to larger convolutional kernels.

[0179] After convolution operations, cross-channel information interaction is carried out through a linear layer. Considering the processing of single-cell data, which involves comparisons between different cells, and the batch size of training samples is small while the distribution consistency is difficult to guarantee, layer normalization can provide a more stable normalization effect compared to batch normalization. Therefore, LN is adopted. Since the task involves complex feature relationships, the activation function Swish is used. Its adaptive gating ability can provide additional flexibility and performance advantages when analyzing single-cell data:

[0180] Swish(x) = x·σ(x) (9);

[0181] where σ(x) is the Sigmoid function

[0182] The features after layer normalization and Swish activation are further fused through 1×1 convolution.

[0183] The cell images pass through the Transformer branch and the CNN branch to obtain global features and local features. This method uses feature fusion to obtain the feature vector of the final cell and generates subsequent feature maps {D2, D3,..., D n}, Meanwhile, the residual connection structure of ResNet50 directly adds the feature map of the previous layer to the feature map after operations such as convolution, that is: D n = D n-1 + f(D n-1 );

[0184] Among them, D n-1 is the feature map of the (n–1)th layer, f(·) represents a series of operation combinations such as convolution, batch normalization, and activation function. The residual connection can effectively solve the problem of gradient disappearance during the training of deep neural networks, enabling the model to effectively learn deeper and more abstract image features to output the generated feature map set {D1, D2,..., D n};

[0185] The Feature Fusion Module inputs the global features extracted by the Transformer branch into the channel attention mechanism. The channel attention mechanism performs global average pooling and global max pooling on the input feature map, and then generates a channel attention map through a network with shared weights. At the same time, the FFM inputs the local features extracted by the CNN branch into the spatial attention mechanism. The spatial attention mechanism first performs max pooling and average pooling on the channel dimension, and then uses a convolutional layer to generate a spatial attention map.

[0186] S33. Feature Fusion and Cell Positive Probability Prediction: Based on the feature map set {D1, D2,..., D n}, a channel attention module and a spatial attention module are constructed. For the channel attention module, first, global average pooling is performed on each feature map D i to obtain the channel feature vector g i . The channel feature vector g i is processed through two fully connected layers to generate the channel attention weight vector w i . The channel attention weight vector w i is multiplied element-wise with the feature map D i to highlight the channel features important for determining the degree of cell lesion, obtaining the feature map enhanced by channel attention

[0187] Next, a spatial attention module is constructed for the feature map . All the feature maps enhanced by channel attention Concatenate on the channel dimension, and then generate a spatial attention map M through a 7×7 convolution operation. Multiply the spatial attention map M element-wise with the concatenated feature map to obtain a feature map with enhanced spatial attention.

[0188] The self-supervised feature F 自监督 is concatenated with the feature vector corresponding to the feature map with enhanced spatial attention during the ResNet50 classification process. The feature vectors of the two are connected in sequence and fused into a new feature vector F. 融合 ;

[0189] The fused feature vector F 融合 is input into a fully connected layer for linear transformation, and then the probability P of positive cells is obtained through the Sigmoid function. The formula of the Sigmoid function is as follows:

[0190]

[0191] where x is the output of the fully connected layer.

[0192] Finally, the probability P of positive cells is combined with the cell positive feature F extracted by the DINO pre-trained ResNet50 model 自监督 again, and linear transformation and fusion are performed through another fully connected layer to obtain the cell-level feature F. 细胞 .

[0193] S40. Prediction of HPV infection status: Adopt a multi-instance learning network to aggregate cell-level features to generate sample-level features, and input them into a classifier to determine the HPV infection status result of the sample-level features.

[0194] The specific steps in step S40 are as follows:

[0195] According to the cell-level feature F in step S33 细胞 obtain the cell instance feature h i (i = 1, 2, …, N);

[0196] where the cell instance feature h i is obtained by successively performing image acquisition, preprocessing, suspicious cell detection and preliminary screening, cell-level feature fusion and disease variable quantification on the cervical cell pathological image.

[0197] According to multi-instance learning, for a suspicious positive cell image c i containing N cell instances, if there is at least one HPV-related lesion cell among them, then determine that the current suspicious positive cell image c i is HPV positive, that is, the HPV infection status label y of the sample. 感染= 1 if and only if there exists at least one i such that the i-th cell instance feature h i corresponds to the HPV-related lesion cell feature, and based on this, an overall framework for converting and classifying from cell-level features to sample-level features is constructed;

[0198] Key instance extraction flow steps: The extracted cell-level feature F 细胞 generates sample-level features through the DSMIL network and performs classification prediction. The DSMIL network contains two parallel feature aggregation channels to more effectively achieve sample-level aggregation and classification prediction of cervical cell instance features h i of.

[0199] The first flow is the key instance extraction flow. This flow scores all cell instance features h i through an instance scoring network, and uses the max-pooling mechanism to select the instance with the highest score as the key instance h m . The selection process of the key instance h m can be expressed as:

[0200] c m (B) = max{W0h i}, i = 1, 2, …, N;

[0201] where W0 is the learnable weight parameter in the instance scoring network, h i is the feature representation of the i-th cell instance, N is the total number of instances, and each cell instance feature is transformed and scored by calculating W0h i .

[0202] Instance attention aggregation flow steps: The second flow of the DSMIL network is the instance attention aggregation flow. This flow assigns dynamic weights to each cell instance feature h i through a non-local attention mechanism;

[0203] In the first step, each cell instance feature h i is mapped into a query vector q q and an information vector v v respectively through the linear transformation matrices W i = W q h i and W i = W v h i .

[0204] In the second step, calculate the attention weight U(h m , h i , h m ) between each instance and the key instance h

[0205]

[0206] Among them, represents the feature similarity between the i-th instance and the key instance, and U(h i , h m ) is the attention weight of the instance.

[0207] According to the attention weight U(h i , h m ), weighted aggregation of the information vectors of all cell instance features h i is performed to form a sample-level feature representation

[0208] Finally, the key instance feature h m obtained by the max-pooling stream is fused with the sample feature b aggregated by the instance attention stream to obtain the final aggregated sample-level feature H 样本 :

[0209]

[0210] Among them, W0 and W b are trainable parameter matrices for post-fusion, and H 样本 is the final aggregated sample-level feature;

[0211] Feature enhancement step: Obtain cell-level statistics, where the cell-level statistics include the positive cell ratio p 阳性 and the highest lesion score s 最高 . The aggregated sample feature H 样本 is concatenated with the cell-level statistics to obtain an enhanced sample-level feature

[0212] The calculation formula for the positive cell ratio is as follows:

[0213]

[0214] Among them, y i indicates that the i-th cell instance is determined to be a positive cell;

[0215] The calculation formula for the highest lesion score is as follows:

[0216] s 最高 = max{s1, s2, …, s n};

[0217] Among them, s i is the lesion score of the i-th cell instance;

[0218] Multi-layer perceptron classification step: Use the enhanced sample-level feature It is input into a multi-layer perceptron (MLP) containing two layers of fully connected layers and the Dropout technique between layers. The first fully connected layer is calculated through the weight matrix W1 and the bias b1, and the calculation formula is as follows:

[0219]

[0220] The second fully connected layer is calculated through the weight matrix W2 and the bias b2, and the calculation formula is as follows:

[0221] y 预测 = σ(W2z1 + b2);

[0222] Among them, ReLU is the activation function, and σ is the Sigmoid function;

[0223] The HPV positive probability y of the sample is obtained as the output 预测 result.

[0224] To verify the effect of the HPV prediction model in actual prediction tasks, two experiments are designed in this embodiment for verification. First, it is to test the result of the HPV positive probability y 预测 of the sample obtained as the output. This embodiment is compared with other WSI classification methods (such as Mean-Pooling, Max-Pooling, ABMIL, CLAM, etc.). Specifically in the experiment, we use four indicators to verify the performance of the method, namely accuracy, precision, recall, and F1-score; the effect is as follows in the figure:

[0225]

[0226]

[0227] The HPV prediction accuracy of this embodiment on ASCUS samples reaches 76.1%, and the F1 value is 0.68, which is better than other multi-instance learning methods.

[0228] Secondly, it is to test other types of cell images to verify the generalization performance of the model.

[0229] Type Accuracy Precision Recall F1 AGC 0.8148 0.333 0.563 0.417 ASCUS 0.761 0.730 0.637 0.680 LSIL 0.743 0.787 0.865 0.824 ASC-H 0.788 0.764 0.826 0.794 HSIL 0.887 0.894 0.988 0.937

[0230] It can be seen that the LSIL test results show that the accuracy of the model is about 78%, and the F1 score is about 82%, clearly indicating good generalization. Compared with ASCUS, the LSIL samples perform slightly better in the test, which may be due to the higher degree of LSIL lesions, making the positive features more obvious. The remaining test results also show an upward trend as the degree of lesions increases.

[0231] Among the 32 patients used for testing, the method in this embodiment correctly predicted the HPV infection status in 24 cases and reduced unnecessary HPV tests in 13 cases.

[0232]

[0233]

[0234] From the test results, it is known that this embodiment directly predicts the HPV infection status, avoiding unnecessary HPV tests for 40% of ASCUS patients and reducing medical costs by about 30%.

[0235] Embodiment 2

[0236] Next, a device for directly predicting HPV using cervical cell pathological images provided by the present invention will be described. The device for directly predicting HPV using cervical cell pathological images described below can be correspondingly referred to the method for directly predicting HPV using cervical cell pathological images described above and will be further described in combination with Embodiment 1;

[0237] As Figure 5 shown in the structure, Figure 5 is a device for directly predicting HPV using cervical cell pathological images provided by an embodiment of the present application, which includes:

[0238] In this embodiment, an image acquisition module is used to collect and standardize the whole-slide image of cervical cells, unify the image brightness and color space, input a Papanicolaou-stained image, and output a standardized image in the LAB color space;

[0239] In this embodiment, a suspicious positive cell detection module is communicatively connected to the image acquisition module and is used to screen positive cells through a YOLOv7 object detection model and output a prediction box of suspicious positive cells containing confidence and class probability;

[0240] Among them, the suspicious positive cell detection module uses an object detection model based on the YOLO series, supports dynamic adjustment of the confidence threshold (0.3 - 0.5), and the output format is a JSON file, including cell coordinates, confidence E(b i ) and class information G(b i );

[0241] In this embodiment, a feature extraction and fusion module is communicatively connected to the suspicious positive cell detection module, including a DINO pre-trained model and an attention mechanism network, and outputs cell-level features after fusing multi-scale features;

[0242] Specifically, the feature extraction and fusion module includes:

[0243] Pre-trained Feature Extraction Unit: The ResNet50 backbone network is used to extract 1024-dimensional cell features;

[0244] Attention Fusion Unit: Calculate channel weights through the channel attention module (Squeeze-and-Excitation), calculate spatial weights through the spatial attention module (Convolutional Block Attention Module), and output the fused feature vector.

[0245] In this embodiment, the multi-instance aggregation module is communicatively connected to the feature extraction and fusion module, and is used to generate sample-level features by extracting key instances and aggregating attention weights in the DSMIL dual-stream network for all cell-level features;

[0246] In this embodiment, the HPV classification decision module is communicatively connected to the multi-instance aggregation module, and is used to calculate and output the HPV infection status result by receiving the sample-level features, including probability values and binary classification labels;

[0247] In this embodiment, the result output module is communicatively connected to the classification decision module, and is used to display the prediction result and visualize the key cell regions.

[0248] Among them, the result output module supports the following functions:

[0249] Visualization display: Superimpose suspicious cell bounding boxes and heatmaps on the whole-slide image;

[0250] Report generation: Output a diagnostic report in PDF format, including HPV infection status, confidence level, and clinical suggestions

[0251] Embodiment 2

[0252] Further described in combination with Embodiment 1, as Figure 8 shown in the structure, Figure 8 This is a schematic structural diagram of the electronic device provided by the embodiment of the present application. The electronic device may include:

[0253] A processor, a memory, a communication bus, and a computer program stored on the memory and executable on the processor.

[0254] The processor can call the computer program in the memory, and when the program is executed, it can implement the method for directly predicting HPV using cervical cell pathological images provided in the above embodiments. The method includes: obtaining the synchronous whole-slide cervical cell pathological image of the examinee who has undergone cervical tissue pathological analysis and whose conclusion assessment is ASCUS; using the trained object detection model to determine the abnormal cell regions in the whole-slide cervical cell pathological image, and screening the suspicious positive cells in the abnormal cell regions; performing image feature extraction on the screened suspicious positive cells based on the self-supervised learning architecture, performing lesion prediction through a convolutional neural network, and fusing image features through a self-attention mechanism to obtain cell-level features; using a multi-instance learning network to aggregate the cell-level features to generate sample-level features, and inputting them into a classifier to determine the HPV infection status result of the sample-level features.

[0255] Further, the electronic device further includes:

[0256] A communication interface, which is used for communication between the memory and the processor.

[0257] A memory, which is used to store the computer program that can run on the processor.

[0258] The memory may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.

[0259] If the memory, the processor, and the communication interface are implemented independently, the communication interface, the memory, and the processor can be interconnected through a communication bus and complete communication with each other. The communication bus can be an Industry Standard Architecture (ISA) communication bus, a Peripheral Component Interconnect (PCI) communication bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 5 only a thick line is used to represent it in the figure, but it does not mean that there is only one communication bus or one type of communication bus.

[0260] In addition, when the logical instructions in the above-mentioned memory can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0261] The processor may include one or more processing units. For example, the processor may include an application processor (abbreviated as AP), an application specific integrated circuit (abbreviated as ASIC), a modem processor, a central processing unit (abbreviated as CPU), an image signal processor (abbreviated as ISP), a controller, a memory, a video codec, a digital signal processor (abbreviated as DSP), a baseband processor, and / or a neural-network processing unit (abbreviated as NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors. Among them, the controller may be the nerve center and command center. The controller can generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching and executing instructions. A memory may also be provided in the processor for storing instructions and data. In some embodiments, the memory in the processor is a cache memory. This memory can save the instructions or data that the processor has just used or recycled. If the processor needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor, and thus improves the efficiency of the system.

[0262] Optionally, in specific implementation, if the memory, the processor, and the communication interface are integrated on a single chip, the memory, the processor, and the communication interface can complete mutual communication through an internal interface.

[0263] On the other hand, an embodiment of the present application also provides a computer non-transitory readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the method for directly predicting HPV using cervical cytopathological images as described above. The method includes: obtaining synchronous whole-slide cervical cytopathological images of the examinees who have undergone cervical tissue pathological analysis and whose conclusion is evaluated as ASCUS; using a trained object detection model to determine the abnormal cell regions in the whole-slide cervical cytopathological images, and screening for suspicious positive cells in the abnormal cell regions; performing image feature extraction on the screened suspicious positive cells based on a self-supervised learning architecture, performing lesion prediction through a convolutional neural network, and fusing image features through a self-attention mechanism to obtain cell-level features; using a multi-instance learning network to aggregate cell-level features to generate sample-level features, and inputting them into a classifier to determine the HPV infection status result of the sample-level features.

[0264] On the other hand, an embodiment of the present application also provides a computer program product. The computer program product includes a computer program, which can be stored on a non-transitory computer-readable storage medium. The computer program can run computer instructions. When the computer program is executed by a processor, the computer can execute the method for directly predicting HPV provided by each of the above methods. The method includes: obtaining synchronous whole-slide cervical cytopathological images of the examinees who have undergone cervical tissue pathological analysis and whose conclusion is evaluated as ASCUS; using a trained object detection model to determine the abnormal cell regions in the whole-slide cervical cytopathological images, and screening for suspicious positive cells in the abnormal cell regions; performing image feature extraction on the screened suspicious positive cells based on a self-supervised learning architecture, performing lesion prediction through a convolutional neural network, and fusing image features through a self-attention mechanism to obtain cell-level features; using a multi-instance learning network to aggregate cell-level features to generate sample-level features, and inputting them into a classifier to determine the HPV infection status result of the sample-level features.

[0265] The logic and / or steps represented in the flowchart or otherwise described herein can, for example, be considered as a definitional sequence of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. As used in this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. More specific examples (but not an exhaustive list) of the computer-readable medium include the following: an electrical connection having one or N wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable medium on which the program can be printed, as the program can be obtained electronically by optically scanning the paper or other medium, followed by editing, interpretation, or otherwise processing as appropriate, and then stored in a computer memory.

[0266] The storage medium mentioned above can be a read-only memory, a magnetic disk, an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A method for directly predicting HPV using cervical cell pathological images, characterized in that, It includes the following steps: S10. Data acquisition: Obtain the synchronous whole-slide cervical cell pathology images of the examinees who have undergone cervical tissue pathological analysis and whose conclusion assessment is ASCUS; S20. Lesion area localization: Use the trained object detection model to determine the abnormal cell areas in the whole-slide cervical cell pathology images, and screen the suspicious positive cells in the abnormal cell areas; S30. Image feature extraction and fusion: Perform image feature extraction on the screened suspicious positive cells based on the self-supervised learning architecture, perform lesion prediction through a convolutional neural network, and fuse the image features through a self-attention mechanism to obtain cell-level features; S40. HPV infection status prediction: Use the multi-instance learning network to aggregate the cell-level features to generate sample-level features, and input them into the classifier to determine the HPV infection status results of the sample-level features.

2. The method for directly predicting HPV using cervical cytopathological images according to claim 1, characterized in that, In step S10, a pathological section scanner with an optical magnification of 20 to 40 times is used to scan the input whole-slide cervical cell images; The output is a digital whole-slide image I, with an image size of n×n pixels, the color space is uniformly in the RGB mode, and the brightness range is [0.8, 1.2]; Among them, the whole-slide cervical cell image is a Papanicolaou-stained whole-slide image, and the Papanicolaou-stained whole-slide image is processed by color standardization to unify the brightness and color distribution of the Papanicolaou-stained whole-slide images collected by different devices. The standardization process includes histogram equalization and color space conversion.

3. The method for directly predicting HPV using cervical cytopathological images according to claim 2, wherein In step 20, the specific steps are as follows: S21. Predict the bounding box: The object detection model is YOLOv7, and the digital whole-slide image I is input into YOLOv7; Perform lightweight CNN operations on the digital whole slide image I based on the YOLOv7 object detection model, and output the predicted bounding box set B = {b1, b2,..., b n}; Among them, each predicted bounding box b i contains the coordinate information of x, y, w, h, the confidence E(b i ), and the category information G(b i ). Here, i represents an index value, and b i represents the i-th predicted bounding box in B = {b1, b2,..., b n}; S22. Initial screening threshold: Set the confidence threshold T 阈值 = 0.

5. According to E(b i ) ≥ T 阈值 Initial screening prediction bounding box b i Obtain the candidate bounding box set Based on the set B of candidate bounding boxes 候选 Mark the abnormal cell regions; S23. Morphological filtering: Based on the candidate bounding boxes calculate specific morphological features of the abnormal cell regions framed; Determine the set of bounding boxes of the final suspicious positive cells by comparing specific morphological features Among them, the specific morphological features include area A and circularity R; S24. Output result: According to the set B of the final suspected positive cell bounding boxes 最终 Crop the corresponding cell images in the digital whole slide image I to obtain the set C of suspected positive cell images C = {c1, c2, …, c i}.

4. The method for directly predicting HPV using cervical cytopathological images according to claim 3, wherein In step 30, the self-supervised learning architecture uses the DINO framework, and the ResNet50 pre-trained by DINO is used as the CNN backbone network. The DINO framework performs contrastive learning pre-training on the unlabeled digital whole-slide image I based on the ResNet50 backbone; Input the suspiciously positive cell image c i Extract multi-layer feature maps. The DINO framework performs global modeling on the multi-layer feature maps extracted by ResNet50 through a Transformer encoder, and uses the self-attention mechanism to enhance the spatial correlation of the features; Among them, the resolutions of the multi-layer feature maps include downsampling at 1 / 8, 1 / 16, and 1 / 32.

5. The method for directly predicting HPV using cervical cytopathological images according to claim 4, wherein The specific steps in step 30 are as follows: S31. Self-supervised feature extraction: Input the suspicious positive cell image c screened in step S24 i into the DINO framework, and perform data augmentation operations on the input suspicious positive cell image c i to generate a multi-view set {I1, I2,..., I n}; Extract the features z of each view I based on the DINO framework i ; i According to the features z i obtain the corresponding feature vectors: {z1, z2,..., z n}; where i represents an index value, I i represents the i-th view in {I1, I2,..., I n}, and z i represents the i-th eigenvector in {z1, z2,..., z n}; The DINO framework minimizes the distance between the feature vectors {z1, z2,..., z n} in multiple views through a contrastive learning mechanism; The final image features F extracted by DINO are obtained through multiple rounds of iterative training 自监督 ; S32. Classification and feature fusion of positive and negative cells: Input the suspicious positive cell image c screened in step S24 i into the hierarchical multi-scale feature fusion network to obtain the positive and negative cell categories; The multi-scale feature fusion network performs feature extraction on the input suspicious positive cell image c i to obtain a global feature vector and a local feature vector, fuse the global feature vector and the local feature vector, and output a set of generated feature maps {D1, D2,..., D n}; S33. Cell positive probability prediction: Based on the set of feature maps {D1, D2,..., D n}, construct a channel attention module and a spatial attention module; For each feature map D i perform global max pooling and global average pooling respectively to obtain the channel feature vector g i ; According to the channel feature vector g i Generate the channel attention weight vector w through the fully connected layer i ; According to the channel attention weight vector w i and the feature map D i multiply to obtain the feature map with enhanced channel attention Concatenate all the feature maps in the channel dimension and generate the spatial attention map M through the convolutional layer; Multiply the spatial attention map M element-wise with the concatenated feature map to obtain the feature map with enhanced spatial attention Stitched self-supervised feature F 自监督 and the feature map obtained by the multi-scale feature fusion network corresponding to the feature vector, and output a new feature vector F 融合 ; Input the feature vector F into the fully connected layer 融合 , and calculate the probability P of positive cells through the Sigmoid function; According to the probability P of positive cells and the positive cell feature F 自监督 Determine the cell level feature F 细胞 .

6. The method for directly predicting HPV using cervical cytopathological images according to claim 5, wherein In step S40, the multi-instance learning network is the DSMIL network, which includes two parallel feature aggregation channels. The first stream is the key instance extraction stream, and the second stream is the instance attention aggregation stream; According to the cell-level feature F in step S33 细胞 Obtain the cell instance feature h i (i = 1, 2, …, N); The key instance extraction scores the cell instance feature h through the instance scoring network i for scoring; Select the cell instance feature h with the highest score using the max pooling mechanism i and output the key instance h m ; The selection formula is as follows: c m (B) = max{W0h i} for i = 1, 2, …, N; Among them, W0 is the learnable weight parameter in the instance scoring network, and h i is the feature representation of the i-th cell instance, and N is the total number of instances; Map each cell instance feature h i to a query vector q i and an information vector v i ; Calculate the feature h of each cell instance i With the key instance h m The attention weight U(h i , h m ), the formula is as follows: Among them, represents the feature similarity between the i-th instance and the key instance; Aggregate the information vectors of cell instance features h i , h m ) according to the attention weight U(h i ) to obtain the sample feature b by weighted aggregation; Fuse the key instance feature h m with the sample feature b to output the final aggregated sample-level feature H 样本 ; The feature fusion formula is as follows: Among them, W0 and W b are trainable parameter matrices for fusion; Obtain cell-level statistics, and splice the cell-level statistics with the aggregated sample-level feature H 样本 to obtain an enhanced sample-level feature Input enhanced sample-level features into the classifier Determine the HPV positive probability y 预测 .

7. An apparatus for directly predicting HPV using cervical cell pathological images, characterized in that, It includes: An image acquisition module, which is used to collect and standardize the whole-slide cervical cell images, unify the image brightness and color space, and the input is a Papanicolaou-stained image, and the output is a LAB color space standardized image; A suspicious positive cell detection module, which is communicatively connected to the image acquisition module, and is used to screen positive cells through the YOLOv7 object detection model and output a suspicious positive cell prediction box containing confidence and class probability; A feature extraction and fusion module, which is communicatively connected to the suspicious positive cell detection module, includes a DINO pre-trained model and an attention mechanism network, and outputs cell-level features after fusing multi-scale features; Multi-instance aggregation module, communicatively connected to the feature extraction and fusion module, is used to generate sample-level features by extracting key instances and aggregating attention weights in the DSMIL dual-stream network for all cell-level features; HPV classification decision module, communicatively connected to the multi-instance aggregation module, is used to calculate and output the HPV infection status result by receiving the sample-level features, including probability values and binary classification labels; Result output module, communicatively connected to the classification decision module, is used to display the prediction result and visualize the key cell regions.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the steps of the method for directly predicting HPV using cervical cell pathological images according to any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program is executed by the processor to implement the steps of the method for directly predicting HPV using cervical cell pathological images according to any one of claims 1 to 6.

10. A computer program product, characterized in that, When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the steps of the method for directly predicting HPV using cervical cell pathological images according to any one of claims 1 to 6.

Citation Information

Cited By

  • Specific image processing method and computer readable storage medium

    CN121280864A

  • Pathological image cervical cancer screening model construction method based on liquid-based cytology

    CN121437462A

  • Nasal polyp pathological section image recognition method and device, electronic equipment and storage medium

    CN121999285A