Detection and differentiation of gynecologic lesions in colposcopy

A deep learning method using transfer learning and CNNs enhances colposcopy accuracy by over 90% in distinguishing LSIL and HSIL lesions, addressing diagnostic challenges in cervical cancer screening.

WO2025211973A1PCT designated stage Publication Date: 2025-10-09GYNOAID ARTIFICIAL INTELLIGENCE DEVELOPMENT SA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/PT2024/050037
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-04
Filing Date
2024-10-23
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Colposcopy for cervical cancer screening has suboptimal diagnostic accuracy and high interobserver variability, particularly in distinguishing between low-grade squamous intraepithelial lesions (LSIL) and high-grade squamous intraepithelial lesions (HSIL), leading to false positives and underdiagnosis.

Method used

A deep learning-based method using transfer learning and semi-active learning to classify colposcopy images, employing convolutional neural networks (CNNs) for lesion detection and differentiation, utilizing images from various phases of the colposcopy exam.

Benefits of technology

Achieves over 90% accuracy in identifying and classifying LSIL and HSIL lesions, reducing interobserver variability and improving diagnostic reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure PT2024050037_09102025_PF_FP_ABST
    Figure PT2024050037_09102025_PF_FP_ABST
Patent Text Reader

Abstract

Detection and differentiation of gynecologic lesions in colposcopy The present invention relates to a computer-implemented method capable of automatically detecting gynecologic lesions in image / videos data, by classifying pixels as lesion or non-lesion, using a convolutional image feature extraction step followed by a classification step and indexing such lesions in the set of one or more classes.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] DESCRIPTION

[0002] DETECTION AND DIFFERENTIATION OF GYNECOLOGIC LESIONS IN

[0003] COLPOSCOPY

[0004] Background of the invention

[0005] The present invention relates to lesion detection and differentiation in medical image data . More particularly, to automated identification and classification of gynecologic lesions in images or videos acquired during a colposcopy exam, to assess the lesion seriousness and subsequent medical treatment .

[0006] Cervical cancer is highly prevalent , with recent estimates suggesting a worldwide incidence of 13.3 cases per 100 000 women-years

[0001] . The cervical cancer burden is particularly significant in low-income countries , with higher incidence and mortality . Human papillomavirus (HPV) infection plays a fundamental role in the pathophysiology of this disease . HPV infection is highly prevalent, occurring in over 70% of sexually active women and men at some point in their lives . Most HPV infections are temporary, nevertheless , HPV persistent infection is commoner with high-risk oncogenic types

[0002] . Cervical squamous cell carcinoma is the main HPV- related cancer , associated with 7 .5% of all female cancer deaths , mainly in less developed countries

[0003] .

[0007] Despite the discouraging prognosis of advanced disease stages , cervical cancer can be easily managed if detected in an early stage

[0004] . The 90-70-90 targets for cervical cancer showcase the importance of an adequate vaccination rate associated with HPV screening and early treatment of invasive disease

[0005] . Therefore , cervical cancer screening is typically performed with HPV testing and / or cytological examination

[0006] . In case of altered results , a colposcopy is indicated .

[0008] Colposcopy is the recommended exam in the setting of a positive HPV test or patients with altered cytology . The direct visualization of cervical mucosa , enabled by the magnification of a colposcopy device with biopsy of suspected lesions , is the mainstay for diagnosis of cervical cancer . Nevertheless , colposcopy presents variable diagnostic accuracy, with low specificity for high-grade squamous lesions . Additionally, this exam depends heavily on the clinician' s skill and has high intra and interobserver variability [ 7 , 8 ] . Recently, a systematic review and metaanalysis pointed to women 50 years of age or older , in a postmenopausal status or with a transformation zone 3 type at risk lesion underdiagnosis by colposcopy -guided biopsies

[0009] .

[0009] Indeed, the discrimination between low and high-grade squamous intraepithelial lesions , from now on referred to as HSIL and LSIL , respectively, is a matter of great clinical importance . Typically, low-grade squamous intraepithelial lesions (LSIL) are associated with less risky HPV types and are not associated with a significantly increased risk for invasive cancer

[0010] .

[0010] Therefore , the main objective of colposcopy in cervical cancer screening is to identify high-grade squamous intraepithelial lesions (HSIL) before they develop into cervical cancer . On the other hand, HSIL have a significant risk of progression to invasive cancer , present more evident dysplastic changes , and are thereby considered precancerous

[0011] . Nevertheless , colposcopy diagnostic accuracy for the identification of HSIL lesions is suboptimal , with a recent study demonstrating an accuracy of nearly 70% for the identification of this lesion type

[0012] .

[0011] Thereby, the accuracy of colposcopic evaluation in cervical cancer screening is highly dependent not only of the clinician' s expertise , but also on the usage of specific agents as acetic acid and lugol ' s iodine . Generally, HSIL lesions are typically whitish lesions enhanced by the acetic acid colorations , while being negative for lugol ' s iodine coloration

[0013] .

[0012] While prospective studies support the sequential application of acetic acid and lugol ’ s staining , increasing the diagnostic performance of colposcopic procedures essentially through heightened sensitivity, there are inherent limitations to this approach

[0013] . Suboptimal specificity and inter-observer variability contribute to a notable number of false positives , even when incorporating staining in colposcopy . Challenges in achieving optimal accuracy persist , emphasizing the need for ongoing refinement to enhance the overall reliability of cervical cancer detection during colposcopy .

[0013] Therefore , in a context of suboptimal diagnostic accuracy and high interobserver variability, artificial intelligence models could have a role in increasing colposcopy costeffectiveness . Additionally, the large numbers of images produced in colposcopy exams favor the use of artificial intelligence tools for image analysis . Indeed, convolutional neural networks (CNN) are a human-based architecture inspired in the human visual cortex suitable for image patterns analysis

[0014] . Indeed, the benefit of CNN implementation in the analysis of imaging patterns has been documented in several medical areas [ 15-17 ] . The present invention provides an automatic optimal method to select a deep learning model for differentiation between LSIL and HSIL in colposcopy exams , using images of all the phases of colposcopy exam, namely non-stained, acetic-acid, lugol-iodine or post-manipulation frames .

[0014] Brief summary of the invention

[0015] The present invention provides a method for deep learningbased detection and classification of gynecologic lesions , namely LSIL or HSIL , in colposcopy images .

[0016] Colposcopies were performed and collected together with follow-up biopsies of the suspected lesions . Each procedure may be divided into four segments : initial non-stained observation , 3% acetic acid solution observation , lugol observation and after therapeutic manipulation (e . g . after laser ablation , plasma coagulation or surgical ablation) . Dataset includes frames from these four categories , and each collected procedure may include any combination of them.

[0017] The following were considered relevant to highlight the problem solved by the present invention from the methods known in the art to detect and classify lesions in colposcopy .

[0018] In the preferred and most generic embodiment of the invention , a method that automatically detects relevant gynecologic lesions in colposcopy images or videos is presented . Lesions identification in gynecologic colposcopy is vital to assess carcinoma probability . In the same preferred embodiment , the invention makes use of transfer learning methods and semi -active learning to optimize the computational efficiency and accuracy of the method . While transfer learning allows feature extraction and high- accuracy classification , using reasonable dataset sizes , semi -active implementation allows a continuous improvement in the classification system. Preferably, the invention uses transfer learning for feature extraction of colposcopic images with overall accuracy over 90% and employs a semiactive learning strategy for colposcopic images .

[0019] In another embodiment of the invention , a method that splits the dataset into a number of stratified folds , where images relative to a given patient are included in one-fold only . Further , additionally or alternatively, such data is trained and validated with patient grouping to a random fold, i . e . , images from an arbitrary patient belong to either the training or the validation set .

[0020] Preferred is a method which uses the chosen training and validation sets to further train a series of network architectures , which include , among others , a feature extraction , and a classification component . The series of convolutional neural networks to train include but are not limited to : VGG16 , InceptionV3 , Xception EfficientNetB5 , EfficientNetB7 , Resnet50 , and Resnetl25 . Preferably, their weights are frozen , with exception to the BatchNormalization layers , and are coupled with a classification component . The classification component comprises at least two dense layers , preferably of sizes 2048 and 1024 , and at least one dropout layer of preferably 0 . 1 in between them.

[0021] Alternatively, but not preferentially, the classification component can be used with more dense layers or with dense layers of different size . Alternatively, but not preferentially, the classification component can also be used without dropout layers .

[0022] Further , additionally, and preferably, the best performing architecture is chosen according to the overall accuracy and sensitivity - Performance metrics include but are not limited to fl-metrics . Further , the method is not limited to two to four dense layers in sequence , starting with 4096 and decreasing in half up to 512 . Between the final two layers there is a dropout layer of 0 . 1 drop rate .

[0023] Lastly, the best performing solution is trained using the complete dataset with patient grouping .

[0024] Further embodiments of the present invention may include similar classification networks , training weights and hyperparameters .

[0025] These may include the usage of any image classification network , new or not yet designed .

[0026] In general , the method includes two modules : prediction and output collector . Prediction reads videos and flags images with findings . Conversely, the output collector passes these images with findings for processing .

[0027] Examples of advantageous effects of the present invention include : training using parameters from machine learning results of cloud-based every-day increasing datasets ; automatically prediction of the colposcopic image by using a deep learning method so that the neoplastic lesions from image input of the colposcopy can be identified and classified into either low or high grade squamous intraepithelial lesion , the usage of transfer learning improves the image classification speed and corresponding classification accuracy .

[0028] Brief description of the drawings

[0029] FIG . 1 illustrates a method for detection of neoplastic lesions in gynecologic colposcopy according to an embodiment of the present invention . FIG . 2 illustrates the method for automatic detection and differentiation of gynecologic lesions in colposcopy .

[0030] FIG . 3 illustrates the major processes for automatic detection and differentiation of gynecologic lesions in colposcopy .

[0031] FIG . 4 illustrates the structure of the classification network for gynecologic lesions .

[0032] FIG . 5 depicts an embodiment of the classification network to classify gynecologic lesions .

[0033] FIG . 6 illustrates a preferred embodiment of the present invention where the accuracy curves for the training on a small subset of images and labelled data are shown .

[0034] FIG . 7 illustrates exemplary accuracy curves during training on a small subset of images and labelled data and according to an embodiment of the present invention .

[0035] FIG . 8a , 8b , 8c illustrate exemplary Receiver-operating characteristic curves (ROC curves) and Area under the curve (AUC) values obtained after training on a small subset of images and labelled data in cervix (8a) , vagina (8b) and vulva (8c) .

[0036] FIG . 9a , 9b and 9c illustrate an exemplary confusion matrix after training on a small subset of images and labelled data in cervix ( 9a) , vagina ( 9b) and vulva ( 9c) .

[0037] FIG . 10a , 10b and 10c show examples of lesion classification , in cervix ( 10a) , vagina ( 10b) and vulva ( 10c) according to an embodiment of the present invention .

[0038] FIG . 11 illustrates a result of performing deep learningbased lesion classification on the data volume 240 and 250 , according to an embodiment of the present invention . FIG . 12 illustrates an example of a classified lesion waiting for expert confirmation .

[0039] FIG . 13 illustrates an example of a generated heatmap showing how CNN distinguishes a precursor cervical squamous cell carcinoma precursor .

[0040] FIG . 14 illustrate colposcopic images .

[0041] Detailed description

[0042] The present invention discloses a new method capable of identifying and differentiating gynecologic lesions in images or videos acquired during a colposcopy exam.

[0043] Some preferable embodiments will be described in more detail with reference to the accompanying drawings , in which the embodiments of the present disclosure have been illustrated . However , the present disclosure can be implemented in various manners , and thus should not be construed to be limited to the embodiments disclosed herein .

[0044] It is to be understood that although this disclosure includes a detailed description on cloud computing , implementation of the teachings recited herein are not limited to a cloud computing environment . Rather , embodiments of the present invention are capable of being implemented in conjunction with any other type of computing environment now known or later developed .

[0045] The term "deep learning" is a machine learning technique that uses multiple data processing layers to classify the data sets with high accuracy . It can be a training network (model or device) that learns based on a plurality of inputs and outputs . A deep learning network can be a deployed network (model or device) generated from the training network and provides an output response to an input .

[0046] The term "supervised learning" is a deep learning training method in which the machine is provided with already classified data from human sources . In supervised learning , features are learned via labeled input .

[0047] The term "convolutional neural networks" or "CNNs" are networks that interconnect data used in deep learning to recognize objects and regions in datasets . CNNs evaluate raw data in a series of stages to assess learned features .

[0048] The term "gynecologic lesions" refers to a biologically diverse group of lesions that have varying degrees of malignant potential . "Gynecologic lesions" include but are not limited to congenital , inflammatory, and neoplastic lesions .

[0049] The term "LSIL" refers to gynecologic lesions which cytology revealed as low grade squamous intraepithelial lesion .

[0050] The term "HSIL" refers to gynecologic lesions which cytology revealed as high grade squamous intraepithelial lesion .

[0051] The present invention relates to a method for deep learning based detection of gynecologic lesions in a colposcopy images or video frames (Fig 1 ) . Preferably, the embodiments of the present invention provide a visual output of the deep learning lesions detection method .

[0052] A method is disclosed for gynecologic lesion classification in colposcopy medical images . Said method comprises an image acquisition module ( 1000) , a storage module (2000) , a processing module (3000) , an exam volume data (5000) , a training module (4000) , a second prediction module ( 6000) , and a output collector module (7000) . In the most generic embodiment of the method, image acquisition module 1000 receives exam input volumes from gynecologic colposcopy medical images providers . Images and corresponding labels are loaded onto the storage module 2000 . The storage module 2000 includes a multitude of classification network architectures 100 , trained convolutional network architectures 110 and hyperparameters for training . The storage module 2000 can be a local or cloud computational device . The storage module contains training input labelled data from colposcopy imagery and the required metadata to run processing module 3000 , training module 4000 , exam volume data 5000 , a second prediction module 6000 , and output collector module 7000 . The input labelled data includes , but not only, images and corresponding lesion classification . The metadata includes , but not only, a multitude of classification networks architectures 100 , as exemplified in FIG . 4 , a multitude of trained convolutional neural networks architectures 110 , training hyperparameters , training metrics , fully trained models and selected fully trained models .

[0053] Images 1000 and labelled data are processed at the processing module 3000 before running the optimized training at the training module 4000 . The processing module normalizes the images according to the deep model architecture , to be trained at 3000 or evaluated at 4000 . By manual or scheduled request , the processing module normalizes the image data at the storage module 2000 , according to the deep model architectures that will run at training module 4000 . Additionally, the processing module generates the data pointers to the storage module 2000 to form the partial or full images and ground-truth labels required to run the training module 3000 . To prepare each training session , a dataset is divided into folds , where patient-specific imagery is exclusive to one and one fold only, for training and testing . The training set is split for model training to generate the data pointers of the all images and groundtruth labels , required to run the training process 9000 . K- fold is applied with stratified grouping by patient in the training set to generate the data pointers of the partial images and ground- truth labels , required to run the model verification process 8000 of the training module 4000 . The split ratios and number of folds are available at the metadata of the storage module . Operators include but are not limited to users , a convolutional neural network trained to optimize the k-fold or a mere computational routine . Merely as an example , the dataset is divided with patient split into 90% for training and 10% for testing . Optionally, images selected for training can be split into 80% for training and 20% for validation during training . A 5-fold stratified grouping by patient is applied in the images selected for training . By manual or scheduled request , the processing module normalizes the exam volume data 5000 according to the deep model architecture to run at the prediction module 6000 .

[0054] As seen in Fig 2 , the training module 4000 has a model verification process 8000 , a model selection step 400 and a model training step 9000 . The model verification part iteratively selects combinations of classification architectures 100 and convolutional networks 110 to train a deep model for gynecological lesion classification . The classification network 100 has Dense and Dropout layers to classify gynecological lesions according to their neoplastic potential . A neural convolutional network 110 trained on large datasets is coupled to the said classification network 100 to train a deep model 300 . Partial training images 200 and ground- truth labels 210 train the said deep model 300 . The performance metrics of the trained deep model 120 are calculated using a plurality of partial training images 220 and ground- truth labels 230 . The model selection step 400 is based on the calculated performance metrics , such as f- 1 . The model training part 9000 trains the selected deep model architecture 130 , at process 310 , using the entire data of training images 240 and ground- truth labels 250 . At the prediction module 6000 , the trained deep model 140 outputs gynecologic lesion classification 270 from a given evaluation image 260 . An exam volume of data 5000 comprising the images from the colposcopy imagery is the input of the prediction module 6000 . The prediction module 6000 classifies image volumes of the exam volume 5000 using the best-performed trained deep model from 4000 (see Fig 3) . An output collector module 7000 receives the classified volumes and load them to the storage module after validation by another neural network or any other computational system adapted to perform the validation task .

[0055] Merely as exemplificative , the invention comprises a server containing training results for architectures in which training results from large cloud-based large datasets such as , but not only, ImageNet , ILSVRC , and JFT . The architecture variants include , but are not limited to , VGG, ResNet , Inception , Xception or Mobile , EfficientNets . All data and metadata can be stored in a cloud-based solution or on a local computer . Embodiments of the present invention also provide various approaches to make a faster deep model selection . FIG . 2 illustrates a method for deep learning gynecologic lesion classification according to an embodiment of the present invention . The method of FIG . 2 includes a pretraining stage 8000 , and a training stage 9000 . The training stage 8000 is performed with early stopping on small subsets of data to select the best-performed deep neural network for gynecologic lesion classification among multiple combinations of convolution and classification parts . For example , a classification network of two dense layers of size 512 is coupled with the Xception model to train on a random set resulting from k-fold cross validation with patient grouping . Another random set is selected as the test set .

[0056] The process of training 8000 with early stopping and testing on random subsets is repeated in an optimization loop for combinations of (i ) classification and transfer-learned deep neural networks ; (ii ) training hyperparameters . The image feature extraction component of the deep neural network is any architecture variant without the top layers accessible from the storage module . The layers of the feature extraction component remain frozen but are accessible at the time of training via the mentioned storage module . The BatchNormalization layers of the feature extraction component are unfrozen , so the system efficiently trains with colposcopy imagery presenting distinct features from the cloud images . The classification component has at least two blocks , each having , among others , a Dense layer followed by a Dropout layer . The final block of the classification component has a BatchNormalization layer followed by a Dense layer with the depth size equal to the number of lesions type one wants to classify .

[0057] The fitness of the optimization procedure is computed to (i ) guarantee a minimum accuracy and sensitivity at all classes , defined by a threshold; (ii ) minimize differences between training , validation , and test losses ; (iii ) maximize learning on the last convolutional layer . For example , if a training shows evidence of overfitting , a combination of a shallow model is selected for evaluation .

[0058] The training stage 9000 is applied on the best performed deep neural network using the whole dataset . The fully trained deep model 140 can be deployed onto the prediction module 6000 . Each evaluation image 260 is then classified to output a lesion classification 270 . The output collector module has means of communication with other systems to perform expert validation and confirmation on newly predicted data volumes , reaching 270 . Such means of communication include a display module for user input , a thoroughly trained neural network for decision making or any computational programmable process to execute such task . Validated classifications are loaded on the storage module to become part of the datasets needed to run the pipelines 8000 and 9000 , either by manual or schedule requests .

[0059] An exemplificative embodiment of the classification network 100 , is shown in Fig 5 . Said network , can classify a group of pixels in the image according to the gynecologic lesion nature as LSIL or HSIL . At a given iteration of method 8000 the optimization pipeline described herein uses accuracy curves , ROC curves and AUC values , and confusion matrix from training on a small subset of images and labeled data to assess each architecture' s performance .

[0060] Among the total dataset (n = 22 , 693) , 8 , 729 frames were classified as HSIL , while the remaining 13 , 964 frames were categorized as LSIL . This classification was consistently determined by assessing the respective histopathological report of the biopsy obtained from the observed lesion during the colposcopic procedure . The complete dataset was split into two parts : a training set comprising 20 , 423 frames ( 90% of total dataset) and a testing set with 2270 frames ( 10% of total dataset) . Each frame was exclusively assigned to either the training or testing category . The testing set served as an independent validation for the CNN .

[0061] CNN using a ResNet model has been constructed . The model ' s weights were pre- trained using ImageNet , a comprehensive image dataset designed for object recognition . We kept the initial convolutional layers to transfer its learned features to our model . The last fully connected layers were removed, and new fully connected layers were attached based on the number of classes needed for classifying colposcopy frames . Model ' s architecture includes two blocks , each comprising a fully connected layer followed by a dropout layer of 0 .3 drop rate . Subsequent to these blocks , we incorporated a dense layer whose size was determined by the binary classification (HSIL or LSIL) . The hyperparameters , including the learning rate (0 . 00015) , batch size ( 128) and number of epochs ( 10) , were fine-tuned through a process of trial and error . Data preparation involved using FFMPEG, Pandas and Pillow libraries , while PyTorch was utilized to run the model .

[0062] The model calculated the probability of each frame being classified as HSIL and LSIL . Their final classification was determined based on the one category with the higher probability . Subsequently, CNN' s classification was compared to the corresponding histopathological one , regarded as the gold standard . The primary outcome measures were sensitivity, specificity, accuracy, positive predictive value (PPV) , negative predictive value (NPV) and area under the receiver operating curve (AUC-ROC) . Heatmaps were also generated to enhance our understanding of the specific frame regions contributing the most to the CNN' s prediction (Figure 2 ) . The computational performance was assessed by measuring the time required to process all frames in the testing set . Sci-Kit learn was used for statistical analysis .

[0063] This Al algorithm demonstrated effective performance in identifying and distinguishing between LSIL and HSIL . Additionally, the model was able to maintain its diagnostic accuracy throughout the colposcopy exam, distinguishing lesions in non-stained, acetic acid, lugol-iodine and postmanipulation images. It’s important to note that this is an initial, retrospective, and single-center study, relying on still frames. Acknowledging the necessity for improved robustness and broader applicability, our ongoing goal is to significantly expand the dataset for the CNN and explore additional anatomical areas.

[0064] Said embodiment presented ROC curves and AUC values as shown in figure 8a obtained after training on a small subset of images and labelled data in cervix, where the light gray curve has LSIL - AUC: 0.94 and dashed black line represents the Random Guessing.

[0065] Conversely, Fig 8b illustrates ROC curves and AUC values obtained after training on a small subset of images and labelled data in vagina, where the light gray curve has LSIL - AUC: 0.98) and dashed black line represents the Random Guessing .

[0066] Lastly, Fig 8c illustrates ROC curves and AUC values obtained after training on a small subset of images and labelled data in vulva, where the light gray curve has HSIL - AUC: 1.00 and dashed black line represents the Random Guessing.

[0067] Fig 9a, 9b and 9c illustrates an exemplary confusion matrix after training on a small subset of images and labelled data in cervix, vagina and vulva. Results used for model selection. Number of images of the small subset of data.

[0068] In cervix (9a) A combined total of 22,693 frames were used for development and validation of the CCN, with 8,729 frames classified as HSIL. From the complete dataset, 90% (n= 20,423 frames) were allocated for training the algorithm, while the remaining portion was used for independent validation. For the independent validation, CNN's sensitivity was 99.7% and specificity was 98.6%. PPV and NPV were 97.8% and 99.8%, respectively (9a) . Overall accuracy was 99.0%. The AUC was 0.98. The CNN processing time was 115 frames per second.

[0069] In vagina (9b) A dataset of 13,372 frames was assessed, with 8,574 frames classified as HSIL and the remaining 5,158 frames categorized as LSIL. These classifications were confirmed by analyzing the corresponding histopathological reports of the biopsies obtained during colposcopic procedures . The dataset was divided into two parts : a training set of 12,358 frames (90%) and a testing set of 1,374 frames (10%) . The testing set was used to independently validate the performance of the CNN (9a) .

[0070] In vulva (9C) A total of 965 frames were utilized to develop and validate the CNN. Out of the dataset, 837 frames were classified as HSIL and 128 as LSIL, confirmed by histopathological reports. Among these, 64% (620 frames) were assigned for algorithm training, while the rest (345 frames) were reserved for independent validation using a Resnetl8 model. Initial layers were kept from pre-trained weights on ImageNet, with subsequent adjustments for classifying colposcopy frames. Hyperparameters were finetuned, utilizing PyTorch and libraries like FFMPEG, Pandas, and Pillow. Computational power was supported by NVIDIA Quadro RTX 8000 GPUs and an Intel Xeon Gold 6130 processor. The CNN for the independent validation set demonstrated a sensitivity of 99.7% and specificity of 100.0%. PPV and PV were both 100.0% and 97.9%, respectively (9c) . The overall accuracy reached 99.7%, with an AUC of 1.00. Additionally, the CNN processed 114 frames per second, showcasing its efficiency in handling the dataset.

[0071] Fig 10a, 10b and 10c shows examples of lesion classification, in cervix (10a) , vagina (10b) and vulva (10c) according to an embodiment of the present invention , where in 500 there is low grade squamous intraepithelial lesion and in 510 there is a high grade squamous intraepithelial lesion .

[0072] Fig 11 shows a result of performing deep learning-based lesion classification on the data volume 240 and 250 , according to an embodiment of the present invention . The results of gynecologic colsposcopy classification using the training method 8000 of the present invention are significantly improved as compared to the results using the existing methods (without method 8000) .

[0073] Fig 12 shows an example of a classified lesion waiting for validation by the output collector module 7000 . By another neural network or any other computational system adapted to perform the validation task , physician expert in colposcopy imagery identifies gynecologic lesions , analyzing the labelled image classified by the deep model 140 . Options for image reclassification on the last layer of the classification network 100 are depicted in figure 5 . Optionally, confirmation or reclassification are sent to the storage module .

[0074] Fig 13 shows an example of a generated heatmap showing how CNN distinguishes a precursor cervical squamous cell carcinoma precursor .

[0075] Fig 14 shows Colposcopies performed between December 2022 and January 2023 collected in Centro Materno Infantil do Norte , in Porto , Portugal .

[0076] The foregoing Detailed Description is to be understood as being in every respect illustrative and exemplary, but not restrictive , and the scope of the invention disclosed herein is not to be determined from the Detailed Description , but rather from the claims as interpreted according to the full breadth permitted by the patent laws. It is to be understood that the embodiments shown and described herein are only illustrative of the principles of the present invention and that various modifications may be implemented by those skilled in the art within the scope of the appended claims.

[0077] REFERENCES

[0078] [1] Singh D, Vignat J, Lorenzoni V, et al. Global estimates of incidence and mortality of cervical cancer in 2020: a baseline analysis of the WHO Global Cervical Cancer Elimination Initiative. Lancet Glob Health 2023 ; 11 : el97- e206.

[0079] [2] Burd EM. Human papillomavirus and cervical cancer. Clin Mier. Rev 2003;16:1-17.

[0080] [3] Bray F, Ferlay J, Soer jomataram I, et al. Global cancer statistics 2018: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J Clin 2018;68:394-424.

[0081] [4] Pimple SA, Mishra GA. Global strategies for cervical cancer prevention and screening. Minerva Ginecol 2019;71:313-320.

[0082] [5] Davies-Oliveira JC, Smith MA, Grover S, et al. Eliminating Cervical Cancer: Progress and Challenges for High-income Countries. Clin Oncol (R Coll Radiol) 2021;33:550-559.

[0083] [6] Perkins RB, Wentzensen N, Guido RS, et al. Cervical Cancer Screening: A Review. Jama 2023;330:547-558.

[0084] [7] Vallikad E, Siddartha PT, Kulkarni KA, et al. Intra and Inter-Observer Variability of Transformation Zone Assessment in Colposcopy: A Qualitative and Quantitative Study. J Clin Diagn Res 2017 ; 11 :XcO4-xcO6. [8] Benkortbi K, Catarino R, Wisniak A, et al. Inter- and intra-observer agreement in the assessment of the cervical transformation zone (TZ) by visual inspection with acetic acid (VIA) and its implications for a screen and treat approach: a reliability study. BMC Womens Health 2023;23:27.

[0085] [9] Ren H, Jia M, Zhao S, et al. Factors Correlated with the Accuracy of Colposcopy-Directed Biopsy: A Systematic Review and Meta-Analysis. J Invest Surg 2022;35:284-292.

[0086]

[0010] Ciavattini A, Serri M, Di Giuseppe J, et al. Long-term observational approach in women with histological diagnosis of cervical low-grade squamous intraepithelial lesion: an Italian multicentric retrospective cohort study. BMJ Open 2019 ; 9 : e024920.

[0087]

[0011] Sankaranarayanan R, Gaffikin L, Jacob M, et al. A critical assessment of screening methods for cervical neoplasia. Int J Gynaecol Obstet 2005; 89 Suppl 2:S4-S12.

[0088]

[0012] Bai A, Wang J, Li Q, et al. Assessing colposcopic accuracy for high-grade squamous intraepithelial lesion detection: a retrospective, cohort study. BMC Womens Health 2022 ;22: 9.

[0089]

[0013] Rezniczek GA, Ertan S, Rehman S, et al. Sequential Application of Lugol ’ s Iodine Test after Acetic Acid for Detecting Cervical Dysplasia: A Prospective Cohort Study. Diagnostics (Basel) 2021; 11.

[0090]

[0014] Richards BA, Lillicrap TP, Beaudoin P, et al. A deep learning framework for neuroscience. Nat Neurosci 2019;22:1761-1770.

[0091]

[0015] Khurshid S, Friedman S, Reeder C, et al. ECG-Based Deep Learning and Clinical Risk Factors to Predict Atrial Fibrillation. Circulation 2022;145:122-133.

[0092]

[0016] Islam MM, Poly TN, Walther BA, et al. Deep Learning for the Diagnosis of Esophageal Cancer in Endoscopic Images: A Systematic Review and Meta-Analysis. Cancers (Basel) 2022; 14.

[0017] Wu X, Chen D. Convolutional Neural Network in Microsurgery Treatment of Spontaneous Intracerebral Hemorrhage. Comput Math Methods Med 2022 ; 2022 : 9701702.

[0093]

[0018] Mascarenhas M, Afonso J, Andrade P, et al. Artificial intelligence and capsule endoscopy: unravelling the future. Ann Gastroenterol 2021;34:300-309.

[0094]

[0019] Ferreira JPS , de Mascarenhas Saraiva M, Afonso JPL, et al. Identification of Ulcers and Erosions by the Novel Pillcam Crohn’s Capsule Using a Convolutional Neural Network: A Multicentre Pilot Study. J Crohns Colitis 2022;16:169-172.

[0095]

[0020] Mascarenhas M, Mendes F, Ribeiro T, et al. Deep Learning and Minimally Invasive Endoscopy: Automatic Classification of Pleomorphic Gastric Lesions in Capsule Endoscopy. Clin Transl Gastroenterol 2023 ; 14 : e00609.

[0096]

[0021] Saraiva MM, Ribeiro T, Gonzalez-Haba M, et al. Deep Learning for Automatic Diagnosis and Morphologic Characterization of Malignant Biliary Strictures Using Digital Cholangioscopy: A Multicentric Study. Cancers (Basel) 2023; 15.

[0097]

[0022] Saraiva MM, Spindler L, Fathallah N, et al. Artificial intelligence and high-resolution anoscopy: automatic identification of anal squamous cell carcinoma precursors using a convolutional neural network . Tech Coloproctol 2022;26:893-900.

Claims

AMENDED CLAIMS received by the International Bureau on 10 MAR 2025 (10.03.2025)CLAIMS1 . A computer-implemented method capable of automatically detecting and di f ferentiating LS IL - low grade squamous intraepithelial lesions and HS IL - high grade squamous intraepithelial lesions , in gynecologic colposcopy images , by classi fying the pixels as LS IL or HSIL lesions wherein the method :- accesses previously labeled as LS IL and HS IL gynecologic colposcopy images or video frames and newly acquired gynecologic colposcopy images or video frames ; selects a number of subsets of labeled gynecologic colposcopy data, wherein patient-speci fic data is comprised in one and only one of said subsets ;- equally splits newly acquired gynecologic colposcopy data into data subsets ;- selects another subset as validation set , wherein the subset does not overlap chosen images on the previously selected subsets ;- pre-trains ( 8000 ) of each of the chosen subsets with one of a plurality of combinations of image feature extraction component , followed by a subsequent classi fication neural network component for pixel classi fication as LS IL or HS IL lesions wherein said pre-training :- early stops when the scores did not improved over a given number of epochs , namely three ;- evaluates the performance of each of the combinations ;- is repeated on new, di f ferent subsets , with another networks combination and training hyperparameters , wherein such new combination considers a higher number of dense layers i f the f l-metric is low and fewer dense layers i f f l-metric suggests overfitting;- selects ( 400 ) the architecture combination that performs best during pre-training;- fully trains and validates during training ( 9000 ) the selected architecture combination using the entire set of gynecologic colposcopy images to obtain an optimi zed architecture combination; receives the classi fication output ( 270 ) of the prediction ( 6000 ) of the fully-trained selected architecture by an output collect module with means of communication to a third-party capable of performing validation by interpreting the accuracy of the classi fication output and of correcting a wrong prediction, wherein the third-party comprises at least one of : another neural network, any other computational system adapted to perform the validation task or, optionally, a physician expert in gynecologic colposcopy imagery; storing the corrected prediction into the storage component .2 . The method of claim 1 , wherein the classi fication network architecture comprises at least two blocks , each having a Dense layer followed by a Dropout layer .3 . The method of claims 1 and 2 , wherein the last block of the classi fication component includes a BatchNormali zation layer followed by a Dense layer where the depth si ze is equal to the number of lesions type one desires to classi fy .

4. The method of claim 1, wherein the set of pre-trained neural networks is the best performing among the following: VGG16, IncpetionV3, Xception, Ef f icientNetB5, Ef f icientNetBV , Resnet50 and Resnetl25.

5. The method of claims 1 and 4, wherein the best performing combination is chosen based on the overall accuracy and on the fl-metrics .

6. The method of claims 1 and 4, wherein the training of the best performing combination comprises two to four dense layers in sequence, starting with 4096 and decreasing in half up to 512.

7. The method of claims 1, 4 and 6, wherein between the final two layers of the best performing combination there is a dropout layer of 0.1 drop rate.

8. The method of claim 1, wherein the training of the samples includes a ratio of training-to-validation of 10%-90%.

9. The method of claim 1, wherein the third-party validation is done by user-input.

10. The method of claims 1 and 9, wherein the training dataset includes images in the storage component that were predicted sequentially performing the steps of such method.

Citation Information

Patent Citations

  • Colposcope image classification computer-aided diagnosis system and method based on deep learning

    CN113139944A