Colposcope image focus analysis method, device and system, and storage medium
By constructing a colposcopic image lesion analysis model based on machine learning and deep learning, the problem of insufficient accuracy and repetition of colposcopy in cervical cancer screening is solved, and efficient location and classification of cervical cancer lesions is achieved, and diagnostic accuracy is improved.
Patent Information
- Application Number
- CN202510312603.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-04
AI Technical Summary
The current colposcopy has limited accuracy and repetition in cervical cancer screening, especially in areas with scarce resources that lead to missed diagnosis and misdiagnosis.
Using AI technology based on machine learning and deep learning, a colposcopic image lesion analysis model is constructed, and the colposcopic images are lesion-located, segmented and classified colposcopic images with an anchor frame model, ResNet50 model and Den-SE-MedNeXt model are used to train the convolutional neural network with a large-scale colposcopic cervical cancer dataset.
Accurate lesions positioning, segmentation and classification of colposcopic cervical cancer images are achieved, the accuracy and efficiency of colposcopic diagnosis are improved, and misdiagnosis is reduced.
Smart Images

Figure CN120259208A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical artificial intelligence, and particularly relates to a method and device, a system, and a storage medium for analyzing lesions in colposcopy images. Background Art
[0003] Currently, colposcopy is the main method for cervical cancer screening. For cytological abnormalities or high-risk HPV virus infections, it is necessary to observe under colposcopy for any abnormalities. The accuracy of colposcopy diagnosis will directly affect the quality of cervical cancer screening. However, its accuracy and repeatability are limited. The lack of experienced colposcopy physicians and the heavy workload of colposcopy physicians exacerbate the inaccuracy of colposcopy diagnosis. Especially in resource-scarce areas, due to the high complexity of cervical lesions and the scarcity of professional colposcopy doctors, missed diagnoses and misdiagnoses occur frequently. Therefore, it is necessary to optimize the colposcopy examination technology according to the existing technology.
[0004] In recent years, artificial intelligence has been widely applied in the field of medical images. AI technologies based on machine learning or deep learning can extract relevant information reflecting the lesion grading and prognostic value of patients from medical image data, so as to realize intelligent risk prediction and prognostic determination of diseases. Currently, the research on the cervical cancer screening stage based on artificial intelligence mainly focuses on the analysis of cytological images at the pathological level, and there is no analysis of real-time microscopy images. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a method and device, a system, and a storage medium for analyzing lesions in colposcopy images.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] A method for analyzing lesions in colposcopy images includes:
[0008] Step S1, obtaining a colposcopy image;
[0009] Step S2, obtaining a cervical cancer image lesion analysis model according to the colposcopy image;
[0010] Step S3, inputting the lesion area image of the patient's cervix into the cervical cancer image lesion analysis model to obtain the cervical cancer image lesion localization, segmentation, and classification results.
[0011] Preferably, step S2 includes:
[0012] Preprocessing the colposcopy image, and the preprocessing includes image enhancement and image scaling;
[0013] Label the suspicious lesion area and the lesion area of the preprocessed colposcopy image to obtain a labeled data set;
[0014] Based on the labeled data set and the preprocessed colposcopy image, obtain a cervical cancer image lesion analysis model.
[0015] Preferably, the cervical cancer image lesion analysis model includes: an anchor box model, a ResNet50 model, and a Den-SE-MedNeXt model; among them, the anchor box model is used to locate the lesion position in the colposcopy image based on the anchor box algorithm of Transformer during object detection; the ResNet50 model uses the gradient-weighted class activation mapping method for feature visualization to realize the classification of cervical lesions; the Den-SE-MedNeXt model is used to segment the local image of the lesion area located and cropped by the anchor box model.
[0016] The present invention also provides a colposcopy image lesion analysis device, including:
[0017] An acquisition module for acquiring colposcopy images;
[0018] A training module for obtaining a cervical cancer image lesion analysis model based on the colposcopy image;
[0019] An analysis module for inputting the lesion area image of the patient's cervical disease into the cervical cancer image lesion analysis model to obtain the results of cervical cancer image lesion location, segmentation, and classification.
[0020] Preferably, the training module includes:
[0021] A preprocessing unit for preprocessing the colposcopy image, and the preprocessing includes image enhancement and image scaling;
[0022] A labeling unit for labeling the suspicious lesion area and the lesion area of the preprocessed colposcopy image to obtain a labeled data set;
[0023] A training unit for obtaining a cervical cancer image lesion analysis model based on the labeled data set and the preprocessed colposcopy image.
[0024] Preferably, the cervical cancer image lesion analysis model includes: an anchor box model, a ResNet50 model, and a Den-SE-MedNeXt model; among them, the anchor box model is used to locate the lesion position in the colposcopy image based on the anchor box algorithm of Transformer during object detection; the ResNet50 model uses the gradient-weighted class activation mapping method for feature visualization to realize the classification of cervical lesions; the Den-SE-MedNeXt model is used to segment the local image of the lesion area located and cropped by the anchor box model.
[0025] The present invention also provides a colposcopy image lesion analysis system, comprising: a memory and a processor, wherein a computer program run by the processor is stored on the memory, and the computer program executes a colposcopy image lesion analysis method when run by the processor.
[0026] The present invention also provides a storage medium, on which a computer program is stored, and the computer program executes a colposcopy image lesion analysis method when running.
[0027] The present invention uses a large-scale labeled colposcopy cervical cancer dataset as a training sample to construct a convolutional neural network model, which can accurately classify and segment colposcopy cervical cancer images. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0029] Figure 1 is a flowchart of the colposcopy image lesion analysis method according to the embodiment of the present invention;
[0030] Figure 2 is a schematic diagram of an anchor box model of an anchor box algorithm based on Transformer;
[0031] Figure 3 is a schematic diagram of a cervical cancer lesion classification model based on ResNet50;
[0032] Figure 4 is a schematic diagram of the Den-SE-MedNeXt model. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.
[0034] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0035] Embodiment 1:
[0036] Such asFigure 1 As shown in Figure 1 , an embodiment of the present invention provides a method for analyzing lesions in colposcopy images, including:
[0037] Step S1: Obtain a colposcopy image;
[0038] Step S2: Obtain a cervical cancer image lesion analysis model based on the colposcopy image;
[0039] Step S3: Input the lesion area image of the patient's cervix into the cervical cancer image lesion analysis model to obtain the results of cervical cancer image lesion localization, segmentation, and classification.
[0040] As an implementation manner of the embodiment of the present invention, the colposcopy image is obtained as follows:
[0041] The subject takes the lithotomy position. The vagina is dilated using a vaginal speculum. After fully exposing the cervix, the surface secretions of the cervix and the vaginal secretions are gently wiped with physiological saline. Then, with the assistance of the position of the speculum and a cotton swab, the position and angle of the colposcope are adjusted until a clear cervical image is presented, and image acquisition is performed;
[0042] The examination device uses an electronic colposcope, and the specific parameters are as follows:
[0043] Imager parameters: Pixel count ≥ 440K (GC - 2000B, TC - 3000B) or ≥ 3270K (TC - 8000B); Optical magnification: ≥ 18X (GC - 2000B, TC - 3000B) or ≥ 20X (TC - 8000B); Working distance: 200 - 300mm; Field of view range: Φ6 - 60mm;
[0044] Light source parameters: Illuminance: ≥ 1600lx (working distance 300mm) or ≥ 2000lx (working distance 200mm) (TC - 3000B), ≥ 2500lx (GC - 2000B) or ≥ 6000lx (TC - 8000B); Color temperature: 3200K - 7000K;
[0045] System resolution: ≥ 430TVL (GC - 2000B), ≥ 510TVL (TC - 3000B) or ≥ 800TVL (TC - 8000B);
[0046] Spatial resolution: ≥ 101pm;
[0047] Color reproduction: Deviation not exceeding -5% to +20%;
[0048] Image geometric distortion: ≤ 3%.
[0049] As an implementation manner of the embodiment of the present invention, step S2 includes:
[0050] Step S21: Preprocess the colposcopy image. The preprocessing includes image enhancement and image scaling.
[0051] Step S22: Label the suspicious lesion area and the lesion area on the preprocessed colposcopy image to obtain a label dataset.
[0052] Step S23: Obtain a cervical cancer image lesion analysis model based on the label dataset and the preprocessed colposcopy image.
[0053] Furthermore, the data augmentation of the colposcopy image is as follows: Horizontally and vertically flip / mirror the input image with a probability of 0.5, rotate -10 to +10 degrees, and shear -4 to +4 degrees. The image scaling of the data augmentation of the colposcopy image is as follows: Scale the image from 5184×3456 pixels to 576×384 pixels to adapt to the input requirements of the convolutional neural network and reduce the use of GPU resources.
[0054] As an implementation manner of the embodiment of the present invention, the cervical cancer image lesion analysis model includes: an anchor box model, a ResNet50 model, and a Den-SE-MedNeXt model. Among them, the anchor box model is used to locate the lesion position in the colposcopy image based on the anchor box algorithm of Transformer during object detection. The ResNet50 model uses the gradient-weighted class activation mapping method for feature visualization to realize the classification of cervical lesions. The Den-SE-MedNeXt model is used to segment the local image of the lesion area located and cropped by the anchor box model.
[0055] As an implementation manner of the embodiment of the present invention, as Figure 2 shown, the upper part presents a partial structure of the original Yolov8 model, which is a multi-layer feature pyramid structure of P1-P5 from bottom to top. Each layer has an arrow pointing to "Detect", indicating the detection operation. The lower part shows the improved model structure. Each layer of P1-P5 first undergoes a Conv convolution operation, and among them, P3-P5 also undergoes a replacement operation of the C2f module (replaced with the SwinV2_CSPB module). The improved process includes multiple operation modules, such as Concat (concatenation), Upsample (upsampling), etc. Through the combination and processing of these modules, it is used to improve the resolution of the feature map so that the model can process target information of different scales.
[0056] The SwinV2_CSPB module is based on the anchor box algorithm of Transformer based on the improved Yolov8 model. During the optimization process of the model architecture, the C2f module is replaced with the SwinV2_CSPB module while maintaining the structure of the Conv module unchanged.
[0057] Specifically, the operating mechanism of the SwinV2_CSPB module is as follows: First, the input image is divided into blocks by the convolutional layer Conv2d, and the basic features are extracted from it. Subsequently, the Window-based Multi-Self-Attention (W-MSA) and Shifted Window-based Multi-Self-Attention (SW-MSA) mechanisms of the Transformer operate alternately to effectively capture the global features of the image. Among them, W-MSA focuses on the feature interaction within the local window, while SW-MSA further expands the scope of feature interaction through window offset operations, thereby enhancing the model's perception ability of global information without significantly increasing the computational complexity. Finally, the Patch Merging operation is used to fuse the extracted multi-scale features to generate a more representative feature representation.
[0058] Compared with the traditional Transformer module, the SwinV2_CSPB module innovatively introduces a local attention mechanism. This mechanism enables the model to focus on the features of local regions when processing images, avoiding the huge computational overhead caused by global calculation of the entire image, thus significantly reducing the computational amount.
[0059] 1), Image input: Assume the input colposcopy image is represented as matrix X ∈ R C×H×W , where H, W, and C are the height, width, and number of channels of the image respectively.
[0060] 2), In the convolutional layer of the anchor box model, the input image is operated on using a convolutional kernel. Assume the convolutional kernel of the l-th layer is where k is the convolutional kernel size, C in is the number of input channels, and C out is the number of output channels. The convolutional operation can be expressed as:
[0061]
[0062] where F l (i, j, m) is the feature of the m-th channel at the position (i, j) of the l-th layer of the output feature map, I is the input feature map, K l is the convolutional kernel, b l is the bias, and σ is the Sigmoid activation function.
[0063] 3), Feature Pyramid Network (FPN) and Path Aggregation Network (PANet)
[0064] In the FPN and PANet structures, the anchor box model cleverly utilizes the feature fusion strategy to achieve the effective integration of multi-scale information. Assuming that the feature map is from level l, the fusion process can be described as follows:
[0065]
[0066] Among them, F (l-1) and F (l) represent the feature maps from levels l-1 and l respectively, UpSample represents the upsampling operation, DownSample represents the downsampling operation, and Concat represents the concatenation operation in the channel dimension.
[0067] 4), Anchor box generation
[0068] Use the predefined Anchor Box to predict the position of the target bounding box. Assuming the size of the anchor box is (ω a , h a ), then the predicted bounding box (x, y, ω, h) can be calculated by the following formula:
[0069] x = σ(t x )·w a +x a
[0070] y = σ(t y )·h a +y a
[0071] w = exp(t w )·w a
[0072] h = exp(t h )·h a
[0073] Among them, (t x , t y , t w , t h ) is the output of the network, and (x a , y a ) is the upper left corner coordinates of the grid cell.
[0074] As an implementation manner of the embodiment of the present invention, such as Figure 3As shown, the process of using the ResNet50 network structure for cervical lesion classification is presented. The ResNet50 network consists of multiple stages, and each stage contains several residual blocks. The residual blocks are composed of two or three convolutional layers, and after each convolutional layer, batch normalization and the ReLU activation function are connected. The residual blocks use skip connections to directly add the input to the output, which enables the network to learn the difference between the input and the output. At the beginning of each stage, there is usually a downsampling layer (implemented by a 1x1 convolutional layer with a stride of 2) to reduce the spatial dimension and expand the number of channels. The colposcopy image is input into the network structure of the ResNet50 model, and after being processed through each layer, the classification of cervical lesions is finally achieved, and the classification results include low-grade squamous intraepithelial lesion and high-grade squamous intraepithelial lesion.
[0075] As an implementation manner of the embodiment of the present invention, as Figure 4 shown, the Den-SE-MedNeXt model is used to accurately segment the local image of the lesion area that has been located and cropped by the anchor box model. The specific process is as follows:
[0076] 1), the encoder part;
[0077] The encoder of the Den-SE-MedNeXt model adopts a combination form of dense connection and inverse residual bottleneck unit (MedNeXtBlock). The input image first enters the "Stem,C" module, which serves as the starting part of the encoder, performs preliminary feature extraction and transformation on the input image, and outputs a feature map with the number of channels C.
[0078] Four inverse residual bottleneck units are embedded inside the encoder, and the number of channels doubles in turn. These units are interconnected through a dense network, that is, the feature map generated by each inverse residual bottleneck unit will be transmitted to all subsequent bottleneck units through dense connection. This dense connection method ensures the full flow of information inside the encoder and the efficient reuse of features, enabling the model to accurately capture the subtle features and structural information of the cervical lesion area.
[0079] 2), the decoder part;
[0080] The combination of the scSE module and the inverse residual bottleneck unit is used as the decoder to upsample the feature map. At the same time, a skip connection structure with the scSE module is used to connect the encoder and decoder at each level. This connection method can effectively fuse the feature information of different levels in the encoder into the decoder, providing rich context information for the decoder to restore the high-resolution output segmentation map. The scSE module further enhances the representation ability of the feature map through its spatial and channel attention mechanisms. The decoder gradually converts the high-level features extracted by the encoder into a high-resolution output segmentation map by restoring the spatial resolution of the feature map layer by layer. This design enables the network to capture key features in the image more effectively while suppressing irrelevant information.
[0081] 3) Model running process;
[0082] The colposcopy image is imported into the Den-SE-MedNeXt model. The image first undergoes a series of feature extraction, downsampling, and dense connection operations in the encoder to generate high-level semantic features. These features are then gradually converted into a high-resolution output segmentation map in the decoder through upsampling, the attention mechanism of the scSE module, and skip connections with the encoder, realizing the fine segmentation of the local image of the lesion area located and cropped by the anchor box model.
[0083] 4) In the Den-SE-MedNeXt model, the main formulas and their parameters;
[0084] 1. Convert the input image into the initial feature map H stem
[0085] The input image is denoted as X ∈ R C×H×W , and the convolution kernel size of the initial convolution layer (stem) is 1×1, and the stride is 1×1 (Conv2d).
[0086] H stem = Conv2d(X, 32, kernel_size=(1, 1), stride=(1, 1)).
[0087] 2. MedNeXtDown Block
[0088] O out(0) = O + H stem ;
[0089] Among them, O out(0) represents the input of the next level of the model;
[0090] O out(1) = Conv2d(GELU(GroupNorm(Conv2d(O out(0) , W expansion, k_size, stride
[0091] = 2))), W compression ) + Conv2d(H stem , W residual , stride = 2)
[0092] Among them, GELU is the GELU activation function, W expansion ∈ R CR×C×1×1 is the depth convolution kernel, R is the expansion ratio, W compression ∈ R C×CR×1×1 is the compression convolution kernel; W residual ∈ R C×C×1×1 .
[0093] 3. Dense Block
[0094] For the l-th layer:
[0095] H l = [O0, O1, …, O l-1
[0096] Among them, H l represents the intermediate state of feature map processing;
[0097] The input of the l-th layer generates a new feature map:
[0098] H l+1 = Conv2d(ReLU(BN(H l )))
[0099] Among them, BN normalizes the input feature map to adjust its distribution, ReLU is the ReLU activation function, introducing non-linearity;
[0100] After generating the new feature map, the feature map is concatenated with the input feature map of this layer:
[0101] O l = [O0, O1, …, O l-1 , H l+1
[0102] Finally, the feature maps are concatenated and passed to the next layer:
[0103] O out = [O0, O1, …, O l
[0104] To control the computational complexity of the model, a Transition Layer is introduced:
[0105] O pool = AvgPool2d(O relu , kernel_size = 2, stride = 2)
[0106] Among them, O pool As the output of the Transition Layer and also as the input of the Dense Block and the bottleneck layer, AvgPool2d represents a 2D average pooling operation, O relu Indicates applying the ReLU activation function to the convolutional feature map.
[0107] 4. scSE
[0108] Channel Squeeze & Excitation (cSE): This process reduces the number of parameters and computational complexity through dimensionality reduction, and at the same time introduces a non-linear transformation to enhance the feature expression ability;
[0109]
[0110] S = ReLU(Conv3D(W1Z C ))
[0111] E = σ(W2S)
[0112] O cSE = E⊙O c,i,j,k(bottle) ;
[0113] Spatial Squeeze & Excitation (sSE): Enhances the important spatial information in the feature map and suppresses the unimportant spatial information,
[0114] O sSE = (σ(Conv2d(X, W sSE )))⊙O c,i,j(bottle)
[0115] Finally, add O cSE and O sSE to obtain the feature map O scSE :
[0116] O scSE = O cSE + O sSE
[0117] Adds a transposed convolution operation for upsampling and a residual connection, enabling effective upsampling operation.
[0118] Example 2:
[0119] The embodiment of the present invention also provides a colposcopy image lesion analysis device, including:
[0120] An acquisition module for acquiring colposcopy images;
[0121] A training module for obtaining a cervical cancer image lesion analysis model based on colposcopy images;
[0122] An analysis module for inputting the lesion area image of a patient's cervix into the cervical cancer image lesion analysis model to obtain the results of cervical cancer image lesion localization, segmentation, and classification.
[0123] As an implementation manner of an embodiment of the present invention, the training module includes:
[0124] A preprocessing unit for preprocessing the colposcopy images, where the preprocessing includes image enhancement and image scaling;
[0125] A labeling unit for labeling the suspicious lesion areas and lesion areas of the preprocessed colposcopy images to obtain a label data set;
[0126] A training unit for obtaining a cervical cancer image lesion analysis model based on the label data set and the preprocessed colposcopy images.
[0127] As an implementation manner of an embodiment of the present invention, the cervical cancer image lesion analysis model includes: an anchor box model, a ResNet50 model, and a Den-SE-MedNeXt model; among them, the anchor box model is used to locate the lesion positions in the colposcopy images based on the anchor box algorithm of Transformer during object detection; the ResNet50 model uses the gradient-weighted class activation mapping method for feature visualization to realize the classification of cervical lesions; the Den-SE-MedNeXt model is used to segment the local images of the lesion areas located and cropped by the anchor box model.
[0128] Example 3:
[0129] The embodiment of the present invention also provides a colposcopy image lesion analysis system, including: a memory and a processor, where a computer program is stored on the memory and run by the processor, and the computer program executes the colposcopy image lesion analysis method when run by the processor.
[0130] Example 4:
[0131] The embodiment of the present invention also provides a storage medium, where a computer program is stored on the storage medium, and the computer program executes the colposcopy image lesion analysis method when running.
[0132] The above-described embodiments are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention should fall within the protection scope determined by the claims of the present invention.
Claims
1. A method for analyzing lesions in colposcopy images, characterized in that, Including: Step S1: Obtain a colposcopy image; Step S2: Obtain a cervical cancer image lesion analysis model based on the colposcopy image; Step S3: Input the lesion area image of the patient's cervix into the cervical cancer image lesion analysis model to obtain the cervical cancer image lesion localization, segmentation, and classification results.
2. The colposcopy image lesion analysis method according to claim 1, wherein, Step S2 includes: Preprocess the colposcopy image, and the preprocessing includes image enhancement and image scaling; Annotate the suspicious lesion area and the lesion area on the preprocessed colposcopy image to obtain a label dataset; Obtain a cervical cancer image lesion analysis model based on the label dataset and the preprocessed colposcopy image.
3. The colposcopy image lesion analysis method according to claim 2, characterized in that, The cervical cancer image lesion analysis model includes: an anchor box model, a ResNet50 model, and a Den-SE-MedNeXt model; among them, the anchor box model is used to locate the lesion position in the colposcopy image based on the Transformer-based anchor box algorithm during object detection; the ResNet50 model uses the gradient-weighted class activation mapping method for feature visualization to achieve the classification of cervical lesions; the Den-SE-MedNeXt model is used to segment the local image of the lesion area located and cropped by the anchor box model.
4. A colposcopy image lesion analysis device, characterized in that, Including: An acquisition module for obtaining a colposcopy image; A training module for obtaining a cervical cancer image lesion analysis model based on the colposcopy image; An analysis module for inputting the lesion area image of the patient's cervix into the cervical cancer image lesion analysis model to obtain the cervical cancer image lesion localization, segmentation, and classification results.
5. The colposcope image lesion analysis device according to claim 4, characterized in that, The training module includes: A preprocessing unit for preprocessing the colposcopy image, and the preprocessing includes image enhancement and image scaling; An annotation unit for annotating the suspicious lesion area and the lesion area on the preprocessed colposcopy image to obtain a label dataset; A training unit for obtaining a cervical cancer image lesion analysis model based on the label dataset and the preprocessed colposcopy image.
6. The colposcope image lesion analysis device according to claim 5, wherein, The cervical cancer image lesion analysis model includes: an anchor box model, a ResNet50 model, and a Den-SE-MedNeXt model; among them, the anchor box model is used to locate the lesion position in the colposcopy image based on the Transformer-based anchor box algorithm during object detection; the ResNet50 model uses the gradient-weighted class activation mapping method for feature visualization to achieve the classification of cervical lesions; the Den-SE-MedNeXt model is used to segment the local image of the lesion area located and cropped by the anchor box model.
7. A colposcopy image lesion analysis system, characterized in that, Including: A memory and a processor, where a computer program is stored on the memory and run by the processor, and the computer program, when run by the processor, executes the colposcopy image lesion analysis method according to any one of claims 1-3.
8. A storage medium, characterized in that, A computer program is stored on the storage medium, and the computer program, when running, executes the colposcopy image lesion analysis method according to any one of claims 1-3.