A colposcopy image segmentation system and method

CN122820749APending Publication Date: 2026-09-25SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611308134.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-27
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0002]在阴道镜图像的计算机辅助诊断中,现有技术通常采用单一或固定方式的层间融合,易丢失宫颈组织的颜色边界、微小纹理变化及细小血管信息等空间细节,导致对边界模糊病灶以及面积较小的高级别病变区域的像素级定位精度不足,分割边界勾勒粗糙,难以同时兼顾范围较大的连续病变区域与细粒度局部病灶,这就使得后续处理过程中对少数类别及细小病灶的识别能力受限,存在高风险病变漏检或误检的固有局限

Benefits of technology

[0014]本发明提供的一种阴道镜图像的分割系统及方法,对阴道镜图像中的宫颈区域进行多尺度特征提取与融合,实现像素级目标区域定位和类型识别,同时输出不同类别区域的精细分割结果。此外,通过多尺度特征融合、类别感知损失优化及难例增强策略,可以提高对特殊目标区域,如小面积目标区域、稀缺样本目标区域等的识别能力。同时,将分割结果进一步转化为结构化表型信息,可以融合HPV、TCT、病理及临床信息的多模态智能诊断模型提供关键病灶先验,实现从病灶发现、精准分型到辅助诊疗决策的智能化闭环。但应当指出,本发明并不涉及疾病的诊断和治疗方法,而是仅仅提供了与医疗相关的信息,所述方法及系统被应用于对阴道镜图像的处理,而不在于疾病的诊断和治疗,而是相应的诊断和治疗应当由医院/医生向用户提供。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820749A_ABST
    Figure CN122820749A_ABST
Patent Text Reader

Abstract

The application discloses a colposcope image segmentation system and method, which comprises an encoding unit, a decoding unit and a classification and identification unit. The encoding unit comprises N feature extraction modules with stepwise down-sampling to obtain feature maps of input colposcope images. The shallow feature extraction modules extract color, target edge, texture and structural features. The deep feature extraction modules extract spatial relationship features between target regions and surrounding tissues. The decoding unit comprises N semantic feature acquisition modules with stepwise up-sampling and one-to-one skip connection with the feature extraction modules, which are used to restore the spatial resolution of the feature maps, obtain semantic features, and fuse the image features and the semantic features of the corresponding levels. The classification and identification unit is connected to the last-level semantic feature acquisition module, which outputs a pixel-level segmentation result corresponding to the size of the input colposcope image based on the fused features. The segmentation system and method can realize pixel-level target positioning and type identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a system and method for segmenting colposcopy images. Background Technology

[0002] In computer-aided diagnosis of colposcopy images, existing technologies typically employ single or fixed interlayer fusion methods, which easily lose spatial details such as color boundaries, subtle texture changes, and small blood vessel information of cervical tissue. This results in insufficient pixel-level localization accuracy for lesions with blurred boundaries and small high-grade lesion areas, and coarse segmentation boundaries. It is difficult to simultaneously take into account large continuous lesion areas and fine-grained local lesions. This limits the ability to identify a few types and small lesions in subsequent processing, and has inherent limitations in the detection of high-risk lesions, such as missed or false detection.

[0003] Furthermore, existing technologies mostly remain at the level of outputting pixel-level segmentation masks or single overall categories, lacking a mechanism to further transform segmentation masks into structured medical phenotypic information such as lesion location, area, boundary morphology, number of lesions, and confidence levels of different lesion categories. This makes it difficult to directly use them as structured inputs for joint analysis with HPV, TCT, pathological results, and other clinical information, and thus cannot provide prior information on lesions for cervical lesion grading, risk assessment, and treatment decisions. Summary of the Invention

[0004] To address some or all of the problems in the existing technology, and in order to achieve pixel-level target localization and type recognition, the first aspect of this invention provides a segmentation system for colposcopy images, comprising: The encoding unit includes N progressively downsampled feature extraction modules to obtain feature maps of the input colposcopy image. The shallow feature extraction module is used to extract color, target edge, texture, and structural features, while the deep feature extraction module is used to extract spatial relationship features between the target region and surrounding tissues. The decoding unit includes N semantic feature acquisition modules that are upsampled step by step. The decoding unit is used to restore the spatial resolution of the feature map and acquire semantic features. Each semantic feature acquisition module is connected to the feature extraction module in a one-to-one skip connection so as to fuse the image features and semantic features of the corresponding level. The classification and recognition unit is connected to the last-level semantic feature acquisition module and is used to output pixel-level segmentation results based on the fused feature output and the size of the input colposcopy image.

[0005] Furthermore, each feature extraction module and semantic feature acquisition module includes two convolutional blocks, each of which includes a convolutional layer (Conv 3×3), a normalization layer (InstanceNorm), and a non-linear activation layer (LeakyReLU).

[0006] Furthermore, the step size (Stride) of the downsampling of the encoding unit and the upsampling of the decoding unit is 2.

[0007] Furthermore, the segmentation system also includes: An image preprocessing unit is used to preprocess the colposcopy image to reduce device differences, lighting variations, reflections, occlusions, and background tissue interference.

[0008] Furthermore, the segmentation system also includes: The post-processing unit, connected to the classification and recognition unit, is used to perform post-processing on the pixel-level segmentation results, including connected component analysis, small-area false positive region filtering, hole filling, boundary smoothing, and merging of adjacent similar regions. It records the spatial location, contour range, area, perimeter, and category probability of each target region, and determines the boundary and category of each target region based on the pixel-level category probability and region continuity.

[0009] Furthermore, the number of layers, the initial number of channels, and the kernel size of the encoding and decoding units are determined based on the training data during the training phase.

[0010] Furthermore, the segmentation system uses expert annotation information as supervision information during the training phase, and the loss calculation includes cross-entropy loss and / or Dice loss, with different weights set for different target regions.

[0011] Based on the segmentation system described above, a second aspect of the present invention provides a method for segmenting colposcopy images, comprising: The input image is downsampled step by step, and color, target edge, texture, structural features, and spatial relationship features between the target region and surrounding tissues are extracted layer by layer to obtain feature maps; The feature map is upsampled level by level to obtain semantic features, which are then fused with the image features of the corresponding level. The output of the final-level fusion features corresponds to the pixel-level segmentation result of the input colposcopy image size.

[0012] Furthermore, the segmentation method also includes: Before downsampling, the input image is preprocessed to reduce interference from device differences, lighting variations, reflections, occlusions, and background structures.

[0013] Furthermore, the segmentation method also includes: The pixel-level segmentation results are subjected to connected component analysis, small-area false positive region filtering, hole filling, boundary smoothing, and merging of adjacent similar regions. Record the spatial location, outline range, area, perimeter, and category probability of each target region; The boundaries and categories of each target region are determined based on pixel-level category probabilities and regional continuity.

[0014] This invention provides a segmentation system and method for colposcopy images. It extracts and fuses multi-scale features of the cervical region in colposcopy images, achieving pixel-level target region localization and type recognition, while outputting fine segmentation results for different categories of regions. Furthermore, through multi-scale feature fusion, category-aware loss optimization, and hard-case enhancement strategies, it can improve the recognition ability for special target regions, such as small-area target regions and scarce sample target regions. Simultaneously, the segmentation results are further transformed into structured phenotypic information, enabling a multimodal intelligent diagnostic model that integrates HPV, TCT, pathological, and clinical information to provide prior knowledge of key lesions, achieving an intelligent closed loop from lesion discovery and accurate classification to assisted diagnostic and treatment decisions. However, it should be noted that this invention does not involve methods for diagnosing and treating diseases, but only provides medically relevant information. The method and system are applied to the processing of colposcopy images, not for the diagnosis and treatment of diseases; the corresponding diagnosis and treatment should be provided to users by hospitals / doctors. Attached Figure Description

[0015] To further illustrate the above and other advantages and features of the various embodiments of the present invention, a more specific description of the various embodiments of the present invention will be presented with reference to the accompanying drawings. It is to be understood that these drawings depict only typical embodiments of the invention and are therefore not intended to limit its scope. In the drawings, identical or corresponding parts will be indicated by identical or similar reference numerals for clarity.

[0016] Figure 1 A schematic diagram of a colposcopy image segmentation system according to an embodiment of the present invention is shown. Figure 2 This diagram illustrates the structure of a convolutional block in a feature extraction module according to an embodiment of the present invention. Figure 3 This diagram illustrates a flowchart of a method for segmenting colposcopy images according to an embodiment of the present invention. Detailed Implementation

[0017] In the following description, the invention is described with reference to various embodiments. However, those skilled in the art will recognize that the embodiments may be practiced without one or more specific details or in conjunction with other alternatives and / or additional methods or components. In other instances, well-known structures or operations are not shown or described in detail so as not to obscure the inventive points of the invention. Similarly, for illustrative purposes, specific numbers and configurations are set forth to provide a comprehensive understanding of embodiments of the invention. However, the invention is not limited to these specific details.

[0018] In this specification, references to "an embodiment" or "this embodiment" mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment of the invention. The phrase "in one embodiment" appearing throughout this specification does not necessarily refer to the same embodiment in all instances.

[0019] It should be noted that this invention does not relate to methods for diagnosing and treating diseases, but merely provides medically relevant information. The methods and systems described are applied to the processing of colposcopic images, and are not intended for the diagnosis and treatment of diseases. Rather, the corresponding diagnoses and treatments should be provided to users by hospitals / doctors.

[0020] To achieve pixel-level lesion localization and type recognition, and to simultaneously output fine segmentation results for different categories of target regions, including normal tissue, low-grade intraepithelial lesions (LSIL), high-grade intraepithelial lesions (HSIL), and cancer, this invention discloses a segmentation system and method for colposcopy images. It employs an encoder-decoder structure based on a convolutional neural network for semantic segmentation of cervical images. The encoder consists of multiple progressively downsampled feature extraction units. Each feature extraction unit performs hierarchical representation of the input image through convolution, normalization, and nonlinear activation operations, gradually expanding the receptive field as the network depth increases. The decoder uses progressive upsampling to restore the spatial resolution of the feature map and fuses the high-resolution features of the encoder's corresponding layer with the high-semantic features of the decoder through skip connections. The fused features are then processed by convolution to gradually form pixel-level classification results corresponding to the original image size. The segmentation categories include at least background or lesion-free areas, low-grade squamous intraepithelial lesion areas, high-grade squamous intraepithelial lesion areas, and cervical cancer areas. Meanwhile, addressing issues such as blurred lesion boundaries, scarcity of small-area high-grade lesion samples, and class imbalance in colposcopy images, this study improves the ability to identify micro-lesions and high-risk lesion areas through multi-scale feature fusion, class-aware loss optimization, and hard-case enhancement strategies. Furthermore, the segmentation results can be further transformed into structured medical phenotypic information such as lesion location, area, boundary morphology, number, and class confidence, providing key lesion priors for subsequent multimodal intelligent diagnostic models that integrate HPV, TCT, pathological, and clinical information. This achieves an intelligent closed loop from lesion detection and accurate classification to assisted diagnostic and treatment decisions.

[0021] The technical solution of the present invention will be further described below with reference to the accompanying drawings of the embodiments.

[0022] Figure 1 This diagram illustrates the structure of a colposcopy image segmentation system according to an embodiment of the present invention. Figure 1As shown, a segmentation system for colposcopy images includes an encoding unit 101, a decoding unit 102, and a classification and recognition unit 103. The encoding unit 101 is used to acquire a feature map of the input colposcopy image. The decoding unit 102 is used to restore the spatial resolution of the feature map, acquire semantic features, and fuse them with image features of the corresponding level to obtain fused features. The classification and recognition unit 103 performs image segmentation based on the fused features to obtain a probability map or category map of the target region.

[0023] like Figure 1 As shown, in one embodiment of the present invention, the encoding unit 101 is composed of multiple feature extraction units 111 that are progressively downsampled. Each feature extraction unit 111 performs hierarchical representation of the input image through convolution, normalization, and nonlinear activation operations. The shallow encoding features mainly preserve the color changes of cervical tissue, lesion edges, fine-grained texture, and local vascular structures. As the network depth increases, the encoder gradually expands its receptive field to extract the range of acetic acid whitening reaction, abnormal epithelial distribution, coarse blood vessel morphology, overall lesion outline, and the spatial relationship between the lesion area and surrounding normal tissue. Through multi-scale encoding, it is possible to simultaneously perceive small, poorly defined local lesions and large, continuous lesion areas.

[0024] In one embodiment of the present invention, the encoding unit 101 comprises five levels of feature extraction units, from level 0 to level 4. Each level of feature extraction unit includes two convolutional blocks, with the number of channels increasing progressively layer by layer. For example, f0 is 32, f1 is 64, f2 is 128, f3 is 256, and f4 is 512 or 320. The stride for each downsampling step is 2. The structure of the convolutional block is as follows: Figure 2 As shown, it includes a convolutional layer (Conv 3×3), a normalization layer (InstanceNorm), and a nonlinear activation layer (LeakyReLU), which are connected in series. In one embodiment of the invention, Dropout regularization terms may optionally be introduced between layers.

[0025] Correspondingly, such as Figure 1As shown, in one embodiment of the present invention, the decoding unit 102 is composed of multiple semantic feature acquisition units 121 that perform progressive upsampling. The decoding unit 102 uses a progressive upsampling method to restore the spatial resolution of the feature map, and fuses the high-resolution features of the corresponding level of the encoding unit with the high-semantic features of the decoding unit through skip connections. Skip connections enable the system to preserve color boundaries, subtle texture changes, and fine blood vessel information in shallow features while restoring the spatial location of the lesion region, reducing the loss of spatial details caused by continuous downsampling. In one embodiment of the present invention, the high-resolution features of the corresponding level of the encoding unit and the high-semantic features of the decoding unit are fused through concatenation.

[0026] In one embodiment of the present invention, the decoding unit 102 comprises five levels of semantic feature acquisition units (levels 4' to 0'), corresponding to feature extraction unit levels 4 to 0 respectively. Each level of semantic feature acquisition unit also includes two convolutional blocks, with the number of channels decreasing layer by layer, consistent with the corresponding level of feature extraction unit. It upsamples using transpose convolution, with a stride of 2 for each upsampling operation.

[0027] In one embodiment of the present invention, the classification and recognition unit 103 performs convolution processing to form a pixel-level segmentation result corresponding to the size of the input colposcopy image. The pixel-level segmentation result includes at least a background or lesion-free area, a low-grade squamous intraepithelial lesion area, a high-grade squamous intraepithelial lesion area, and a cervical cancer area. In another embodiment of the present invention, the classification and recognition unit 103 can output a probability map corresponding to each category and determine the category with the highest probability as the predicted category of the pixel, thereby obtaining a multi-category lesion segmentation mask.

[0028] In one embodiment of the present invention, the segmentation system is preferably implemented using the nnU-Net framework. Based on the image size, spatial resolution, number of categories, and memory conditions of the training data, the system adaptively determines the input image size, network layers, convolutional kernel size, number of feature channels, and training parameters, thereby improving the model's adaptability to colposcopy images from different sources, on different devices, and under different imaging conditions. During the training phase, lesion regions and lesion types labeled by clinical experts are used as supervisory information. To address the problem that background regions are often significantly more numerous than lesion regions and that the number of samples from different lesion categories is unbalanced, cross-entropy loss, Dice loss, or a combination of both can be used to allow the model to simultaneously focus on pixel classification accuracy and the overall overlap of lesion regions. For smaller high-grade lesions or cancer regions, category weights or hard case weights can also be set to enhance the model's ability to identify a few categories and small lesions.

[0029] To reduce interference from device differences, lighting variations, reflections, occlusion, and background tissue, in one embodiment of the present invention, the segmentation system further includes an image preprocessing unit 104, which preprocesses the acquired colposcopy image, and the preprocessed image serves as the input to the encoding unit 101. The preprocessing may include common image processing techniques such as grayscale normalization, histogram matching, color calibration, noise normalization, camera bias correction, gamma correction, color space transformation, filtering, and noise reduction.

[0030] In one embodiment of the present invention, the segmentation system is further provided with a post-processing unit 105, which is connected to the classification and recognition unit 103, for post-processing the pixel-level segmentation results, including connected component analysis, small-area false positive region filtering, hole filling, boundary smoothing and merging of adjacent similar regions, recording the spatial location, contour range, area, perimeter and category probability of each target region, and determining the boundary and category of each target region based on the pixel-level category probability and region continuity, etc.

[0031] Based on the segmentation system described above Figure 3 This diagram illustrates a flowchart of a method for segmenting colposcopy images according to an embodiment of the present invention. Figure 3 As shown, a method for segmenting colposcopy images includes: First, in step 301, image features are extracted. The input image is downsampled layer by layer, and color, target edges, texture, structural features, and spatial relationship features between the target area and surrounding tissues are extracted layer by layer to obtain a feature map. In one embodiment of the present invention, before the downsampling, step 300, image preprocessing, is performed to preprocess the input image to reduce interference from device differences, lighting changes, reflections, occlusions, and background tissues. Next, in step 302, feature fusion is performed. The feature map is upsampled level by level to obtain semantic features, which are then fused with the image features of the corresponding level. Finally, in step 303, image segmentation is performed. Based on the final-level fusion features, a pixel-level segmentation result corresponding to the size of the input colposcopy image is output.

[0032] In one embodiment of the present invention, after segmentation is completed, step 304, post-processing of the segmentation results, is performed. Post-processing of the initial segmentation results includes connected component analysis, small-area false positive region filtering, hole filling, boundary smoothing, and merging of adjacent regions of the same type. For cases where multiple independent lesions exist in the same image, the spatial location, contour range, area, perimeter, and category probability of each lesion are recorded. For cases where different lesion categories are adjacent or overlapping, the boundaries and categories of each lesion region are determined based on pixel-level category probabilities and regional continuity. Finally, the lesion segmentation mask of the cervical image and the type information corresponding to each lesion are output, and the lesion center coordinates, area proportion, boundary morphology, number of lesions, and confidence levels of different lesion categories can be further calculated. These results can be used to delineate the lesion range on the original colposcopy image, and can also serve as structured input for subsequent multimodal fusion diagnostic models, performing joint analysis with HPV, TCT, pathological results, and other clinical information to provide prior information on lesions for cervical lesion grading, risk assessment, and treatment decisions.

[0033] Although various embodiments of the invention have been described above, it should be understood that they are presented by way of example only and not as limitations. It will be apparent to those skilled in the art that various combinations, modifications, and alterations can be made without departing from the spirit and scope of the invention. Therefore, the breadth and scope of the invention disclosed herein should not be limited by the exemplary embodiments disclosed above, but should be defined solely by the appended claims and their equivalents.

Claims

1. A segmentation system for colposcopy images, characterized in that, include: The encoding unit includes N progressively downsampled feature extraction modules. The encoding unit is configured to acquire feature maps of the input colposcopy image. The shallow feature extraction module is configured to extract color, target edge, texture, and structural features, and the deep feature extraction module is configured to extract spatial relationship features between the target region and surrounding tissues. N is a natural number. The decoding unit includes N semantic feature acquisition modules that are upsampled step by step. The decoding unit is configured to restore the spatial resolution of the feature map and acquire semantic features. Each semantic feature acquisition module is connected to the feature extraction module in a one-to-one skip connection so as to fuse the image features and semantic features of the corresponding level. The classification and recognition unit is connected to the last-level semantic feature acquisition module and is configured to perform pixel-level segmentation based on the fused feature output and the size of the input colposcopy image.

2. The segmentation system as described in claim 1, characterized in that, Each feature extraction module and semantic feature acquisition module includes two convolutional blocks, each of which includes a convolutional layer, a normalization layer, and a non-linear activation layer.

3. The segmentation system as described in claim 1, characterized in that, The number of layers, initial number of channels, and convolution kernel size of the encoding and decoding units are adaptively determined based on the image size, spatial resolution, number of categories, and memory conditions of the training data during the training phase.

4. The segmentation system as described in claim 1, characterized in that, During the training phase, expert-annotated information is used as supervision information. The loss calculation includes cross-entropy loss and / or Dice loss, and different weights are set for different target regions.

5. The segmentation system as described in claim 1, characterized in that, The step size for downsampling in the encoding unit and upsampling in the decoding unit is 2.

6. The segmentation system as described in claim 1, characterized in that, Also includes: An image preprocessing unit is configured to preprocess the colposcopy image to reduce device differences, lighting variations, reflections, occlusions, and background tissue interference.

7. The segmentation system as described in claim 1, characterized in that, Also includes: The post-processing unit is connected to the classification and recognition unit and is configured to perform post-processing on the pixel-level segmentation results, including connected component analysis, small-area false positive region filtering, hole filling, boundary smoothing and merging of adjacent similar regions, recording the spatial location, contour range, area, perimeter and category probability of each target region, and determining the boundary and category of each target region based on the pixel-level category probability and region continuity.

8. A method for segmenting colposcopy images, characterized in that, include: The input image is downsampled step by step, and color, target edge, texture, structural features, and spatial relationship features between the target region and surrounding tissues are extracted layer by layer to obtain feature maps; The feature map is upsampled level by level to obtain semantic features, which are then fused with the image features of the corresponding level. The output of the final-level fusion features corresponds to the pixel-level segmentation result of the input colposcopy image size.

9. The segmentation method as described in claim 8, characterized in that, Also includes: Before downsampling, the input image is preprocessed to reduce interference from device differences, lighting variations, reflections, occlusions, and background structures.

10. The segmentation method as described in claim 8, characterized in that, Also includes: The pixel-level segmentation results are subjected to connected component analysis, small-area false positive region filtering, hole filling, boundary smoothing, and merging of adjacent similar regions. Record the spatial location, outline range, area, perimeter, and category probability of each target region; The boundaries and categories of each target region are determined based on pixel-level category probabilities and regional continuity.