An image sequencing system and method for colposcopy images

CN122821171APending Publication Date: 2026-09-25SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611327658.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-31
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]针对现有技术中的部分或全部问题,为了对同一受检对象数量不固定、顺序不确定的多张宫颈检查图像和存在部分缺失的临床文本进行联合处理,并提升相邻候选类别的判别能力,本发明第一方面提供一种用于阴道镜图像的图像排序系统,包括:

Benefits of technology

[0015]本发明提供的一种用于阴道镜图像的图像排序系统及方法,将语义特征聚合限定于目标区域、并以多方向高频分量补充边界与纹理信息,既能够分析宫颈整体组织表现,又能够重点关注目标区域内部及边界的细粒度变化,从而增强对目标区域整体表型、边缘和细粒度纹理的综合表征能力。同时,所述系统及方法以自适应权重替代固定比例或简单平均,全局语义、局部语义以及多方向高频特征得以按病例实际情况动态调节贡献,并通过图文门控残差结构融合年龄、HPV和细胞学报告等阴道镜图像相应的临床文本信息,使图像特征保持主体地位、临床文本作为受控补充,进而降低了固定比例融合或单一信息来源造成的结果偏差。所述图像排序系统可以接收数量不固定、拍摄顺序不统一、视野范围和图像质量存在差异的宫颈图像,并结合年龄、HPV和细胞学报告等临床信息进行处理,进而减少图像筛选、顺序整理和数据格式统一方面的工作量。其依次完成宫颈主体定位、有效图像筛选、重点筛查区域提取、候选类别生成及关键图像排序,将分散的图像和临床信息转化为可供医生直接复核的病例级结果,可以有效地提高阅图、病例分流和重点病例审核效率。但应当指出,本发明并不涉及疾病的诊断和治疗方法,而是仅仅提供了与医疗相关的信息,所述方法及系统被应用于对阴道镜图像的处理,而不在于疾病的诊断和治疗,而是相应的诊断和治疗应当由医院/医生向用户提供。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821171A_ABST
    Figure CN122821171A_ABST
Patent Text Reader

Abstract

The application discloses an image sorting system and method for colposcope images, which obtains global semantic features and target region local semantic features based on a target region mask through a visual coding module, performs wavelet transform on the colposcope images through a wavelet transform module to obtain high-frequency features in multiple directions, calculates weights through a gate fusion module, fuses the global semantic features, the target region local semantic features and the high-frequency features in each direction to obtain image instance features, converts corresponding clinical text information of the colposcope images into a specific format through a text coding module to obtain text features, fuses the image instance features and the text features through a picture-text fusion module, and obtains candidate category information and corresponding confidence, and the image instance features of multiple groups of single images are weighted and aggregated through a weighted aggregation module to generate a composite image sorting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to an image sorting system and method for colposcopy images. Background Technology

[0002] In actual cervical examinations, the same subject typically corresponds to a variable number of images. For example, some cases may have only a few images, while others may contain dozens. Furthermore, the shooting angles, ranges, and quality of each image vary, and they usually do not have a strict, reliable chronological order. However, image segmentation and target recognition often require a fixed number of images or input in a preset order with a fixed length. This results in some cases not being fully incorporated into the processing flow.

[0003] Furthermore, during actual examinations, information relevant to the final diagnosis, such as age, HPV test results, and cytology reports, may be missing. For example, the absence of HPV information does not necessarily mean a negative HPV result, and the absence of a cytology report does not necessarily mean a normal cytology result. Directly filling in zeros for missing fields or treating them as negative information can easily interfere with subsequent data processing.

[0004] Furthermore, existing cervical image classification models primarily rely on global semantic features for category discrimination, lacking specific representation methods for the boundary orientation and local texture differences within lesion areas. Since adjacent candidate categories such as LSIL and HSIL are highly similar in color and macroscopic morphology, global semantic features alone are insufficient to effectively distinguish these categories, resulting in inadequate discriminative ability of existing models between adjacent candidate categories. Summary of the Invention

[0005] To address some or all of the problems in existing technologies, and in order to jointly process multiple cervical examination images of the same subject with varying numbers and uncertain order, along with partially missing clinical text, and to improve the ability to distinguish adjacent candidate categories, the first aspect of this invention provides an image sorting system for colposcopy images, comprising: The visual encoding module, based on the target region mask, pools only the patches within the target region to extract the global token (Cls Token) and local token (Patch Token) of the colposcopy image, thus obtaining global semantic features and local semantic features of the target region. The wavelet transform module is used to perform wavelet transform on colposcopy images to obtain high-frequency components, and pool each high-frequency component based on the same target region mask to obtain high-frequency features in multiple directions. The Token Gate module is used to calculate the weights of the global semantic features, the local semantic features of the target region, and the high-frequency features in each direction, and then fuse them to obtain image instance features. The text encoding module (HealthGPT) is used to convert the text information corresponding to the colposcopy image into a specific format to obtain text features; The image-text fusion module is used to fuse the image instance features and text features to obtain the category information and corresponding confidence level of the target region in the colposcopy image. The weighted aggregation module (Attention MIL) is used to weight and aggregate the features of multiple single images to generate candidate categories and composite image ranking.

[0006] Furthermore, the visual encoding module includes several stacked Transformer Blocks.

[0007] Furthermore, the image-text fusion module includes a multilayer perceptron, a sigmoid activation function, and a residual fusion submodule.

[0008] Based on the image sorting system described above, a second aspect of the present invention provides an image sorting method for colposcopy images, comprising: For each colposcopy image of the same subject, global semantic features, local semantic features of the target region, and wavelet frequency domain features in multiple directions are extracted based on the target region mask. Gated fusion of global semantic features, local semantic features of the target region, and wavelet frequency domain features from multiple directions is performed to obtain image instance features of each colposcopy image. The clinical text information corresponding to the colposcopy images is converted into a specific format to obtain text features; By fusing the text features and image instance features through a text-image gating mechanism, the category information and corresponding confidence level of the target region of each colposcopy image are obtained; The image instance features of multiple sets of single images are weighted and aggregated to generate a composite image ranking.

[0009] Furthermore, based on the target region mask, global semantic features and local semantic features of the target region are extracted, including: Map the target region mask onto a tile grid; Set the weights of tiles outside the target area to zero, and only pool the tiles within the target area; If the target region mask is empty, it will automatically switch to full-image tile pooling.

[0010] Furthermore, converting the corresponding clinical text information of the colposcopy images into a specific format includes: Convert the clinical text information corresponding to the colposcopy images into standard text; The standard text is converted into high-dimensional text features using a pre-trained text encoding network. The high-dimensional text features are projected into low-dimensional features.

[0011] Furthermore, the fusion of the text features and image instance features through the image-text gating mechanism includes: The text features and image instance features are concatenated to obtain the concatenated features; Text gating weights are generated using a multilayer perceptron and a sigmoid activation function. The image instance features are used as the main body, and the text features are used as controlled supplementary information. The text features and image instance features are fused together using a residual method.

[0012] Furthermore, the image sorting method also includes: For each colposcopy image of the same subject, the effective region is screened to identify the main body of the cervix in order to filter out invalid images.

[0013] Furthermore, the image sorting method also includes: Each colposcopy image of the same subject was classified to remove invalid or inadequately phenotypical images.

[0014] Furthermore, the image sorting method also includes: The filtered image is segmented to obtain the target region mask.

[0015] This invention provides an image sorting system and method for colposcopy images. It aggregates semantic features to a target region and supplements boundary and texture information with multi-directional high-frequency components. This allows for the analysis of the overall cervical tissue appearance while focusing on fine-grained changes within and around the target region, thereby enhancing the comprehensive representation of the overall phenotype, edges, and fine-grained texture of the target region. Simultaneously, the system and method replace fixed proportions or simple averaging with adaptive weights, allowing global semantics, local semantics, and multi-directional high-frequency features to dynamically adjust their contributions according to the actual case. Furthermore, it integrates clinical textual information from colposcopy images, such as age, HPV, and cytology reports, through an image-gated residual structure. This ensures that image features remain dominant while clinical text serves as a controlled supplement, reducing the bias caused by fixed-proportion fusion or single information sources. The image sorting system can receive cervical images of varying numbers, shooting orders, and differences in field of view and image quality, and processes them in conjunction with clinical information such as age, HPV, and cytology reports, thereby reducing the workload of image screening, sorting, and data format standardization. It sequentially completes the main cervical localization, effective image screening, key screening area extraction, candidate category generation, and key image sorting, transforming scattered images and clinical information into case-level results that can be directly reviewed by doctors. This can effectively improve the efficiency of image reading, case triage, and review of key cases. However, it should be noted that this invention does not involve methods for diagnosing and treating diseases, but only provides medically related information. The method and system are applied to the processing of colposcopy images, and are not for the diagnosis and treatment of diseases. The corresponding diagnosis and treatment should be provided to users by hospitals / doctors. Attached Figure Description

[0016] To further illustrate the above and other advantages and features of the various embodiments of the present invention, a more specific description of the various embodiments of the present invention will be presented with reference to the accompanying drawings. It is to be understood that these drawings depict only typical embodiments of the invention and are therefore not intended to limit its scope. In the drawings, identical or corresponding parts will be indicated by identical or similar reference numerals for clarity.

[0017] Figure 1 This diagram illustrates the structure of an image sorting system for colposcopy images according to an embodiment of the present invention. Figure 2 The diagram shows a flowchart of an image sorting method for colposcopy images according to an embodiment of the present invention. Detailed Implementation

[0018] In the following description, the invention is described with reference to various embodiments. However, those skilled in the art will recognize that the embodiments may be practiced without one or more specific details or in conjunction with other alternatives and / or additional methods or components. In other instances, well-known structures or operations are not shown or described in detail so as not to obscure the inventive points of the invention. Similarly, for illustrative purposes, specific numbers and configurations are set forth to provide a comprehensive understanding of embodiments of the invention. However, the invention is not limited to these specific details.

[0019] In this specification, references to "an embodiment" or "this embodiment" mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment of the invention. The phrase "in one embodiment" appearing throughout this specification does not necessarily refer to the same embodiment in all instances.

[0020] It should be noted that this invention does not relate to methods for diagnosing and treating diseases, but merely provides medically relevant information. The methods and systems described are applied to the processing of colposcopic images, and are not intended for the diagnosis and treatment of diseases. Rather, the corresponding diagnoses and treatments should be provided to users by hospitals / doctors.

[0021] To rationally filter and sort a variable number of examination images of the same subject, providing a reliable input source for subsequent diagnosis and improving diagnostic accuracy while reducing workload, this invention provides an image sorting system and method for colposcopy images. This system maps candidate abnormal regions onto a patch grid of visual transforms, aggregating only local semantic features within the lesion region. When the mask is empty, it automatically switches to full-image patch pooling. Simultaneously, wavelet transform is used to extract high-frequency components (LH, HL, HH) to obtain the target region, such as edge and texture features in different directions within the lesion region. Global Cls features, local features of the target region, and high-frequency features are dynamically fused using a token gate and then gated residual fusion with corresponding medical text features such as age, HPV, and cytology reports. Finally, for multiple cervical images of the same subject with varying numbers and uncertain order, a patient-level feature aggregation is performed using weighted aggregation (Attention MIL). This outputs category information and corresponding confidence scores for different candidate target categories, such as normal regions, low-grade lesions (LSIL), high-grade lesions (HSIL), and cancer (Ca). Based on the information contribution of each image, key images are reviewed and ranked to help doctors prioritize the review of images with obvious abnormalities or high information value.

[0022] The technical solution of the present invention will be further described below with reference to the accompanying drawings of the embodiments.

[0023] Figure 1 This diagram illustrates the structure of an image sorting system for colposcopy images according to an embodiment of the present invention. Figure 1 As shown, an image sorting system for colposcopy images includes a visual encoding module 101, a wavelet transform module 102, a token gate module 103, a text encoding module (HealthGPT) 104, an image-text fusion module 105, and a weighted aggregation module (Attention MIL) 106. The visual encoding module 101 extracts global tokens (Cls Token) and local tokens (Patch Token) from the colposcopy image based on a target region mask, obtaining global semantic features and local semantic features of the target region. The wavelet transform module 102 performs wavelet transform on the colposcopy image to obtain high-frequency components, and pools each high-frequency component based on the same target region mask to obtain high-frequency features in multiple directions. The token gate module 103 calculates the weights of the global semantic features, the local semantic features of the target region, and the high-frequency features in each direction, and fuses them to obtain image instance features. The text encoding module 104 converts the corresponding text information of the colposcopy image into a specific format to obtain text features. The image-text fusion module 105 is used to fuse the image instance features and text features, which can be used to determine the category information and corresponding confidence level of the target region. The weighted aggregation module 106 is used to weighted aggregate the image instance features obtained from multiple single images to generate a composite image ranking. Figure 1 As shown, in one embodiment of the present invention, the visual encoding module employs a visual Transformer encoding network, which includes several stacked Transformer Blocks. In one embodiment of the present invention, each colposcopy image is input into the visual Transformer encoding network to extract global Cls tokens and local Patch tokens. A target region mask, such as a candidate abnormal region mask, is mapped to a Patch grid. The Patch weights outside the target region are reset to zero, and pooling is performed only on the Patch within the target region. When the mask is empty, it automatically switches to full-image Patch pooling, thereby taking into account the stable feature representation of both abnormal and normal images. For example, during a cervical examination, the target region includes lesion regions.

[0024] In one embodiment of the present invention, the gated fusion module 103 includes a normalization layer, several fully connected layers, and a GELU activation function. In another embodiment of the present invention, the global semantic features, the local semantic features of the target region, and the high-frequency features in each direction are each denoted as 384-dimensional features. The projection features are obtained by passing the layers through a normalization layer, a fully connected layer from 384 dimensions to 384 dimensions, and a GELU activation function, respectively. After concatenating the five projected features into a 1920-dimensional feature, the feature is sequentially passed through a normalization layer, a fully connected layer from 1920 to 384 dimensions, a GELU activation function, and a fully connected layer from 384 to 5 dimensions to obtain five gated scores. Then according to Softmax normalization is performed to obtain the dynamic weights of each feature. Based on these dynamic weights, the 384-dimensional image instance features of each image can be obtained. .

[0025] In one embodiment of the present invention, the text encoding module 104 includes a pre-trained text encoding network, a normalization layer, a fully connected layer from 4096 dimensions to 384 dimensions, and a GELU activation function. The clinical text information corresponding to the colposcopy image is first used to generate 4096-dimensional text features through the pre-trained text encoding network, and then sequentially passed through the normalization layer, the fully connected layer from 4096 dimensions to 384 dimensions, and the GELU activation function to obtain the text features. .

[0026] In one embodiment of the present invention, the image-text fusion module 105 includes a multilayer perceptron, a sigmoid activation function, and a residual fusion submodule, and image instance features. Text features After concatenation, text-gated weights are generated using a multilayer perceptron and a sigmoid algorithm, and fusion is completed using a residual method, making image instance features the primary element and text features the controlled supplementary information. In one embodiment of the invention, the multilayer perceptron includes a normalization layer, fully connected layers of 768 to 384 dimensions, a GELU activation function, and fully connected layers of 384 to 384 dimensions. The image instance features... Text features The features are concatenated into 768 dimensions, and then processed through the multilayer perceptron and sigmoid activation function to generate 384-dimensional text gating weights with values ​​ranging from 0 to 1. Based on the text gating weights Perform residual fusion: In one embodiment of the invention, the weighted aggregation module 106 includes two fully connected branches, which are activated by Tanh and Sigmoid, respectively. In one embodiment of the invention, image instance features are... After passing through a normalization layer, the input is fed into two fully connected branches. The outputs of the two branches are then multiplied element-wise and passed through a fully connected layer from 128-dimensional to 1-dimensional to obtain the instance score. By normalizing the instance scores of all colposcopy images of the same subject, the instance importance weight of each image can be obtained. Based on the instance importance weight, the image instance features of each image are weighted and aggregated to obtain patient-level image features. The patient-level image features can be fused with text features and thus used as the basis for target region identification and classification. At the same time, the images can be sorted according to the instance importance weight.

[0027] Based on the image sorting system described above Figure 2 This diagram illustrates a flowchart of an image sorting method for colposcopy images according to an embodiment of the present invention. Figure 2 As shown, an image sorting method for colposcopy images includes: First, in step 201, semantic features are extracted. After obtaining the target region mask, each colposcopy image is input into a visual Transformer encoding network to extract global Cls tokens and local Patch tokens, resulting in 384-dimensional global semantic features and 384-dimensional local semantic features of the target region. The candidate abnormal region mask is mapped to a Patch grid, with the Patch weights outside the lesion region set to zero, and pooling performed only on the Patch within the lesion region; when the mask is empty, it automatically switches to full-image Patch pooling, thus taking into account the stable feature representation of both abnormal and normal images. The target region mask is obtained through image segmentation. In one embodiment of the present invention, the colposcopy image is segmented by fusing shallow boundary information and deep lesion semantic information through an encoder-decoder segmentation network to obtain the target region mask.

[0028] Simultaneously, in step 202, high-frequency features are extracted. Wavelet transform is performed on the colposcopy image to extract three high-frequency components: LH, HL, and HH. Each high-frequency component is divided into frequency domain patches corresponding to the visual patches, and pooling is performed using the same candidate abnormal region mask to obtain edge, texture, and local detail features in different directions within the lesion region, where each feature is 384-dimensional.

[0029] Next, in step 203, gated fusion is performed. The weights of various features are dynamically calculated using the Token Gate gating module. The global semantic features, local semantic features of the target region, and three sets of high-frequency features generated for each image are adaptively fused to obtain single-image instance features, avoiding the use of fixed proportions or simple averaging for different features. As mentioned earlier, in one embodiment of the present invention, the global semantic features, local semantic features of the target region, and three sets of high-frequency features are sequentially passed through a normalization layer, a 384-dimensional to 384-dimensional fully connected layer, and a GELU activation function to obtain projection features. These five projection features are concatenated into a 1920-dimensional feature, which is then sequentially passed through a normalization layer, a 1920-dimensional to 384-dimensional fully connected layer, a GELU activation function, and a 384-dimensional to 5-dimensional fully connected layer to obtain five gate scores. Softmax normalization is then performed to obtain the dynamic weights of each feature. Based on these dynamic weights, the 384-dimensional image instance features of each image can be obtained.

[0030] Simultaneously, in step 204, text features are extracted. Clinical text information of the examinee corresponding to the colposcopy image, such as age, HPV, and cytology reports, is converted into a specific format to obtain text features. In one embodiment of the invention, the text information is first converted into a unified standard text, then a high-dimensional text feature (e.g., 4096-dimensional) is generated via a pre-trained medical text coding network. This feature is then sequentially passed through a normalization layer, a fully connected layer from 4096-dimensional to 384-dimensional, and a GELU activation function, projected to a low dimension (e.g., 384-dimensional).

[0031] Next, in step 205, image-text fusion is performed. The text features and image instance features are fused using an image-text gating mechanism to obtain candidate category information and corresponding confidence scores for each colposcopy image. In one embodiment of the invention, the text features and image instance features are first concatenated into 768-dimensional features. Then, a 384-dimensional text gating weight of 0 to 1 is generated using a multilayer perceptron and a sigmoid algorithm. Fusion is then performed using a residual method, with image features as the primary feature and text features as controlled supplementary information. The fused features can be used for target region identification and classification of colposcopy images, and corresponding confidence scores are generated.

[0032] Finally, in step 206, image sorting is performed. The image instance features of multiple sets of single images are weighted and aggregated to generate candidate categories and composite image sorting. In one embodiment of the invention, multiple images of the same subject are used to form an unordered, variable-length instance set. Specifically, the Attention MIL module assigns instance importance weights to each image and weights and aggregates all image features to form patient-level comprehensive features. Specifically, in one embodiment of the invention, the image instance features of each image are... After passing through a normalization layer, the input is fed into two fully connected branches. The outputs of the two branches are then multiplied element-wise and passed through a fully connected layer from 128-dimensional to 1-dimensional to obtain the instance score. Then, the instance scores of all colposcopy images of the same subject are normalized to obtain the instance importance weight of each image: ,in Indicates the first Instance score of the image The temperature parameter, for example, can be 1.0. Then, based on the instance importance weight, the image instance features of each image are weighted and aggregated to obtain patient-level image features: The patient-level image features are ultimately used to output four types of candidate information and their corresponding confidence scores: Normal, LSIL, HSIL, and Ca. The instance importance weight is also used to generate a ranking of key images within the case for medical staff to review first.

[0033] To further reduce computational load, in one embodiment of the invention, colposcopy images are screened before feature extraction. In one embodiment, effective region screening is performed on each colposcopy image of the same subject. For example, complementary modeling of cervical local texture, overall layout, irregular contours, and fine-grained structures is performed using ordinary convolution, dilated convolution, deformable convolution, and depthwise convolution. This is combined with dynamic fusion, coordinate attention, and multi-scale detection to improve the detection capability of small, incompletely displayed, and morphologically diverse cervical regions, reducing interference from background tissue and instruments. In yet another embodiment, each colposcopy image of the same subject is also classified. For example, a collaborative relationship is established between color channels, and the attention intensity is dynamically adjusted based on the current image. Simultaneously, local geometric compensation branches are used to preserve small acetic acid-white regions, color boundaries, and local textures to filter out invalid or insufficiently phenotypic images, reducing the workload of doctors examining low-value images one by one.

[0034] Although various embodiments of the invention have been described above, it should be understood that they are presented by way of example only and not as limitations. It will be apparent to those skilled in the art that various combinations, modifications, and alterations can be made without departing from the spirit and scope of the invention. Therefore, the breadth and scope of the invention disclosed herein should not be limited by the exemplary embodiments disclosed above, but should be defined solely by the appended claims and their equivalents.

Claims

1. An image sorting system for colposcopy images, characterized in that, include: The visual encoding module is configured to pool only the patches within the target region based on the target region mask, extract global and local tokens from the colposcopy image, and obtain global semantic features and local semantic features of the target region. The wavelet transform module is configured to perform wavelet transform on the colposcopy image to obtain high-frequency components, and pool each high-frequency component based on the same target region mask to obtain high-frequency features in multiple directions. The gated fusion module is configured to calculate the weights of the global semantic features, the local semantic features of the target region, and the high-frequency features in each direction, and then fuse them to obtain image instance features. A text encoding module is configured to convert clinical text information corresponding to the colposcopy image into a specific format to obtain text features; The image-text fusion module is configured to fuse the image instance features and text features to obtain the category information and corresponding confidence level of the target region in the colposcopy image. The weighted aggregation module is configured to weight and aggregate the features of multiple single images and generate a composite image ranking.

2. The image sorting system as described in claim 1, characterized in that, The visual encoding module includes several stacked transformer blocks.

3. The image sorting system as described in claim 1, characterized in that, The image-text fusion module includes a multilayer perceptron, a sigmoid activation function, and a residual fusion submodule.

4. An image sorting method for colposcopy images, characterized in that, include: For each colposcopy image of the same subject, global semantic features, local semantic features of the target region, and wavelet frequency domain features in multiple directions are extracted based on the target region mask. Gated fusion of global semantic features, local semantic features of the target region, and wavelet frequency domain features from multiple directions is performed to obtain image instance features of each colposcopy image. The clinical text information corresponding to the colposcopy images is converted into a specific format to obtain text features; The text features and image instance features are fused using a text-image gating mechanism to obtain the category information and corresponding confidence level of the target region in each colposcopy image; The image instance features of multiple sets of single images are weighted and aggregated to generate a composite image ranking.

5. The image sorting method as described in claim 4, characterized in that, Global semantic features and local semantic features of the target region are extracted based on the target region mask, including: Map the target region mask onto a tile grid; Set the weights of tiles outside the target area to zero, and only pool the tiles within the target area; If the target region mask is empty, it will automatically switch to full-image tile pooling.

6. The image sorting method as described in claim 4, characterized in that, Converting clinical text information corresponding to colposcopy images into a specific format includes: Convert the clinical text information corresponding to the colposcopy images into standard text; The standard text is converted into high-dimensional text features using a pre-trained text encoding network. The high-dimensional text features are projected into low-dimensional features.

7. The image sorting method as described in claim 4, characterized in that, The fusion of text features and image instance features through a text-image gating mechanism includes: The text features and image instance features are concatenated to obtain the concatenated features; Text gating weights are generated using a multilayer perceptron and a sigmoid activation function. The image instance features are used as the main body, and the text features are used as controlled supplementary information. The text features and image instance features are fused together using a residual method.

8. The image sorting method as described in claim 4, characterized in that, Also includes: For each colposcopy image of the same subject, the effective region is screened to identify the main body of the cervix in order to filter out invalid images.

9. The image sorting method as described in claim 4, characterized in that, Also includes: Each colposcopy image of the same subject was classified to remove invalid or inadequately phenotypical images.

10. The image sorting method as described in claim 8 or 9, characterized in that, Also includes: The filtered image is segmented to obtain the target region mask.