Cervical cell image intelligent diagnosis system based on multi-modal visual language large model
Through the intelligent diagnosis system of cervical cell images based on a large multimodal visual language model, the problems of low accuracy, strong subjectivity and insufficient efficiency in cervical cell image diagnosis have been solved, and efficient and explainable automatic diagnosis of cervical cell abnormalities has been achieved, thereby improving the accuracy and efficiency of early screening of cervical cancer.
Patent Information
- Application Number
- CN202511177937.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-10-10
AI Technical Summary
Existing technologies in cervical cell image diagnosis have problems such as low diagnostic accuracy, strong subjectivity, insufficient efficiency and lack of interpretability. In particular, there is a lack of processing methods for high-resolution cervical pathology WSI images, making it difficult to achieve accurate cellular-level diagnosis and natural language interpretation of diagnostic results.
An intelligent diagnosis system for cervical cell images based on a multimodal visual language large model is adopted, including modules such as image block acquisition, preprocessing, effective tissue area screening, cell area extraction and post-processing, image and text end processing, visual language large model training, cell description generation in the reasoning stage, risk judgment and classification, explainability verification and credibility assessment, and structured diagnosis report output. Combined with specific fine-tuning strategies, efficient and automated diagnosis can be achieved.
It has greatly improved the automated and accurate diagnosis and cell-level natural language description of cervical cell abnormalities, increased screening efficiency and the consistency and interpretability of diagnostic results, reduced diagnostic costs, and promoted the popularization and clinical application of early screening for cervical cancer.
Smart Images

Figure CN120766940A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an intelligent diagnosis system for cervical cell images based on a multimodal visual language large model. Background Art
[0002] Cervical cancer is one of the most common malignant tumors in women worldwide, with over 600,000 new cases and over 300,000 deaths each year, severely impacting women's health and quality of life. Cervical cancer often lacks obvious symptoms in its early stages and progresses rapidly, making early detection and effective intervention crucial. Early screening and diagnosis, particularly cervical cytology (Pap smear) and human papillomavirus (HPV) testing, can effectively reduce the morbidity and mortality of cervical cancer.
[0003] Currently, cervical cytology and HPV testing are widely used in clinical screening. However, both methods rely heavily on manual interpretation and diagnostic judgment by pathologists, which are subject to significant subjective factors. This leads to inconsistent diagnostic results between different physicians and significant human variability in the diagnostic process. The traditional manual interpretation process is not only labor-intensive and prone to fatigue for pathologists, affecting diagnostic accuracy and consistency, but also significantly limits screening efficiency, making it difficult to meet large-scale clinical screening needs.
[0004] To overcome these issues, computer-aided diagnosis (CAD) has been widely explored and gradually applied to cervical cytology. Early CAD techniques often employed traditional machine learning algorithms, such as support vector machines (SVMs) and random forests, and typically relied on manually designed feature engineering to extract image features. While these methods have some diagnostic capabilities, they are limited by the subjective and incomplete nature of feature design and are unable to effectively describe the complex and diverse cellular morphological changes.
[0005] In recent years, with the development of deep learning technology, especially the widespread application of convolutional neural networks (CNN), CAD methods based on CNN have gradually become dominant, showing diagnostic performance superior to traditional methods. However, these deep learning models are usually designed for small-scale images or local cell features. When faced with panoramic cervical pathology images (Whole Slide Image, WSI), due to the extremely high image resolution and the large number of cells involved, traditional small-scale models are difficult to capture both the overall field of view and the fine cell-level lesion features. In addition, due to the black box nature of the model itself, its diagnostic process lacks the necessary explainability, making it difficult for physicians to understand the basis for the model's judgment, which limits the trust and widespread application of deep learning models in clinical practice.
[0006] At the same time, the rise of Visual-Language Models (VLMs) has provided a new approach to solving the above problems. Through deep learning architecture, Visual-Language Models effectively realize multimodal knowledge transfer between images and texts, and can demonstrate outstanding capabilities in image understanding and natural language generation. For example, visual language models such as BLIP, DAM, LLaVA, and GPT-4V proposed in recent years have performed well in tasks such as image description, anomaly detection, and medical report generation. However, the application of these visual language models in the medical field, especially for high-resolution cervical pathology WSI images, still faces huge challenges. WSI images usually have ultra-high resolutions of billions of pixels, data annotations are extremely scarce, and cervical cancer cells often show subtle and irregular morphological changes, which places extremely high demands on the detection and description of cell abnormalities.
[0007] Current medical diagnostic technologies based on large visual language models still have significant shortcomings, including a lack of effective processing methods for high-resolution images, a lack of fine-tuning strategies specifically tailored to medical image features, and difficulties in achieving accurate cellular-level diagnosis and providing natural language interpretation of diagnostic results. Furthermore, the efficient integration of deep learning methods into the diagnostic process in medical practice to ensure efficient and interpretable diagnostic results remains a critical issue that needs to be addressed. Summary of the Invention
[0008] The purpose of the present invention is to solve the problems of low diagnostic accuracy, strong subjectivity, insufficient efficiency and lack of interpretability in the existing technology, and to propose an intelligent diagnosis system for cervical cell images based on a large multimodal visual language model.
[0009] An intelligent diagnosis system for cervical cell images based on a multimodal visual language large model includes:
[0010] Image block acquisition module, image preprocessing and color normalization module, effective tissue area screening module, cell area extraction and post-processing module, image and text processing module, visual language large model acquisition and training module, cell description text generation module in the reasoning stage, risk judgment and classification module, explainability verification and credibility assessment module, and structured diagnostic report output module;
[0011] The image block acquisition module is used to acquire a whole slice image WSI, and divide the acquired whole slice image WSI into blocks to obtain image blocks;
[0012] The image preprocessing and color normalization module is used to convert the image block from RGB space to optical density space, construct the dye matrix of the image based on the optical density space, obtain the concentration matrix based on the dye matrix of the image, obtain the reconstruction matrix based on the concentration matrix and the target dye matrix, and convert the reconstruction matrix back to RGB space;
[0013] The effective tissue area screening module is used to screen each image block converted back to the RGB space obtained by the image preprocessing and color normalization module to obtain a valid image block;
[0014] The cell region extraction and post-processing module is used to extract the cell region from the effective image block obtained by the effective tissue region screening module, and perform post-processing on the extracted cell region to obtain the segmented cell region image;
[0015] The image-side and text-side processing modules are used to perform image-side and text-side processing on the segmented cell region images obtained by the cell region extraction and post-processing modules;
[0016] The visual language model acquisition and training module is used to acquire the visual language model and train the acquired visual language model to obtain a trained visual language model.
[0017] The cell description text generation module in the inference phase is used to input a single cell image to be tested and a task prompt into the trained visual language model, and obtain the cell description text based on the beam search strategy;
[0018] The risk assessment and classification module is used to perform risk assessment and classification on the cell description text obtained by the cell description text generation module in the reasoning phase;
[0019] The explainability verification and credibility assessment module is used for interactive question-answering verification and credibility determination;
[0020] The structured diagnostic report output module is used to aggregate the full-slice analysis results and generate structured reports and recommendations.
[0021] The beneficial effects of the present invention are:
[0022] Based on the above problems, the present invention proposes an innovative method to use a large visual language model combined with a specific fine-tuning strategy to process high-resolution cervical pathology WSI, realize the automated and accurate diagnosis of cervical cell abnormalities and cell-level natural language description, greatly improve the screening efficiency and the consistency and interpretability of the diagnostic results, ultimately reduce the diagnostic cost, and promote the popularization and clinical application of early screening for cervical cancer.
[0023] The present invention aims to provide a high-precision, automated, and interpretable cervical cell screening and diagnosis system, which mainly solves the following technical problems:
[0024] (1) Insufficient diagnostic accuracy: Currently, manual image reading or traditional machine learning models are limited by the method itself and often have difficulty accurately identifying and classifying cervical cell lesions. In particular, the ability to capture subtle features of early lesions is significantly insufficient, resulting in a certain proportion of missed diagnoses or misdiagnoses. This invention effectively utilizes a large visual language model to perform detailed cellular descriptions and accurately identify lesion features, significantly improving diagnostic accuracy.
[0025] (2) Subjective diagnosis: Traditional diagnostic methods rely heavily on the pathologist's personal experience, resulting in a lack of consistency and standardization in diagnostic results. This paper proposes an automated description generation mechanism that achieves objectivity and consistency in the diagnostic process by generating cell-level natural language descriptions and objectively expressing abnormal features.
[0026] (3) Low diagnostic efficiency: Manual image reading is a huge workload and easily causes physician fatigue. Traditional computer-aided diagnosis methods are unable to effectively process large-scale, high-resolution WSI image data, which seriously restricts diagnostic efficiency. This paper proposes an automated, scalable, end-to-end diagnostic process that can effectively process high-resolution image data and achieve fast, efficient, and automated diagnosis.
[0027] (4) Insufficient diagnostic interpretability: Currently, widely used deep learning models are often black-box structures, making it difficult to provide clear diagnostic evidence and explanations, which reduces the trust of doctors and patients in the results. This invention is based on a large visual language model and combines a multi-round question-and-answer verification mechanism to significantly improve the interpretability of the diagnostic process, clarify the basis for lesion identification, and increase the credibility of clinical practice.
[0028] By solving the above-mentioned technical problems, the present invention can significantly improve the accuracy, efficiency and credibility of cervical cancer screening, and promote the development of medical diagnostic technology towards automation, intelligence and explainability.
[0029] By deeply integrating a large visual language model with digital pathology workflows, this invention successfully implements a cervical cell intelligent diagnosis system that combines high precision, strong interpretability, full process automation, and high scalability. Compared with traditional computer-aided diagnosis (CAD) methods, this invention has the following four significant technical advantages:
[0030] 1. Diagnostic accuracy and sensitivity reach new heights, achieving precise identification of cellular-level microscopic lesions. This method breaks through the limitations of traditional models that rely only on coarse features or patch-level classification. First, by using advanced instance segmentation models such as Hover-Net or SAM in S4, accurate contour extraction of individual cell pixels is achieved, ensuring the purity and accuracy of feature analysis from the source. Secondly, the unique multi-task learning framework in S6 not only Optimize diagnostic accuracy and introduce text generation loss and cell attribute regression loss , forcing the model to deeply understand and quantify subcellular morphological features that are closely related to diagnosis (such as nuclear membrane morphology, chromatin texture, nuclear-cytoplasmic ratio, etc.). This refined feature learning and multi-task collaborative optimization mechanism effectively improves the model's ability to capture subtle morphological differences such as early lesions and atypical cells. Experimental data show that the area under the receiver operating characteristic curve (AUC) of this method in the normal / abnormal binary classification task can stably exceed 0.95, and the accuracy rate in multi-classification tasks involving normal, ASC-US, LSIL, HSIL, etc. exceeds 89%, demonstrating its clinical application potential as an efficient screening and auxiliary diagnostic tool.
[0031] 2. Give the diagnostic process unprecedented explainability and credibility, and build an AI-physician trust bridge. In response to the pain points of black-box decision-making of traditional deep learning models, the present invention has constructed a complete diagnostic logic traceability path. The core lies in: First, the model in S7 can generate a natural language description that conforms to pathological standards for each suspicious cell, and convert complex image features into text evidence that can be understood and evaluated by human experts, so that what you see can be described. Second, S9's original multi-round question-and-answer (QA) verification module further explores and verifies whether the model's diagnostic basis is reliable through automated and progressive questioning of suspicious cells. Finally, the system generates a diagnostic credibility index (ConfidenceIndex, ), providing doctors with a quantitative reference for the certainty of model judgments. This dual mechanism of description and verification ensures that every diagnostic conclusion of the model is verifiable and justifiable, greatly enhancing clinicians' trust in AI results and promoting the transformation of AI's role from a tool to a partner.
[0032] 3. Realize the automation of the entire process from sample to report, and significantly improve the efficiency and throughput of pathological diagnosis. The present invention integrates all steps from S1 to S10 into a seamless automated workflow, realizing unmanned intervention from WSI original image loading, color standardization, tissue area screening, cell segmentation, intelligent analysis to structured report output. This can liberate pathologists from a large amount of repetitive and mechanical film reading work, especially the screening of tens of thousands of normal cells. By automatically generating heat maps of diseased cells and reports with pre-filled diagnostic elements, the system enables physicians to focus their precious time and energy on the review and final diagnosis of difficult and positive cases marked by the system. This fully automated process significantly shortens the average reading time of a single slice, providing a technical possibility for large-scale, high-throughput cervical cancer screening, and is particularly suitable for primary medical institutions, regional pathology diagnosis centers and third-party AI-assisted diagnosis service platforms.
[0033] 4. Adopting modular and standardized design, it has excellent clinical deployment flexibility and system expansion capabilities, such as Figure 5 and Figure 6 As shown. The complexity of the real clinical environment was fully considered at the beginning of the design of this system. First, at the data access end (S1), by adopting OpenSlide, it natively supports multiple mainstream WSI file formats, ensuring wide compatibility with scanner equipment of different brands. Secondly, the system architecture itself has a highly modular feature, which facilitates the independent upgrade and maintenance of any functional module (such as segmentation model, language model) in the future. Furthermore, as shown in the flow chart, the system supports Docker container packaging and Kubernetes (K8s) cluster orchestration, which can achieve fast and standardized private deployment or cloud deployment, and can elastically scale computing resources according to business volume. Finally, the standardized report output (JSON / CSV / Word) in S10 and support for medical information standards such as DICOM and HL7FHIR ensure that this system can be easily integrated with the hospital's existing PACS, LIS / HIS systems, and realize the interconnection and interoperability of data flow and workflow.
[0034] In summary, this invention not only surpasses existing technologies in diagnostic performance but, more importantly, reshapes the paradigm of human-machine collaborative pathology diagnosis by innovatively introducing the interpretable interactive capabilities of a large visual language model. This brings a highly accurate, efficient, and reliable AI pathology expert assistant into clinical practice, with profound implications for addressing uneven distribution of medical resources, improving screening coverage, and ensuring diagnostic quality. It demonstrates broad clinical application prospects and significant industrial value. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 It is the overall framework diagram of the method of the present invention; Figure 2To achieve the process and processing step diagram; Figure 3 This is the architecture diagram of the visual language model; Figure 4a Fine-tuning the architecture diagram for LoRa; Figure 4b is the attention mechanism diagram; Figure 4c It is a multimodal fusion graph; Figure 4d This is a diagram of the multi-task loss function architecture; Figure 5 This is the architecture diagram of the key modules; Figure 6 Develop a flow chart for clinical deployment. DETAILED DESCRIPTION
[0036] Specific implementation method 1: Combination Figure 1 、 Figure 2 This embodiment describes an intelligent diagnosis system for cervical cell images based on a multimodal visual language model, including:
[0037] Image block acquisition module, image preprocessing and color normalization module, effective tissue area screening module, cell area extraction and post-processing module, image and text processing module, visual language large model acquisition and training module, cell description text generation module in the reasoning stage, risk judgment and classification module, explainability verification and credibility assessment module, and structured diagnostic report output module;
[0038] The image block acquisition module is used to acquire a high-resolution whole-slice image WSI, and divide the acquired high-resolution whole-slice image WSI into blocks to obtain image blocks;
[0039] The image preprocessing and color normalization module is used to convert the image block from RGB space to optical density space, construct the dye matrix of the image based on the optical density space, obtain the concentration matrix based on the dye matrix of the image, obtain the reconstruction matrix based on the concentration matrix and the target dye matrix, and convert the reconstruction matrix back to RGB space;
[0040] The effective tissue area screening module is used to screen each image block converted back to the RGB space obtained by the image preprocessing and color normalization module to obtain a valid image block;
[0041] The cell region extraction and post-processing module is used to extract the cell region from the effective image block obtained by the effective tissue region screening module, and perform post-processing on the extracted cell region to obtain the segmented cell region image;
[0042] The image-side and text-side processing modules are used to perform image-side and text-side processing on the segmented cell region images obtained by the cell region extraction and post-processing modules;
[0043] The visual language model acquisition and training module is used to acquire the visual language model and train the acquired visual language model to obtain a trained visual language model.
[0044] The cell description text generation module in the inference phase is used to input a single cell image to be tested and a task prompt (for example, please describe the abnormal characteristics of the cell) into the trained visual language model, and obtain the cell description text based on the beam search strategy;
[0045] The risk determination and classification module is used to perform risk determination and classification on the cell description text obtained by the cell description text generation module in the reasoning phase;
[0046] The explainability verification and credibility assessment module is used for interactive question-answering verification and credibility determination;
[0047] The structured diagnostic report output module is used to aggregate the full-slice analysis results and generate structured reports and recommendations.
[0048] Specific embodiment 2: This embodiment differs from specific embodiment 1 in that the image block acquisition module is used to acquire a high-resolution whole-slice image WSI, and divide the acquired high-resolution whole-slice image WSI into blocks to obtain image blocks; the specific process is as follows:
[0049] S11. Use the OpenSlide open source library to obtain high-resolution whole-slide images (WSIs) and metadata of the whole-slide images (WSIs).
[0050] Whole-slide image (WSI) metadata includes: physical size (micrometers / pixel), scanning magnification (e.g., 20x, 40x), and resolution at each level in the image pyramid. These metadata are essential for subsequent accurate quantitative analysis (e.g., calculation of cell nuclear area) and for ensuring consistency in multi-scale processing.
[0051] The OpenSlide open source library is used to ensure broad compatibility with WSI formats generated by various mainstream digital pathology scanners (such as Aperio SVS, Leica SCN, Hamamatsu NDPI, and standard TIFF).
[0052] S12, dividing the whole slice image WSI obtained in S11 into blocks (for example, each image block Patch is 512×512 or 1024×1024 pixels);
[0053] Each image patch contains the following information:
[0054] The spatial coordinates of the whole slice image WSI obtained by each image block Patch in S11 , the level to which each image patch belongs in the whole slice image WSI obtained in S11 and the size of each image patch, ensuring that spatial context information is not lost, which provides the necessary conditions for subsequent lesion area positioning, heat map generation and visual tracing of diagnostic results;
[0055] Spatial coordinates of whole-slice image WSI Refers to the position of the upper left corner of the image patch in the WSI global pixel coordinate system (in pixels, with the origin at the upper left corner of the WSI);
[0056] The process is expressed as:
[0057]
[0058] in, represents an image block, represents the whole slice image WSI acquired by S11, Indicates the highest resolution level, express The coordinates of the upper left corner, express The width, express height.
[0059] Since the native resolution of a single WSI is extremely high (for example, 100,000 × 80,000 pixels at 40x magnification), it cannot be fully loaded at once.
[0060] Other steps and parameters are the same as those in the first embodiment.
[0061] Specific embodiment three: This embodiment differs from specific embodiment one or two in that: the image preprocessing and color normalization module is used to convert the image block from RGB space to optical density space, construct the dye matrix of the image based on the optical density space, obtain the concentration matrix based on the dye matrix of the image, obtain the reconstruction matrix based on the concentration matrix and the target dye matrix, and convert the reconstruction matrix back to RGB space; the specific process is as follows:
[0062] S21. To eliminate color inconsistencies caused by differences in different medical institutions, different batches of stains (such as hematoxylin-eosin, H&E staining), and scanning equipment, this method uses the color normalization algorithm proposed by Macenko et al. The process is: each image block obtained in S1 is converted from RGB color space to optical density (OD) space. This conversion is more consistent with the physical principle of dye absorption of light (Beer-Lambert law). The conversion formula is:
[0063]
[0064] in, represents the pixel intensity value of the image block obtained by S1 (range 0-255), Indicates the reference intensity value of the background white light (usually 255). Adding 1 to the numerator is to avoid the logarithm value being 0.
[0065] S22. Settings Threshold value;
[0066] In S21, select greater than Threshold Vector, based on the selected greater than in S21 Threshold Vector constructs the dye matrix of the image (the selected vector in S21 is greater than Threshold The vector is used as the element of the dye matrix of the image. The matrix is a matrix with 3 rows and 2 columns, and 1 column in the matrix is 1 Vector, 2 columns for hematoxylin and eosin, 3 rows for 3 channels); perform color deconvolution on each element in the dye matrix of the image, decompose the color of each element into the concentration value of hematoxylin (mainly marking the cell nucleus) and eosin (mainly marking the cytoplasm and matrix), and obtain a concentration matrix; multiply the concentration matrix with the set target dye matrix to obtain a reconstruction matrix, and convert the reconstruction matrix back to RGB space;
[0067] When converted back to RGB space, we get a bunch of normalized patches.
[0068] Because if the subsequent task requires the entire image (such as visualizing a heat map), these patches will be pieced back together.
[0069] This process ensures that all images input into the deep learning model have high consistency in color representation, thereby significantly improving the robustness and generalization ability of the model.
[0070] The effective tissue area screening module is used to screen each image block converted back to the RGB space obtained by the image preprocessing and color normalization module to obtain a valid image block; the specific process is as follows:
[0071] The Otsu algorithm is used to process each image block converted back to the RGB space obtained by the image preprocessing and color normalization module to obtain a binary image of each image block. The proportion of foreground pixels in the binary image of each image block is calculated, and only valid image blocks whose proportion of foreground pixels in the binary image of each image block exceeds a preset threshold (for example, 20%) are retained.
[0072] In order to greatly improve the processing efficiency and avoid the interference of blank background areas on model training, all patches need to be screened for effectiveness. This method uses Otsu's method, which is an adaptive threshold segmentation algorithm that can automatically find an optimal threshold on the grayscale histogram of the image. , this threshold can maximize the variance between the foreground (tissue) and background (blank) categories. Its core objective function is to maximize the inter-class variance :
[0073]
[0074] in, Represents the pixel ratio of each category, Represents the average grayscale value of each category. By applying the Otsu algorithm to each patch to obtain a binary image, the proportion of foreground pixels is calculated, and only those valid patches whose proportion exceeds a preset threshold (for example, 20%) are retained.
[0075] Other steps and parameters are the same as those in the first or second embodiment.
[0076] Specific embodiment 4: This embodiment differs from any one of specific embodiments 1 to 3 in that: the cell region extraction and post-processing module is used to extract the cell region from the effective image block obtained by the effective tissue region screening module, and post-process the extracted cell region to obtain a segmented cell region image; the specific process is as follows:
[0077] S41, inputting each valid image block obtained by the valid tissue region screening module into the Hover-Net network model, and the Hover-Net network model outputting a cell region image of each valid image block;
[0078] Each valid image block obtained by the valid tissue area screening module is input into the Cellpose network model, and the Cellpose network model outputs the cell area image of each valid image block;
[0079] Each valid image block obtained by the valid tissue area screening module is input into the Segment Anything Model network model, and the Segment Anything Model network model outputs the cell area image of each valid image block;
[0080] For each valid image block selected, this method uses an advanced instance segmentation model to complete the most critical task of cell localization and extraction. The goal of this step is not only to detect cells but also to accurately outline the pixel-level boundaries of each individual cell.
[0081] Hover-Net: A network designed specifically for nucleus segmentation and separation, which can effectively separate densely adherent nuclei by simultaneously predicting a nucleus pixel map and a horizontal / vertical distance map from each nucleus pixel to its centroid.
[0082] Cellpose: A deep learning model that predicts a gradient flow field representing the geometric center of objects, and reconstructs the complete contour of individual cells by tracking this flow field, showing excellent segmentation performance for cells with highly irregular shapes.
[0083] Segment Anything Model (SAM): A powerful, promptable base segmentation model that can segment any object with zero or few-shot samples, with strong generalization ability and flexibility. The model finally generates a set of segmentation masks for each patch, each mask corresponding to an identified independent cell instance.
[0084] S42, morphological filtering is performed on the cell region image of each valid image block output by the network model to obtain a morphologically filtered image; attribute filtering is performed on the morphologically filtered image to obtain a filtered image; the process is as follows:
[0085] An opening operation is performed on the cell region image of each valid image block output by the network model, and a closing operation is performed on each image after the opening operation to obtain a morphologically filtered image;
[0086] The geometric properties of each morphologically filtered image are calculated, and images that do not meet the geometric property threshold are removed to obtain a filtered image;
[0087] The geometric properties are area, perimeter, circularity and compactness;
[0088] In order to improve the quality and reliability of the segmentation results, a series of automatic post-processing operations are performed on the initially generated masks:
[0089] Morphological Filtering: Opening operation (first erosion, then dilation) is applied to remove isolated small noise points and fine edge burrs, and then closing operation (first dilation, then erosion) is applied to fill small holes in the cell interior, making the cell morphology smoother and more complete.
[0090] Property Filtering: Based on prior pathological knowledge, geometric properties such as area, perimeter, circularity, and solidity are calculated for each segmented object. By setting appropriate attribute thresholds, objects that are too small (possibly cell debris or artifacts), too large (possibly unsegmented cell clumps), or have extremely abnormal shapes can be effectively eliminated.
[0091] The other steps and parameters are the same as those in the first to third embodiments.
[0092] Specific embodiment 5: This embodiment differs from any one of specific embodiments 1 to 4 in that: the image-side and text-side processing modules are used to perform image-side and text-side processing on the segmented cell region image obtained by the cell region extraction and post-processing module; the specific process is as follows:
[0093] Each segmented cell region image obtained by the cell region extraction and post-processing module is uniformly scaled to the standard input size required by the Visual Language Model (VLM) (e.g., 224×224 pixels);
[0094] Construct a text prompt (Prompt) for each segmented cell region image obtained by the cell region extraction and post-processing module;
[0095] The text prompts are:
[0096] Comprehensive descriptive prompt: Please describe this image of cervical cells in detail from the perspective of a pathologist. Your analysis should include the size and shape of the nucleus, the thickness and uniformity of the chromatin distribution, and the nuclear-cytoplasmic ratio. Finally, based on the size and shape of the nucleus, the thickness and uniformity of the chromatin distribution, and the nuclear-cytoplasmic ratio, please provide a preliminary judgment on whether the cervical cells show signs of pathology.
[0097] Specific feature query type prompt: whether the nuclear membrane of cervical cells is smooth and regular;
[0098] Direct classification task prompt: Please classify cervical cells into one of the following options according to The Bethesda System (TBS) classification criteria: normal squamous cells, atypical squamous cells (ASC-US), low-grade squamous intraepithelial lesion (LSIL), or high-grade squamous intraepithelial lesion (HSIL).
[0099] By building a large number of (images, prompts) The training sample pairs of expected answers lay the foundation for subsequent model fine-tuning;
[0100] (1) Image side: According to the accurate segmentation mask optimized by S4, the minimum circumscribed rectangle region of each single cell is cropped from the corresponding Patch image. In order to retain the local context information around the cell, the cropping frame can be moderately expanded outward (for example, expand the margin by 20%). All cropped single cell images will be uniformly scaled to the standard input size (for example, 224x224 pixels) required by the visual language large model (VLM). During the model training phase, a series of data augmentation operations such as random rotation, flipping, color jittering and Gaussian blur will also be applied to these images to expand the data set and improve the generalization performance of the model.
[0101] (2) Text side: For each cell image, one or more carefully designed text prompts are paired. These prompts are key instructions that guide the model to think and answer specifically.
[0102] Prompt engineering is one of the cores of this method, which decomposes complex medical diagnosis tasks into natural language instructions that the model can understand and execute. The designed prompt templates cover multiple levels from general observation to specific diagnosis.
[0103] The other steps and parameters are the same as those in the first to fourth embodiments.
[0104] Embodiment six: This embodiment is different from one of the first to fifth embodiments in that the visual language large model acquisition and training module is used to acquire a visual language large model and train the acquired visual language large model to obtain a trained visual language large model; the specific process is as follows:
[0105] S61, fine-tune the language large model (VLM) using LoRA technology to obtain a visual language large model; the specific process is as follows:
[0106] The forward propagation process of the language large model at a specific layer becomes:
[0107]
[0108] wherein, is a low-rank matrix, ; is a low-rank matrix, ; is a Token vector input to a certain layer of the language large model; is a scaling constant; is a weight matrix of the language large model; is a rank, much smaller than the dimension and (for example, r=16); The Token vector dimension of the input of a certain layer; The dimension of the Token vector output by a certain layer; is the output vector after adding LoRA low-rank update;
[0109] In this way, large models can be efficiently adapted to the specific task of cervical cytology diagnosis with very few trainable parameters (typically <1%).
[0110] S62. In order to enable the large visual language model to master multiple capabilities such as description, classification and quantification at the same time and promote each other, a multi-task joint loss function for the large visual language model is constructed. (See Figure 4d ):
[0111]
[0112] in, Represents text generation loss, which is cross entropy loss; express The weight hyperparameters of represents the classification loss; express The weight hyperparameters of Represents regression loss, which is the mean square error loss; express The weight hyperparameters of
[0113] Text Generation Loss : Use standard cross-entropy loss to penalize the difference between the text generated by the model and the standard medical description, guiding the model to learn to generate fluent and accurate language.
[0114]
[0115] Classification loss : When the data comes with clear cell category labels, add a classification head and use cross-entropy loss for supervision to directly optimize the classification accuracy of the model.
[0116]
[0117] in, Indicates the classification results output by the visual language model, represents the truth value;
[0118] Regression loss : Guide the model to learn to predict some key, quantifiable pathological indicators (such as nuclear area and nuclear-cytoplasmic ratio), usually using mean squared error (MSE) loss. This task forces the model to focus on lower-level visual features that are directly relevant to diagnosis.
[0119]
[0120] in, Represents the output prediction value of the visual language model, represents the truth value;
[0121] Each loss term is determined by the weight hyperparameter Balance is achieved and through joint optimization, the model becomes a versatile diagnostic assistant.
[0122] This method chooses a powerful pre-trained Vision-Language Model (VLM) as its basis, such as BLIP-2 or LLaVA. The general architecture of these models is as follows Figure 3 As shown in the figure, it usually contains three core components: a visual encoder (such as Vision Transformer, ViT), which is responsible for converting the input cell image into a series of high-dimensional feature vectors; a large language model (LLM, such as LLaMA or GPT series), which is responsible for understanding text prompts and generating natural language descriptions; and a cross-modal fusion module (such as Q-Former), which serves as a bridge to connect the two modalities of vision and language;
[0123] Directly fine-tuning the entire VLM with billions of parameters requires huge computational resources and is prone to catastrophic forgetting. To solve this problem, this method uses Low-Rank Adaptation (LoRA) technology for efficient fine-tuning (see Figure 4a ). The core idea of LoRA is to freeze the original weight matrix of the pre-trained model , and next to it is connected a matrix consisting of two low-rank matrices and During training, only these two small matrices are updated and Parameters.
[0124] The key to the model's image-text understanding lies in the cross-modal attention mechanism (Cross-modal Attention), such as Figure 4b and Figure 4c As shown in Figure 2. In this mechanism, the feature vector from the textual prompt serves as the query (Query, Q), and attention is paid to the sequence of feature vectors (Key, K and Value, V) representing different regions of the image generated by the visual encoder. By calculating the similarity between the query and all keys to assign weights to the values, the model can dynamically incorporate the visual information most relevant to the current text, thereby performing accurate joint image-text reasoning. This calculation follows the standard scaled dot product attention formula.
[0125] Other steps and parameters are the same as those in Specific Implementations 1 to 5-1.
[0126] Specific embodiment 7: This embodiment differs from any one of specific embodiments 1 to 6 in that the cell description text generation module in the inference stage is used to input a single cell image to be tested and a task prompt (for example, please describe the abnormal characteristics of the cell) into a trained visual language model, and obtain the cell description text based on the beam search strategy; the specific process is as follows:
[0127] A single-cell image to be tested and a task prompt (for example, please describe the abnormal characteristics of the cell) are input into the trained visual language model. The trained visual language model will retain k candidate sequences with the highest probability (k is the beam width) at each step, and predict the output of the next step based on the retained k candidate sequences with the highest probability (k is the beam width). Finally, the trained visual language model outputs the cell description text with the best overall probability.
[0128] After fine-tuning, the model enters the inference phase. Here, the model is fed with any single-cell image to be examined and a prompt (e.g., "Describe the abnormal characteristics of this cell"). The model uses an autoregressive approach, generating descriptive text token by token. To ensure the quality, coherence, and medical expertise of the generated descriptions, rather than simply selecting the token with the highest probability at each step (a greedy search), this method employs a beam search strategy. At each step, beam search retains the k (k is the beam width) candidate sequences with the highest probability, expanding them further to ultimately select the complete sequence with the highest overall probability. This effectively avoids generating incoherent or factually incorrect descriptions, producing more stable and high-quality text. For example, the following example demonstrates: "This cell is an atypical squamous cell, exhibiting a slightly increased nuclear-cytoplasmic ratio and slightly hyperchromatic nuclei, but with a smooth nuclear membrane and no clear features of a koilocyte."
[0129] The other steps and parameters are the same as those in the first to sixth embodiments.
[0130] Specific embodiment eight: This embodiment differs from any one of specific embodiments one to seven in that: the risk determination and classification module is used to perform risk determination and classification on the cell description text obtained by the cell description text generation module in the reasoning phase; the specific process is as follows:
[0131] S81. Construct a knowledge base of medical keywords and phrases, the knowledge base of medical keywords and phrases including abnormal indicators such as nuclear atypia, nuclear hyperchromasia, rough chromatin, koilocytes, and cornified beads, and normal indicators such as clear cell borders and normal nuclear-cytoplasmic ratio;
[0132] Performing regular expression processing on the cell description text obtained by the cell description text generation module in the inference phase to obtain the processed cell description text;
[0133] Matching the processed cell description text in the medical keyword and phrase knowledge base;
[0134] When the processed cell description text matches the normal indicator words in the medical keyword and phrase knowledge base, the processed cell description text is normal;
[0135] When the processed cell description text matches one abnormal indicator word in the medical keyword and phrase knowledge base, the processed cell description text is low risk;
[0136] When the processed cell description text matches two abnormal indicator words in the medical keyword and phrase knowledge base, the processed cell description text is medium risk;
[0137] When the processed cell description text matches more than or equal to three abnormal indicator words in the medical keyword and phrase knowledge base, the processed cell description text is considered high risk;
[0138] This is a rapid and highly interpretable screening method. The system first analyzes the generated cell description text. Using a built-in knowledge base of medical keywords and phrases (e.g., including abnormality indicators such as nuclear atypia, hyperchromasia, rough chromatin, koilocytes, and keratinized beads, as well as normal indicators such as clear cell borders and normal nuclear-cytoplasmic ratio), it detects the presence of these keywords in the description using efficient string matching or regular expression algorithms. Based on the type, number, and combination of matched keywords, the cells are initially classified for abnormality risk.
[0139] S82, inputting the cell description text obtained by the cell description text generation module in the inference phase into the classification head at the end of the trained visual language large model, and the classification head outputting the probability distribution of the cell belonging to each predefined category;
[0140] If the maximum probability is greater than or equal to 0.7, the category corresponding to the maximum probability output by the classification head is the classification result, and the classification result is normal NILM or atypical squamous cells ASC-US or low-grade squamous intraepithelial lesion LSIL or high-grade squamous intraepithelial lesion HSIL;
[0141] If the maximum probability is less than 0.7, the category output by the classification head is uncertain;
[0142] The classification head is a feed-forward network or a linear layer + Softmax activation function.
[0143] For example: the model infers the input image and obtains the probability scores of 4 categories (Softmax output)
[0144] For example: NILM = 0.82, ASC-US = 0.10, LSIL = 0.05, HSIL = 0.02;
[0145] Take the maximum category as the predicted category Here NILM = 0.82 Maximum → Preliminary predicted category = NILM.
[0146] To achieve more accurate and robust classification results, this method leverages the final fused feature embedding vector generated by the VLM during inference. This high-dimensional vector (e.g., 768 dimensions) represents the model's deep understanding and condensed representation of the cell image and input prompt. This vector is directly fed into a lightweight classifier head, attached to the VLM during training. This head can be a simple feed-forward network (FFN) or a linear layer with a softmax activation function. The classifier directly outputs a probability distribution for whether the cell belongs to one of the predefined categories (e.g., normal / ASC-US / LSIL / HSIL), or a binary probability value P(y=abnormal | Image, Prompt) for normal / abnormal. This method directly leverages the deep multimodal features learned by the model and generally achieves higher accuracy than simple text analysis.
[0147] Other steps and parameters are the same as those in Specific Embodiments 1 to 7-1.
[0148] Specific embodiment 9: This embodiment differs from any one of specific embodiments 1 to 8 in that the explainability verification and credibility assessment module is used for interactive question-answering and credibility determination; the specific process is as follows:
[0149] If the results of S81 and S82 are the same, the answer of S9 will be consistent with the results of S81 and S82;
[0150] If the results of S81 and S82 are inconsistent, the answer to S9 will produce the correct result;
[0151] S91. Interactive Q&A:
[0152] To enhance the transparency and reliability of the model's diagnostic conclusions, cells initially screened as suspicious or positive will automatically trigger a multi-turn QA verification module. This module simulates the pathologist's thought process when encountering a difficult case, probing the model's diagnostic logic chain through a series of follow-up questions. Based on the pre-set diagnostic decision tree, it automatically initiates a series of questions, from macro to micro, from general to specific, such as:
[0153] Q1: What is the abnormality of the overall morphology of the cervical cell? A1: The nuclear-cytoplasmic ratio of the cervical cell is significantly increased;
[0154] Q2: Based on A1, since the nuclear-cytoplasmic ratio is increased, please describe the nuclear membrane of the cervical cell. A2: The nuclear membrane edge presents irregular indentation and shrinkage;
[0155] Q3: Based on A2, what is the morphology of the chromatin? A3: The chromatin is condensed into coarse granular shape, and the distribution is uneven, and obvious nuclear clearing zone can be seen;
[0156] This process not only verifies the preliminary judgment, but also accumulates more rich textual evidence for the final diagnosis;
[0157] Q1 represents the first question, A1 represents the first answer, Q2 represents the second question, A2 represents the second answer, Q3 represents the third question, and A3 represents the third answer;
[0158] S92, calculate the credibility index :
[0159] In each round of question and answer, the credibility index is obtained while generating the answer; after completing all the question and answers, the credibility indexes of all rounds are comprehensively calculated by weighted average to calculate a final credibility index , ; set the credibility threshold;
[0160] When the final credibility index is greater than or equal to the credibility threshold, A3 is a high credibility result;
[0161] When the final credibility index is less than the credibility threshold, A3 is a low credibility result;
[0162] For example: in the first round of question and answer, the credibility index is obtained while generating the answer ; in the second round of question and answer, the credibility index is obtained while generating the answer ; in the third round of question and answer, the credibility index is obtained while generating the answer ; the credibility indexes of all rounds are comprehensively calculated by weighted average ; 、 、 is the weight ;
[0163] In each round of Q&A, while the model generates an answer, its internal probability distribution also reflects its confidence in the answer. This confidence level is quantified and combined with the importance of the question in the diagnostic process to give a score for each round of Q&A. After all Q&A rounds are completed, the scores of all rounds are combined and a final diagnostic confidence index (CI) is calculated through weighted averaging or other aggregation algorithms. Its value range is Only when If the confidence level is higher than a preset threshold (e.g. 0.85), the diagnostic label is considered a high confidence result. The system will automatically mark the results as requiring manual review, achieving an efficient human-computer collaborative workflow.
[0164] The other steps and parameters are the same as those in the specific implementation modes 1 to 8-1.
[0165] Specific embodiment 10: This embodiment differs from any one of specific embodiments 1 to 9 in that the structured diagnostic report output module is used to aggregate the full-slice analysis results and generate a structured report and suggestions; the specific process is as follows:
[0166] S101, aggregated full-slice analysis results:
[0167] After completing the analysis of all valid cells in the WSI, the report generation stage begins.
[0168] First, the analysis results of all cells were aggregated and statistically analyzed based on S42, S7, S81, S82, S91, and S92.
[0169] Calculate the total number of cells, the total number of abnormal cells, and the number and percentage of each abnormal subtype (ASC-US, LSIL, HSIL, etc.);
[0170] Then, based on the original coordinates of each abnormal cell in the WSI, a heatmap of the spatial distribution of diseased cells is generated. The areas where abnormal cells are concentrated are highlighted on the thumbnail of the whole film based on the heatmap.
[0171] This heat map greatly helps pathologists quickly locate areas that require special attention, significantly improving film reading efficiency.
[0172] S102. Generate a structured report:
[0173] All analytical data is collated and output into a standardized, structured diagnostic report. The report supports multiple formats, including JSON, CSV, or Word, to facilitate seamless integration into the hospital's existing laboratory information system (LIS) or electronic medical record system (EMR). The core content of the report is clearly structured, and the structured report includes:
[0174] Basic information: patient ID, sample number, WSI file name, analysis date;
[0175] Analysis summary: Total cell count, total abnormal cell count, and number and percentage of each abnormal subtype (ASC-US, LSIL, HSIL, etc.) calculated by S101;
[0176] High-risk cell details list: This is the most critical part of the report. For each cell judged as high-risk or positive, there is a detailed entry, including:
[0177] The unique ID of the cervical cell obtained by S1 and the coordinates of the cervical cell in WSI;
[0178] Snapshot of high-resolution image of cervical cells obtained by S1 (S1 acquired);
[0179] S81 obtained cell description text;
[0180] The predicted category label (e.g., HSIL) obtained by S82;
[0181] The interactive Q&A log in S9;
[0182] Final credibility index in S9 ;
[0183] S103, Automatically generate preliminary diagnostic opinions and suggestions:
[0184] If the abnormal level (diagnostic result) is normal NILM, the final credibility index If the number of abnormal cells is 0, the preliminary diagnosis is that there is no abnormality found, and the recommended operation (rule logic) is routine review with a follow-up interval of 3 years (according to TBS / NCCN);
[0185] If the abnormal grade (diagnosis result) is atypical squamous cells ASC-US, the final confidence index , the number of abnormal cells The preliminary diagnosis is atypical squamous cells, which are of unknown significance. The recommended procedure (rule logic) is to recommend HPV testing. If HPV (+) → colposcopy; if HPV (-) → retest in 12 months.
[0186] If the abnormal grade (diagnosis result) is low-grade squamous intraepithelial lesion LSIL, the final confidence index , the number of abnormal cells The preliminary diagnosis is low-grade squamous intraepithelial lesion; the recommended procedure (rule logic) is colposcopy; young women can be followed up (flexible treatment according to the guidelines);
[0187] If the abnormal grade (diagnosis result) is high-grade squamous intraepithelial lesion HSIL, the final confidence index , abnormal cell count The preliminary diagnosis is high-grade squamous intraepithelial lesion; the recommended operation (rule logic) is immediate colposcopy + directed biopsy; conization is performed if necessary;
[0188] If the abnormality level (diagnostic result) is uncertain, the final credibility index is arbitrary, and the number of abnormal cells is arbitrary; the preliminary diagnosis opinion is that the result credibility is low; the recommended action (rule logic) is to mark it as requiring manual review, and no clear conclusion is automatically generated;
[0189] 1. Multi-condition trigger: consider the diagnosis category, Credibility, number of abnormal cells;
[0190] For example: If there is only 1 ASC-US, =0.68 → Output requires manual review
[0191] 2. Guideline-based mapping: The table content corresponds to the Bethesda + NCCN cervical cancer screening guidelines and is converted into IF–THEN rules;
[0192] Example: IF diagnosis = HSIL AND → Recommendation = colposcopy + biopsy;
[0193] 3. Flexible Updates: The knowledge base is built into the model and can be quickly modified based on the latest guideline updates or local hospital diagnostic and treatment protocols. At the end of the report, the system automatically generates a preliminary diagnosis and follow-up recommendations based on a comprehensive analysis of the entire slide data (for example, the number of abnormal cells, the most severe lesion grade, and the proportion of high-confidence results) and the built-in clinical guideline knowledge base.
[0194] For example: Preliminary Diagnostic Opinion: Screening result positive. Multiple cells consistent with high-grade squamous intraepithelial lesion (HSIL) were found in the sample, giving the diagnosis high confidence. Recommendation: Based on the patient's specific circumstances, the clinician should consider colposcopy and tissue biopsy for final confirmation and recommend high-risk HPV typing testing. This comprehensive, interpretable, and traceable report provides strong decision support for the pathologist's final sign-off.
[0195] The other steps and parameters are the same as those in the specific implementation modes 1 to 9-1.
[0196] This invention, through the deep integration and innovation of cutting-edge artificial intelligence technologies such as large-scale visual language models, instance segmentation, and multimodal fusion, has built an intelligent diagnostic system for cervical cytopathology that has achieved substantial breakthroughs in diagnostic accuracy, process automation, and decision-making interpretability. The core originality and technical protection scope of this invention cover the following key points:
[0197] 1. Efficient fine-tuning architecture and methods for large-scale visual language models for pathological diagnosis;
[0198] Key Technical Points: To address the challenges of high training costs and scarce domain data when migrating large general visual language models to the medical field, this paper proposes an efficient fine-tuning method centered around low-rank adaptation (LoRA). This method freezes the main parameters of the pre-trained model and only injects and trains a very small number of trainable parameters in the attention layers of key modules such as ViT and the language decoder. This method retains the powerful generalization capabilities of the large model while rapidly adapting to the detailed description and diagnosis tasks of cervical cytology images with extremely low computing power and data requirements.
[0199] Core protection points: The contents to be protected include: (1) the specific network topology of integrating the LoRA module into the visual language model (especially the cross-modal fusion layer); (2) the multi-task joint training strategy and loss function design that drives the model to simultaneously perform text generation, lesion classification and key feature regression under the LoRA framework (as described in S6); (3) a serialized prompt engineering template for cervical cell diagnosis that is deeply coupled with the fine-tuning architecture.
[0200] 2. A high-precision, adaptive cell instance segmentation algorithm that integrates multi-source information;
[0201] Key Technical Points: To address the segmentation challenges of densely packed, overlapping, and morphologically diverse cells in cervical pathology images, this paper designs an adaptive instance-based segmentation pipeline. This pipeline not only integrates multiple advanced segmentation models, such as HoverNet, Cellpose, and SAM (as described in S4), but also includes a series of specially optimized post-processing steps, such as an adhesion cell separation algorithm based on morphology and horizontal / vertical distance maps of cell nuclei, and an artifact filtering mechanism based on multi-dimensional features such as area and circularity.
[0202] Core protection points: The contents to be protected include: (1) Adaptive segmentation strategies that dynamically select or fuse multiple segmentation models to adapt to different image features; (2) Methods that combine the multi-scale features of the cell nucleus and cytoplasm with spatial context information to improve the accuracy of segmentation boundaries; (3) Methods that use a small amount of weak label (weak labels) data to train lightweight classifiers based on the feature embedding vector output by the model to achieve efficient abnormal cell detection.
[0203] 3. A highly interpretable diagnostic reasoning process based on a description-verification closed loop;
[0204] Key technical points: This invention overturns the black box model of traditional CAD systems that input images and output labels, and establishes a set of traceable and verifiable diagnostic reasoning paths. First, through the natural language generation module of S7, the model's understanding of cell morphology is converted into a text description readable by pathologists. Then, through the unique multi-round question-answering (QA) mechanism of S9, the model's preliminary judgment is automatically and logically questioned, and based on the consistency and certainty of the model's answers, a quantitative diagnostic credibility index is finally generated. .
[0205] Core protection points: The content to be protected includes: (1) a complete generative diagnostic method from images to natural language descriptions and then to diagnostic labels; (2) an automated, multi-round question-answering (QA) verification engine based on a diagnostic decision tree, and its algorithm for calculating the final diagnostic credibility; (3) a method for constructing a structured diagnostic knowledge chain formed by this process that can be used for clinical decision support, physician training, and medical record archiving.
[0206] 4. Full-process automated processing and generation technology from raw images to multi-format reports;
[0207] Key Technical Points: This invention achieves end-to-end automation, from automatically loading and preprocessing WSI raw images (S1-S3) to ultimately generating a structured diagnostic report rich in multi-dimensional information (S10). The generated report not only includes the final diagnostic recommendation but also integrates a heat map of diseased cells, an image snapshot of each suspicious cell, coordinate location, natural language description, classification labels, and QA verification records.
[0208] Core protection points: The contents to be protected include: (1) a fully automated, unmanned data processing pipeline from WSI to structured diagnostic reports; (2) a method for dynamically aggregating cell-level analysis results and generating comprehensive diagnostic reports containing images, text, quantitative data and diagnostic recommendations, supporting multiple formats such as JSON, CSV, and DOCX.
[0209] 5. Open multimodal information fusion framework and task scalability;
[0210] Key Technical Points: The system architecture of this invention boasts excellent openness. In addition to core image and text processing capabilities, it also incorporates interfaces that, through prompt engineering, can incorporate other dimensions of clinical information (such as patient age, high-risk HPV test results, and medical history) into the diagnostic model, enabling more accurate and personalized diagnosis. Furthermore, this core model framework can be easily extended to other related pathology analysis tasks, such as automated grading of cervical inflammatory responses and assessment of cell proliferation activity.
[0211] Core protection points: The contents to be protected include: (1) a method for injecting non-image structured / unstructured clinical data into a visual language model through a specific prompt template to achieve multimodal joint diagnosis; (2) a methodology for migrating and expanding this core technology framework to other cytopathology or histopathology analysis tasks.
[0212] 6. Enterprise-level deployment and parallel inference architecture supporting high-throughput analysis;
[0213] Key Technical Points: To meet the needs of large-scale clinical applications, this system utilizes a microservices and modular design in its engineering implementation. Encapsulated through Docker containerization technology and supporting Kubernetes (K8s) cluster management and scheduling, it enables fast and flexible deployment on private or public clouds. The system's inference engine has been deeply optimized to support efficient parallel slicing and asynchronous inference of WSI images, capable of processing data at the level of tens or even hundreds of millions of cells.
[0214] Core protection points: The contents to be protected include: (1) the modular and containerized deployment architecture of the entire diagnostic system; (2) the system scheduling and task management methods that can realize efficient parallel processing and distributed reasoning of massive WSI data on large-scale computing clusters.
[0215] Regarding the technical solution in item 3, are there any other alternative solutions that can also achieve the purpose of the invention?
[0216] To objectively evaluate the innovation and technical advantages of the present invention, this section constructs and analyzes an alternative technical solution. This solution, instead of employing a visual language model (VLM), utilizes a traditional cascaded deep learning pipeline of segmentation, feature extraction, and classification, combined with rule-based text generation, to simulate and achieve similar objectives. Its specific structure and limitations are analyzed as follows:
[0217] This alternative solution is mainly composed of the following technical modules in series:
[0218] 1. Image preprocessing and cell segmentation: In steps S1-S2, the OpenSlide library is also used to load WSI and perform color normalization. However, in the cell segmentation stage of S4, the scheme uses classic semantic segmentation networks (such as U-Net, FCN). Such networks are designed to classify each pixel in the image (for example, into nucleus, cytoplasm, background), but they themselves do not have the ability to distinguish between two closely connected similar objects (ie, instances). Therefore, when dealing with cell overlap or dense areas that are common in cervical smears, the segmentation accuracy will drop significantly, often resulting in cell under-segmentation (misclassifying multiple cells as one) or over-segmentation, which fundamentally affects the accuracy of subsequent single-cell analysis.
[0219] 2. Image feature extraction and abnormality classification: After segmenting the cell area, the scheme uses an independent image classification network (such as ResNet, InceptionV3) as a feature extractor in S3 (corresponding to S8 of the present invention) to convert each cell image into a feature vector of fixed dimension. Subsequently, these feature vectors are input into another independent traditional machine learning classifier (such as support vector machine SVM, random forest) or a shallow fully connected neural network for normal / abnormal binary classification or multi-classification. This two-stage model of feature extraction + classification has obvious defects: first, the feature extraction network and the final classifier are trained separately, and end-to-end collaborative optimization cannot be performed; second, the performance of the model is highly dependent on large-scale, high-quality pixel-level annotation data, and the generalization ability is limited, and the recognition effect of atypical cells with variable morphology is poor.
[0220] 3. Pseudo-natural language description generation: In order to simulate the natural language description function in S7 of the present invention, this alternative adopts a text generation system based on templates and rules. The system first obtains the predicted label (such as HSIL) from the classification model and some preset quantitative indicators (for example, the calculated cell nucleus area is X, the average grayscale value is Y) from the feature extractor, and then rigidly fills these discrete information into a preset sentence template to generate a description such as the cell is classified as HSIL, its cell nucleus area is X, and the average chromatin concentration is Y. The language generated by this method is stiff and unnatural, and cannot describe the complex morphological features outside the template, and does not have the ability to understand the context and perform logical reasoning.
[0221] 4. Simple threshold verification and report generation: In the abnormality verification link, this scheme does not have the multi-round question and answer (QA) mechanism in S9 of the present application. It can only compare the confidence score (for example, Softmax probability value) output by the classifier with a fixed threshold to judge the reliability of the result. This verification method is one-way and non-interactive, and cannot explore the specific reasons for the model's judgment. The final generated report only contains simple statistical data and a list of classification results, lacking the rich and traceable diagnostic evidence chain in S10 of the present application.
[0222] Through the above analysis, it can be seen that although this alternative scheme can theoretically build a basic diagnostic process, in many core dimensions, its effect and ability are far inferior to the innovative scheme based on visual language large model proposed by the present application. The specific comparison is shown in Table 1:
[0223]
[0224] Conclusion: Due to its cascading architecture, heavy dependence on large-scale accurate labeling, and lack of deep semantic understanding ability, this alternative scheme has insurmountable bottlenecks in diagnostic accuracy, explainability, automation level, and system scalability. The present application fundamentally solves these pain points by introducing advanced visual language large models, supplemented by a series of innovative methods such as efficient fine-tuning, multi-task learning, and interactive verification, and its comprehensive technical effect far surpasses that of traditional alternative schemes.
[0225] What are the advantages of the present application compared with the closest prior art?
[0226] The cervical cell intelligent screening and diagnosis method based on visual language large model (VLM) fine-tuning proposed by the present application substantially surpasses the prior art in terms of diagnostic paradigm, model performance, process intelligence, and application ecosystem. Its core advantages can be summarized as follows:
[0227] 1. In terms of diagnostic accuracy and model generalization ability: leap from feature engineering to deep semantic understanding.
[0228] Limitations of existing technology: Traditional diagnostic methods rely heavily on CNN-based architectures such as ResNet and U-Net. These models can only extract local and pre-defined image features when dealing with high-resolution and high-complexity pathological images. They have limited ability to capture subtle but crucial morphological differences between cells, such as atypical lesions. Their performance is highly dependent on large-scale, pixel-level accurate labeled datasets. Advantages of the present invention: By introducing a visual language large model (VLM) and combining it with a multi-task learning framework in S6, the invention achieves a transition from seeing to understanding. The model not only analyzes pixels but also generates semantic descriptions consistent with image content. This cross-modal joint modeling capability enables it to capture deeper and more subtle pathological features. More importantly, through the LoRA efficient fine-tuning strategy, the invention eliminates the dependence on massive labeled data and quickly adapts to a small amount of domain data, greatly improving the model's generalization ability and deployment efficiency in different medical institutions and under different staining conditions.
[0229] 2. In the explainability of the diagnostic process: from black box output to transparent and traceable diagnostic dialogue.
[0230] Limitations of existing technology: Most existing CAD systems are black box models that can only output a classification label or probability value, making it difficult for clinicians to establish trust and conduct effective quality control and review. Advantages of the present invention: The invention creates a complete human-machine collaborative diagnostic logic chain. First, the natural language generation module of S7 presents the model's thinking process in a doctor-readable text format, such as "nucleus is large," "chromatin is dense," and "nuclear membrane is irregular." Second, S9's innovative multi-round question and answer (QA) verification mechanism builds a clear and traceable reasoning path through automated follow-up and confirmation, and finally quantifies it into a diagnostic confidence index. This not only makes the AI decision-making process transparent, but also transforms the diagnostic process from a one-way notification to an interactive dialogue, providing unprecedented powerful tools for clinical review, physician training, and medical quality control.
[0231] 3. In the automation and intelligence level of the system process: from auxiliary tool to workflow engine.
[0232] Limitations of existing technologies: Traditional AI tools usually only solve an isolated link in the diagnostic process (such as segmentation or classification), with low integration between modules. A large amount of manual operation is still required for series connection, and the degree of automation is limited. Advantages of the present invention: The present invention constructs an end-to-end, full-process automated diagnostic architecture from S1 to S10. It is not only an analysis tool, but also an intelligent workflow engine that covers everything from WSI image loading, precise cell positioning, multi-dimensional intelligent analysis to the automatic generation of rich-text structured reports in S10. This high degree of automation and pipeline operation can free pathologists from heavy and repetitive screening work, allowing them to focus on decision-making on difficult cases, thereby subversively improving the work efficiency of the entire pathology department.
[0233] 4. In terms of recall and accuracy of screening: from single criterion to joint decision-making based on multi-dimensional evidence.
[0234] Limitations of existing technologies: Traditional methods usually rely only on a single criterion of the classifier to identify abnormal cells, which can easily lead to missed diagnosis (low recall rate) or misdiagnosis (low accuracy rate) due to atypical feature expression. Advantages of the present invention: The present invention establishes a dual-channel, multi-evidence joint discrimination mechanism of visual features + text semantics in S8. When making a judgment, the model not only relies on the deep visual features extracted from the image, but also synchronously evaluates the text description content generated by it. Through multiple means such as keyword analysis and semantic similarity matching, combined with the confirmation of the QA verification link, the system can make more reliable judgments on cells with blurred boundaries or atypical morphology, thereby significantly improving the recall rate of early lesions while ensuring high accuracy, and effectively reducing the risk of missed diagnosis.
[0235] 5. Standardization and data reuse of diagnostic results: From instantaneous results to traceable structured knowledge assets.
[0236] Limitations of existing technologies: The output results of existing models (such as a heat map or a classification label) are often discrete and unstructured, making them difficult to manage, query, and reuse systematically. Advantages of the present invention: The present invention solidifies the entire process information of each diagnosis - including cell coordinates, image snapshots, diagnostic descriptions, classification labels, QA records, and credibility, etc. - in a unified JSON structure through the report generation module of S10. This not only generates a report, but also creates standardized, machine-readable, and traceable electronic pathology records. These structured data assets can be easily retrieved and reviewed, and can flow back into the model training library to form a closed-loop ecosystem of continuous learning and self-optimization.
[0237] 6. In terms of the scalability of the technical framework: from fixed functions to prompt-driven task generalization.
[0238] Limitations of existing technologies: CNN-based systems are usually hard-coded for specific tasks and have a single function. If new functions are to be added (such as assessing the level of inflammation), a new model almost needs to be developed. Advantages of the present invention: The present invention is based on Prompt's configurable VLM design, which gives the system unprecedented task generalization capabilities. By simply modifying the input natural language prompts, the model can be guided to perform new tasks, such as cell subtype discrimination, cervical inflammation degree grading, and even analysis of specific molecular pathology markers. In addition, the framework can also easily integrate text-based multimodal data such as HPV test results and clinical history to achieve a more comprehensive individualized diagnosis, which is difficult to achieve with traditional architectures.
[0239] In summary, this invention is not a simple improvement on existing technologies, but rather a fundamental innovation in core concepts and technical approaches. It successfully advances intelligent pathology diagnosis from a perceptual intelligence stage centered on pixel recognition to a new level of cognitive intelligence with preliminary logical reasoning and language interaction capabilities. This paves the way for truly efficient, reliable, and trustworthy human-computer collaborative diagnosis, and holds significant technological leadership and clinical application prospects.
[0240] The present invention may have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art may make various corresponding changes and modifications based on the present invention, but these corresponding changes and modifications should all fall within the scope of protection of the claims attached to the present invention.
Claims
1. An intelligent diagnosis system for cervical cell images based on a multimodal visual language model, characterized by: The system comprises: Image block acquisition module, image preprocessing and color normalization module, effective tissue area screening module, cell area extraction and post-processing module, image and text processing module, visual language large model acquisition and training module, cell description text generation module in the reasoning stage, risk judgment and classification module, explainability verification and credibility assessment module, and structured diagnostic report output module; The image block acquisition module is used to acquire a whole slice image WSI, and divide the acquired whole slice image WSI into blocks to obtain image blocks; The image preprocessing and color normalization module is used to convert the image block from RGB space to optical density space, construct the dye matrix of the image based on the optical density space, obtain the concentration matrix based on the dye matrix of the image, obtain the reconstruction matrix based on the concentration matrix and the target dye matrix, and convert the reconstruction matrix back to RGB space; The effective tissue area screening module is used to screen each image block converted back to the RGB space obtained by the image preprocessing and color normalization module to obtain a valid image block; The cell region extraction and post-processing module is used to extract the cell region from the effective image block obtained by the effective tissue region screening module, and perform post-processing on the extracted cell region to obtain the segmented cell region image; The image-side and text-side processing modules are used to perform image-side and text-side processing on the segmented cell region images obtained by the cell region extraction and post-processing modules; The visual language model acquisition and training module is used to acquire the visual language model and train the acquired visual language model to obtain a trained visual language model. The cell description text generation module in the inference phase is used to input a single cell image to be tested and a task prompt into the trained visual language model, and obtain the cell description text based on the beam search strategy; The risk determination and classification module is used to perform risk determination and classification on the cell description text obtained by the cell description text generation module in the reasoning phase; The explainability verification and credibility assessment module is used for interactive question-answering verification and credibility determination; The structured diagnostic report output module is used to aggregate the full-slice analysis results and generate structured reports and recommendations.
2. The intelligent diagnosis system for cervical cell images based on a multimodal visual language large model according to claim 1, characterized in that: The image block acquisition module is used to acquire a whole slice image WSI, and divide the acquired whole slice image WSI into blocks to obtain image blocks; The specific process is: S11. Use the OpenSlide open source library to obtain the whole-slide image WSI and its metadata; The metadata of the whole-slide image WSI includes: the physical size of the whole-slide image WSI, the scanning magnification, and the resolution of each level in the image pyramid; S12, dividing the whole slice image WSI obtained in S11 into blocks; Each image patch contains the following information: The spatial coordinates of the whole slice image WSI obtained by each image block Patch in S11 , the level to which each image block Patch belongs in the whole slice image WSI obtained in S11 and the size of each image block Patch; The process is expressed as: in, represents an image block, represents the whole slice image WSI acquired by S11, Indicates the highest resolution level, express The coordinates of the upper left corner, express The width, express height.
3. The intelligent diagnosis system for cervical cell images based on a multimodal visual language large model according to claim 2, characterized in that: The image preprocessing and color normalization module is used to convert the image block from the RGB space to the optical density space, construct a dye matrix of the image based on the optical density space, obtain a concentration matrix based on the dye matrix of the image, obtain a reconstruction matrix based on the concentration matrix and the target dye matrix, and convert the reconstruction matrix back to the RGB space; The specific process is: S21, convert each image block obtained in S1 from RGB color space to optical density space; the conversion formula is: in, represents the pixel intensity value of the image block obtained by S1, Indicates the reference intensity value of background white light; S22. Settings Threshold; In S21, select greater than Threshold Vector, based on the selected greater than in S21 Threshold The vector constructs the dye matrix of the image; performs color deconvolution on each element in the dye matrix of the image, decomposes the color of each element into the concentration value of hematoxylin and eosin, and obtains a concentration matrix; multiplies the concentration matrix with the set target dye matrix to obtain a reconstruction matrix, and converts the reconstruction matrix back to RGB space; The effective tissue area screening module is used to screen each image block converted back to the RGB space obtained by the image preprocessing and color normalization module to obtain a valid image block; The specific process is: The Otsu algorithm is used to process each image block converted back to the RGB space obtained by the image preprocessing and color normalization module to obtain a binary image of each image block. The proportion of foreground pixels in the binary image of each image block is calculated, and only valid image blocks whose proportion of foreground pixels in the binary image of each image block exceeds a preset threshold (for example, 20%) are retained.
4. The intelligent diagnosis system for cervical cell images based on a multimodal visual language large model according to claim 3 is characterized by: The cell region extraction and post-processing module is used to extract the cell region from the effective image block obtained by the effective tissue region screening module, and post-process the extracted cell region to obtain a segmented cell region image; The specific process is: S41、 Each valid image block obtained by the effective tissue area screening module is input into the Hover-Net network model, and the Hover-Net network model outputs the cell area image of each valid image block; Each valid image block obtained by the valid tissue area screening module is input into the Cellpose network model, and the Cellpose network model outputs the cell area image of each valid image block; Each valid image block obtained by the valid tissue area screening module is input into the Segment Anything Model network model, and the Segment Anything Model network model outputs the cell area image of each valid image block; S42、 Perform morphological filtering on the cell area image of each valid image block output by the network model to obtain each image after morphological filtering; perform attribute screening on each image after morphological filtering to obtain a filtered image; the process is: Perform an opening operation on the cell area image of each valid image block output by the network model, and perform a closing operation on each image after the opening operation to obtain each image after morphological filtering; Calculate the geometric properties of each image after morphological filtering, eliminate images that do not meet the geometric property threshold, and obtain the filtered image; The geometric properties are area, perimeter, circularity and density.
5. The intelligent diagnosis system for cervical cell images based on a multimodal visual language large model according to claim 4 is characterized by: The image-side and text-side processing modules are used to perform image-side and text-side processing on the segmented cell region image obtained by the cell region extraction and post-processing module; The specific process is: Each segmented cell region image obtained by the cell region extraction and post-processing module is uniformly scaled to the standard input size required by the visual language large model; Construct text prompts for each segmented cell region image obtained by the cell region extraction and post-processing module; The text prompts are: Comprehensive descriptive prompt: Please describe this image of cervical cells in detail from the perspective of a pathologist. Your analysis should include the size and shape of the nucleus, the thickness and uniformity of the chromatin distribution, and the nuclear-cytoplasmic ratio. Finally, based on the size and shape of the nucleus, the thickness and uniformity of the chromatin distribution, and the nuclear-cytoplasmic ratio, please provide a preliminary judgment on whether the cervical cells show signs of pathology. Specific feature query type prompt: whether the nuclear membrane of cervical cells is smooth and regular; Direct Classification Task-Based Prompt: Please classify cervical cells as one of the following based on The Bethesda System classification criteria: normal squamous cells, atypical squamous cells, low-grade squamous intraepithelial lesion, or high-grade squamous intraepithelial lesion.
6. The intelligent diagnosis system for cervical cell images based on a multimodal visual language large model according to claim 5, characterized in that: The visual language large model acquisition and training module is used to acquire the visual language large model and train the acquired visual language large model to obtain a trained visual language large model; The specific process is: S61. Use LoRA technology to fine-tune the language model to obtain a visual language model. The specific process is: The forward propagation process at a specific layer of the language model becomes: in, is a low-rank matrix, ; is a low-rank matrix, ; The token vector input to a certain layer of the language model; is a scaling constant; is the weight matrix of the language model; For order; The Token vector dimension of the input of a certain layer; The dimension of the Token vector output by a certain layer; is the output vector after adding LoRA low-rank update; S62. Constructing a multi-task joint loss function for a large visual language model : in, Represents text generation loss, which is cross entropy loss; express The weight hyperparameters of represents the classification loss; express The weight hyperparameters of Represents regression loss, which is the mean square error loss; express The weight hyperparameters.
7. The intelligent diagnosis system for cervical cell images based on a multimodal visual language large model according to claim 6, characterized in that: The cell description text generation module in the reasoning stage is used to input a single cell image to be tested and a task prompt into the trained visual language model, and obtain the cell description text based on the beam search strategy; The specific process is: A single-cell image to be tested and a task prompt are input into the trained visual language model. The trained visual language model will retain the k candidate sequences with the highest probability at each step, and predict the output of the next step based on the retained k candidate sequences with the highest probability. Finally, the trained visual language model outputs the cell description text with the best overall probability.
8. The intelligent diagnosis system for cervical cell images based on a multimodal visual language large model according to claim 7, characterized in that: The risk determination and classification module is used to perform risk determination and classification on the cell description text obtained by the cell description text generation module in the reasoning stage; The specific process is: S81、 Construct a knowledge base of medical keywords and phrases, which includes abnormal indicators such as nuclear atypia, nuclear hyperchromasia, rough chromatin, koilocytes, and cornified beads, as well as normal indicators such as clear cell borders and normal nuclear-cytoplasmic ratio; Performing regular expression processing on the cell description text obtained by the cell description text generation module in the inference phase to obtain the processed cell description text; Matching the processed cell description text in the medical keyword and phrase knowledge base; When the processed cell description text matches the normal indicator words in the medical keyword and phrase knowledge base, the processed cell description text is normal; When the processed cell description text matches one abnormal indicator word in the medical keyword and phrase knowledge base, the processed cell description text is low risk; When the processed cell description text matches two abnormal indicator words in the medical keyword and phrase knowledge base, the processed cell description text is medium risk; When the processed cell description text matches more than or equal to three abnormal indicator words in the medical keyword and phrase knowledge base, the processed cell description text is considered high risk; S82、 The cell description text obtained by the cell description text generation module in the inference phase is input into the classification head at the end of the trained visual language model. The classification head outputs the probability distribution of cells belonging to each predefined category. If the maximum probability is greater than or equal to 0.7, the category corresponding to the maximum probability output by the classification head is the classification result, and the classification result is normal NILM or atypical squamous cells ASC-US or low-grade squamous intraepithelial lesion LSIL or high-grade squamous intraepithelial lesion HSIL; If the maximum probability is less than 0.7, the category output by the classification head is uncertain; The classification head is a feed-forward network or a linear layer + Softmax activation function.
9. The intelligent diagnosis system for cervical cell images based on a multimodal visual language large model according to claim 8, characterized in that: The explainability verification and credibility assessment module is used for interactive question answering and credibility determination; The specific process is: S91. Interactive Q&A: Q1: What are the abnormalities in the overall morphology of cervical cells? A1: The nuclear-cytoplasmic ratio of cervical cells increases; Q2: Based on A1, since the nuclear-to-cytoplasmic ratio is increased, please describe the nuclear membrane of cervical cells; A2: The nuclear membrane edge shows irregular depressions and wrinkles; Q3: Based on A2, what is the chromatin form? A3: Chromatin is condensed into granules and unevenly distributed, with nuclear clearing areas visible; Q1 indicates the first question, A1 indicates the first answer, Q2 indicates the second question, A2 indicates the second answer, Q3 indicates the third question, A3 indicates the third answer; S92. Calculate the credibility index : In each round of question-answering, a credibility index is obtained while generating answers; After all questions and answers are completed, a final credibility index is calculated by weighted average of all rounds of credibility index. , ; Set credibility thresholds; When the final credibility index If it is greater than or equal to the confidence threshold, then A3 is a high-confidence result; When the final credibility index If it is less than the confidence threshold, A3 is a low-confidence result.
10. The intelligent diagnosis system for cervical cell images based on a multimodal visual language large model according to claim 9, characterized in that: The structured diagnostic report output module is used to aggregate the full-slice analysis results and generate a structured report and suggestions; The specific process is: S101, aggregated full-slice analysis results: First, the total number of cells, the total number of abnormal cells, and the number and percentage of each abnormal subtype were calculated based on S42, S7, S81, S82, S91, and S92; Then, based on the original coordinates of each abnormal cell in the WSI, a heat map of the spatial distribution of diseased cells is generated. The area where abnormal cells are concentrated is highlighted on the thumbnail of the whole film based on the heat map; S102. Generate a structured report: Structured reports include: Basic information: patient ID, sample number, WSI file name, analysis date; Analysis summary: S101 calculated total cell count, total abnormal cell count, and number and percentage of each abnormal subtype; A detailed list of high-risk cells, including: The unique ID of the cervical cell obtained by S1 and the coordinates of the cervical cell in WSI; S1 obtained images of cervical cells; S81 obtained cell description text; The predicted category label obtained by S82; Full transcript of the interactive Q&A session in S9; Final credibility index in S9 ; S103, Automatically generate preliminary diagnostic opinions and suggestions: If the abnormal level is normal NILM, the final credibility index If the number of abnormal cells is 0, the preliminary diagnosis is that there is no abnormality found, and the recommended operation is a routine review with a follow-up interval of 3 years; If the abnormal grade is atypical squamous cells ASC-US, the final confidence index , the number of abnormal cells ; The preliminary diagnosis is atypical squamous cells; the recommended procedure is to recommend HPV testing; If the abnormal grade is low-grade squamous intraepithelial lesion LSIL, the final confidence index , the number of abnormal cells ; The preliminary diagnosis is low-grade squamous intraepithelial lesion; the recommended operation is colposcopy; If the abnormal grade is high-grade squamous intraepithelial lesion HSIL, the final confidence index , the number of abnormal cells The preliminary diagnosis is high-grade squamous intraepithelial lesion; the recommended operation is immediate colposcopy + directed biopsy; If the abnormality level is uncertain, the final credibility index is arbitrary, and the number of abnormal cells is arbitrary; the preliminary diagnosis is that the credibility of the result is low; the recommended action is to mark it as requiring manual review and not automatically generate a clear conclusion.
Citation Information
Cited By
Chromosome anomaly detection method based on multi-modal large model
CN120954014A
Blood cell image automatic classification system based on deep learning
CN120976656A
Digestive tract pathological diagnosis visual language large model construction method based on reinforcement learning and application thereof
CN121354882A
Construction method and application of a visual language large model for diagnosis of digestive tract pathology based on reinforcement learning
CN121354882B
Pathological report generation and interaction system based on generative artificial intelligence
CN121583439A