Full-process pathological image analysis method and device
By combining multi-task deep learning models and multimodal feature libraries, the problems of low efficiency and insufficient accuracy in traditional gastric cancer pathology analysis have been solved, and full-process intelligent pathology image analysis has been realized, improving analysis efficiency and accuracy.
Patent Information
- Application Number
- CN202510895479.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-10
AI Technical Summary
Traditional gastric cancer pathology analysis relies on manual observation, which is inefficient and easily affected by subjective factors. Existing technical auxiliary tools have single functions and cannot comprehensively improve the accuracy and efficiency of analysis.
A multi-task deep learning model is used to analyze gastric cancer WSI slice images, build a multimodal feature library, realize multimodal retrieval of pathological information, and generate pathology reports. Knowledge distillation and data balancing techniques are combined to improve model performance.
It has realized the full process intelligence of gastric cancer pathology analysis, improved analysis efficiency and accuracy, and shortened the diagnosis time of difficult cases.
Smart Images

Figure CN120766022A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of medical artificial intelligence, in particular to a whole-process pathological image analysis method and device. BACKGROUND
[0002] Traditional gastric cancer pathological analysis mainly relies on pathologists to observe slides under a microscope, which is low in efficiency and easily affected by subjective factors. Some existing single-point technologies, such as simply using computer vision to perform preliminary image feature extraction or simple case image database retrieval tools, although to some extent, assist in gastric cancer pathological analysis, but have the problem of single function and cannot improve the comprehensiveness and accuracy of gastric cancer pathological analysis as a whole. SUMMARY
[0003] Therefore, the purpose of the present application is to provide a whole-process pathological image analysis method and device, which can realize the whole-process intelligentization from image analysis to pathological report generation.
[0004] The whole-process pathological image analysis method provided by the present application comprises the following steps: Obtain a whole field digital WSI slice image derived from a gastric cancer patient, and analyze the obtained WSI slice image using a multi-task deep learning model to obtain pathological morphology data; wherein the multi-task deep learning model adopts a multi-task architecture comprising a tumor task branch model, a margin task branch model and a lymph node task branch model, and is used for tumor diagnosis recognition, margin diagnosis recognition and lymph node diagnosis recognition of the obtained WSI slice image; Based on the image features of a large number of WSI slice images and the semantic features of a large number of pathological texts, a multi-modal feature library containing image-text pairs is constructed, and multi-modal retrieval of pathological information is performed on the to-be-retrieved texts or to-be-retrieved images based on the multi-modal feature library; Integrate the obtained pathological morphology data and the pathological information, and fill in the content according to a preset template to generate a pathological report.
[0005] In some embodiments, the tumor diagnosis recognition, margin diagnosis recognition and lymph node diagnosis recognition of the obtained WSI slice image comprise the following steps: The tumor task branch model first locates a tumor region from the input WSI slice image based on an anchor box mechanism, then uses a feature pyramid network FPN to perform multi-scale feature extraction on the tumor region, recognizes different tumor histological types, and performs morphological feature extraction based on a convolutional neural network CNN to evaluate the degree of tissue differentiation; The surgical margin line is located from the input WSI slice image based on an edge detection algorithm, and the cancer cells in the range of the surgical margin line are identified based on pixel-level semantic segmentation to determine the residual state of the cancer cells at the surgical margin by the surgical margin task branch model. The lymph node contour is identified from the input WSI slice image based on pixel-level semantic segmentation, and the lymph node contour is detected for micrometastasis by a target detection network combined with a focal loss function to obtain metastasis parameters.
[0006] In some embodiments, the multi-task deep learning model is trained in a knowledge distillation manner, including the following steps: A WSI slice image sample dataset is obtained, which includes a plurality of WSI slice image samples and pathological morphology data labeled for each WSI slice image sample. The pathological morphology data includes tumor histological type and its histological differentiation degree, surgical margin line position and its surgical margin cancer cell residual state, lymph node contour and its metastasis parameters. First, the teacher model is trained based on the WSI slice image sample dataset to obtain a trained teacher model. Then, the knowledge is transferred to the student model through the set knowledge distillation parameters for training to obtain a trained student model, which is used as the final multi-task deep learning model for WSI slice image analysis.
[0007] In some embodiments, the training of the teacher model based on the WSI slice image sample dataset includes the following steps: The number of samples of different categories of pathological morphology data in the WSI slice image sample dataset is counted, and the minority class samples and the majority class samples are determined. The minority class samples and the majority class samples are processed according to the set balancing strategy. The minority class samples are subjected to data augmentation and oversampling, the majority class samples are subjected to undersampling, and the loss function during model training is set according to the sample quantity proportion of different categories.
[0008] In some embodiments, the multi-modal retrieval of pathological information based on the multi-modal feature library for the text to be retrieved or the image to be retrieved includes the following steps: The image features of the image to be retrieved are extracted and compared with the image features in the multi-modal feature library to obtain an image list related to the image to be retrieved and corresponding pathological information. Extracting semantic features of the text to be retrieved and mapping them to the same space as the image features, and calculating the similarity between the mapped semantic features and the image features in the multimodal feature library to obtain a list of images related to the text to be retrieved and the corresponding pathological information; Extracting image features of the image to be retrieved and converting them into text descriptions, and matching the converted text with pathology text in the multimodal feature library to obtain pathology information related to the image to be retrieved; The semantic features of the text to be retrieved are extracted and similarity calculation is performed between the semantic features in the multimodal feature library to obtain pathological information related to the text to be retrieved.
[0009] In some embodiments, integrating the obtained pathological morphological data and the pathological information, and filling in the content according to a preset template to generate a pathology report includes the following steps: Set different report templates and their corresponding formats according to the needs of gastric cancer pathology analysis; Structuring the acquired pathological morphological data and pathological information, and filling the selected report template with content according to a set format to obtain a pathology report; Obtain feedback data from the attending physician, and modify the pathology report based on the feedback data and then archive it.
[0010] In some embodiments, the method further comprises the following steps: The multi-task deep learning model is optimized based on the pathological information obtained from the multimodal feature library.
[0011] In some embodiments, a full-process pathology image analysis device is further provided, comprising: An image analysis module is configured to acquire full-field digital WSI slice images from gastric cancer patients and analyze the acquired WSI slice images using a multi-task deep learning model to obtain pathological morphological data; wherein the multi-task deep learning model adopts a multi-task architecture including a tumor task branch model, a resection margin task branch model, and a lymph node task branch model, and is configured to perform tumor diagnosis and identification, resection margin diagnosis and identification, and lymph node diagnosis and identification on the acquired WSI slice images; A pathology information retrieval module is used to construct a multimodal feature library containing image-text pairs based on the image features of a large number of WSI slice images and the semantic features of a large number of pathology texts, and to perform multimodal retrieval of pathology information on the text to be retrieved or the image to be retrieved based on the multimodal feature library; The pathology report generating module is used to integrate the obtained pathological morphology data and the pathological information, and fill in the content according to the preset template to generate a pathology report.
[0012] In some embodiments, an electronic device is also provided, comprising: a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor and the memory communicate through the bus, and when the machine-readable instructions are executed by the processor, the steps of any one of the above-described full-process pathology image analysis methods are performed.
[0013] In some embodiments, a computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the above-mentioned full-process pathology image analysis methods are executed.
[0014] The present application describes a full-process pathological image analysis method and device, which obtains full-field digital WSI slice images from gastric cancer patients and analyzes the obtained WSI slice images using a multi-task deep learning model to obtain pathological morphological data; wherein the multi-task deep learning model adopts a multi-task architecture including a tumor task branch model, a resection margin task branch model, and a lymph node task branch model, for performing tumor diagnosis and identification, resection margin diagnosis and identification, and lymph node diagnosis and identification on the obtained WSI slice images; based on the image features of a large number of WSI slice images and the semantic features of a large number of pathological texts, a multimodal feature library containing image-text pairs is constructed, and multimodal retrieval of pathological information is performed on the text to be retrieved or the image to be retrieved based on the multimodal feature library; the obtained pathological morphological data and the pathological information are integrated, and content is filled in according to a preset template to generate a pathology report. Thus, through pathological image analysis based on multi-task deep learning, cross-modal interactive image and text retrieval, and automatic generation of pathology reports, the full process of pathological analysis is intelligentized, and the efficiency and accuracy of gastric cancer pathological analysis are improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0016] Figure 1 A flowchart of the full-process pathological image analysis method according to an embodiment of the present application is shown; Figure 2 A flowchart of training the multi-task deep learning model according to an embodiment of the present application is shown; Figure 3A flowchart of the multi-modal retrieval of pathological information of the to-be-retrieved text or the to-be-retrieved image based on the multi-modal feature library according to the embodiments of the present application is shown. Figure 4 A flowchart of the integration of the obtained pathological morphological data and the pathological information, and the content filling according to a preset template to generate a pathological report according to the embodiments of the present application is shown. Figure 5 A structural schematic diagram of the whole-process pathological image analysis device according to the embodiments of the present application is shown. Figure 6 A structural schematic diagram of the electronic device according to the embodiments of the present application is shown. DETAILED DESCRIPTION
[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below in detail with reference to the drawings in the embodiments of the present application. It should be understood that the drawings in the present application only serve the purpose of description and illustration, and are not used to limit the scope of protection of the present application. In addition, it should be understood that the schematic drawings are not drawn according to the actual proportions. The flowchart used in the present application shows the operations implemented according to some embodiments of the present application. It should be understood that the operations of the flowchart can not be implemented in sequence, and the steps without logical context relationship can be reversed in sequence or implemented simultaneously. In addition, one or more other operations can be added to the flowchart or one or more operations can be removed from the flowchart under the guidance of the content of the present application.
[0018] In addition, the described embodiments are only some of the embodiments of the present application, not all the embodiments. The components of the embodiments of the present application described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0019] It should be noted that the term "comprising" will be used in the embodiments of the present application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.
[0020] In view of the technical problems proposed in the background art, the present application provides a whole-process pathological image analysis method, device, electronic device, and storage medium, which can realize the whole-process intelligentization from image analysis to pathological report generation.
[0021] Reference is made to the drawings accompanying the specification Figure 1, a full-process pathology image analysis method provided by this application includes the following steps: S1. Obtaining full-field digital WSI slice images from a gastric cancer patient, and analyzing the acquired WSI slice images using a multi-task deep learning model to obtain pathological morphological data; wherein the multi-task deep learning model adopts a multi-task architecture including a tumor task branch model, a resection margin task branch model, and a lymph node task branch model, for performing tumor diagnosis and identification, resection margin diagnosis and identification, and lymph node diagnosis and identification on the acquired WSI slice images; S2. Based on the image features of a large number of WSI slice images and the semantic features of a large number of pathology texts, a multimodal feature library containing image-text pairs is constructed, and multimodal retrieval of pathology information of the text to be retrieved or the image to be retrieved is performed based on the multimodal feature library; S3. Integrate the obtained pathological morphological data and the pathological information, and fill in the content according to a preset template to generate a pathology report.
[0022] In step S1, WSI (Whole Slide Imaging) slices, i.e., full-field digital slices, are obtained by scanning glass pathology slices with a professional microscope scanner (such as 3D HISTECH or Leica Aperio) to convert optical images into high-resolution digital image files. In this application, the WSI slice images are derived from pathological samples of gastric cancer patients. In this application, the multi-task deep learning model adopts a multi-task architecture that combines a shared backbone network with specific task branches. The specific task branches include a tumor task branch model, a resection margin task branch model, and a lymph node task branch model, which are used to perform tumor diagnosis and identification, resection margin diagnosis and identification, and lymph node diagnosis and identification, respectively. Through task decoupling, efficient and collaborative multi-dimensional analysis of gastric cancer pathology is achieved.
[0023] Specifically, when the tumor task branch model is used to diagnose and identify tumors in the input WSI slice images, the anchor mechanism of the YOLO v8 architecture can be used to pre-calculate the anchor size suitable for gastric cancer tissue (such as 128×128, 256×256 pixels) based on the clustering algorithm to quickly locate the tumor area; then, FPN (feature pyramid network) is used to extract multi-scale features (16×16, 32×32, 64×64 pixel blocks) from the located tumor area to identify different tumor histological types such as tubular adenocarcinoma and signet ring cell carcinoma; finally, based on the convolutional neural network (CNN), morphological characteristics such as nuclear atypia index and glandular structural integrity are analyzed to evaluate the degree of tissue differentiation (classification as high / medium / low grade).
[0024] When performing surgical margin diagnosis and recognition on the input WSI slice image through the surgical margin task branch model, the surgical margin line can be identified through the edge detection algorithm. Then, an attention mechanism is introduced to focus on the suspicious cell area for pixel-level semantic segmentation, identifying cancer cells within 1mm of the surgical margin line, and then determining the residual cancer cell status at the surgical margin. When performing lymph node diagnosis and identification through the lymph node task branch model, the DeepLab v3+ model can be used to segment the lymph nodes and surrounding tissues from the input WSI slice image based on pixel-level semantic segmentation, identify the lymph node contours, and then use the target detection network combined with the focal loss function to detect micrometastasis lesions on the identified lymph node contours to obtain metastasis lesion parameters (such as metastasis lesion coordinates, number, maximum diameter, etc.).
[0025] It should be noted that in this application, the teacher-student model training method, also known as knowledge distillation, is used to obtain the final deployed multi-task deep learning model. Figure 2 , the multi-task deep learning model is trained in the following manner, comprising the following steps: S101. Acquire a WSI slice image sample dataset; the WSI slice image sample dataset includes multiple WSI slice image samples and pathological morphological data annotated for each WSI slice image sample; the pathological morphological data includes tumor histological type and degree of tissue differentiation, surgical margin position and residual cancer cell status at the margin, lymph node contour and metastasis parameters; S102. First, the constructed teacher model is trained based on the WSI slice image sample dataset to obtain a trained teacher model. Then, the knowledge is transferred to the constructed student model for training through the set knowledge distillation parameters to obtain a trained student model, and the student model is used as the final multi-task deep learning model for WSI slice image analysis.
[0026] In step S101, the main task is to construct a high-quality annotated dataset. This involves collecting a large number of WSI slice images from gastric cancer patients and annotating them with pathology experts. This dataset is then used to train the multi-task deep learning model. The annotated information includes tumor histological type and differentiation, surgical margin location and residual cancer cells, lymph node contours, and metastasis parameters.
[0027] In step S102, a multi-task deep learning model is first constructed, a shared backbone network is determined (e.g., ResNet-50 is selected as the basic feature extraction network), and branch models for the tumor task, the resection margin task, and the lymph node task are separately constructed, with the output layer of each branch model defined. Next, a high-performance model (e.g., ResNet-152) is selected as the teacher model, a lightweight student model (e.g., MobileNet v3) is determined, and a mapping relationship between the corresponding layers of the teacher and student models is established to prepare for knowledge distillation. The teacher model is first iteratively trained using the WSI slice image sample dataset. Once performance evaluation passes, the pre-trained ResNet-152 model parameters are saved and used as the teacher model for knowledge distillation. Performance evaluation can include factors such as the accuracy of tumor histological typing and the recall rate of residual cancer cells at the resection margin. The student model is then iteratively trained based on the set knowledge distillation parameters, including the distillation temperature and loss weight, gradually improving its ability to learn the teacher's knowledge and its task performance. After passing performance evaluation, the optimal parameters are saved and used as the final multi-task deep learning model for tumor, margin, and lymph node diagnosis. Furthermore, through the distillation mechanism, the cross-task pathological morphological features and task-related knowledge from the teacher model are compressed into the student model, reducing reliance on high-performance computing resources and enabling real-time analysis of WSI slice images.
[0028] In addition, it should be noted that the WSI slice image sample dataset collected in this application suffers from data imbalance. In gastric cancer pathology diagnosis, data imbalance primarily refers to significant differences in the number of samples across different pathological feature categories. Specifically, the sample size of certain rare pathological types (such as signet ring cell carcinoma and mucinous adenocarcinoma) is extremely small, while the sample size of common types (such as tubular adenocarcinoma) is extremely high. This imbalance can lead to a bias in model training towards the majority class, seriously affecting the diagnostic accuracy of the minority class.
[0029] In one embodiment, to address the aforementioned data imbalance, the number of pathology data samples of different categories in the WSI slice image sample dataset is first counted to determine minority and majority category samples. Oversampling and data augmentation are then performed on the minority category samples. Oversampling primarily involves synthesizing new samples through interpolation, thereby increasing the number of minority and majority category samples. For example, feature vectors are extracted from minority category samples, and the Euclidean distance between each sample and the other samples is calculated. A sample is randomly selected from its k-nearest neighbors, and new samples are randomly generated along the line connecting the two. Data augmentation can generate new samples through techniques such as geometric transformation (rotation, scaling), color transformation, and local occlusion. Undersampling is then performed on the majority category samples. Undersampling primarily involves removing the majority category samples that are closest to the minority category samples, thereby reducing the number of majority category samples and balancing the training data distribution.
[0030] In other embodiments, the loss function during model training can be further set according to the ratio of the number of samples of different categories, wherein a higher weight is given to minority category samples so that the model pays more attention to the classification accuracy of minority category samples during training; a lower weight is given to majority category samples to avoid the model from over-learning the majority category features.
[0031] See the instructions attached Figure 3 In step S2, the multimodal retrieval of pathological information of the text to be retrieved or the image to be retrieved based on the multimodal feature library includes the following steps: S201, extracting image features of the image to be retrieved and performing similarity calculation on the features with the image features in the multimodal feature library to obtain an image list related to the image to be retrieved and corresponding pathological information; S202, extracting semantic features of the text to be retrieved and mapping them to the same space as the image features, and calculating the similarity between the mapped semantic features and the image features in the multimodal feature library, to obtain a list of images related to the text to be retrieved and corresponding pathological information; S203, extracting image features of the image to be retrieved and converting them into text descriptions, and matching the converted text with the pathology text in the multimodal feature library to obtain pathology information related to the image to be retrieved; S204 , extracting semantic features of the text to be retrieved and performing similarity calculation on the semantic features with those in the multimodal feature library to obtain pathological information related to the text to be retrieved.
[0032] That is, multimodal retrieval of pathological information from the text or image to be retrieved based on the multimodal feature library includes four methods: image search, image search, image search, and text search. Specifically, by extracting image features from a large number of WSI slice images and semantic features from a large number of pathological texts, and mapping the extracted image features and semantic features into a unified feature space via bilinear mapping or a multi-layer perceptron, a multimodal feature library containing image-text pairs is constructed.
[0033] In step S201, for image search, the same deep feature extraction operation as that used in constructing the multimodal feature library is performed on the input image to be retrieved, and the cosine similarity or Euclidean distance between the extracted image features and the image features in the multimodal feature library is calculated. The images are sorted in descending order according to the similarity scores, and a list of images related to the image to be retrieved is screened out, along with corresponding pathological information, such as the diagnostic label "tubular adenocarcinoma, well-differentiated".
[0034] In step S202, for text-to-image search, the input text to be retrieved, such as "gastric tubular adenocarcinoma, moderately differentiated, negative resection margin", is first converted into semantic features, and the medical terms and semantic relationships therein are parsed. Then, the text semantic features are projected into the same space as the image features through a pre-trained cross-modal mapping model, and the similarity between the mapped text features and the image features in the multimodal feature library is calculated to find the pathological image that matches the text description, and the relevant pathological information or diagnostic information is annotated.
[0035] In step S203, for image-based text search, after extracting image features from the input image to be searched, an image-to-text mapping model (such as a Transformer-based encoding-decoding model) is used to convert the image features into text description features. The generated text is then used as a query keyword to search the text data in the multimodal feature library to match relevant text reports and analysis content. In step S204, for text search, the input text to be retrieved, such as "cases of gastric cancer lymph node metastasis", is subjected to semantic analysis and feature extraction through a pre-trained language model, and then a semantic matching retrieval using cosine similarity calculation is performed in the text data of the multimodal feature library to find other text information related to the semantics of the input text to assist the attending physician in making a diagnostic reference.
[0036] Therefore, through cross-modal retrieval, the image-text association information of similar cases can be quickly matched, providing a reference for the diagnosis of difficult cases and shortening the analysis or diagnosis time of difficult cases.
[0037] Furthermore, the pathology information obtained through multimodal retrieval is not only used for subsequent pathology report generation, but can also be used to expand the WSI slide image sample dataset and optimize multi-task deep learning models. For example, if a case only has image data but lacks text descriptions, multimodal retrieval can be used to find paired text for similar images, which can serve as labels for the case, thereby assisting multi-task model learning and improving prediction accuracy.
[0038] See the instructions attached Figure 4 In step S3, the obtained pathological morphological data and the pathological information are integrated and filled with content according to a preset template to generate a pathology report, including the following steps: S301. Setting different report templates and corresponding formats according to gastric cancer pathology analysis requirements; S302: Structuring the acquired pathological morphological data and pathological information, and filling the selected report template with content according to a set format to obtain a pathology report; S303: Obtain feedback data from the attending physician, and modify the pathology report according to the feedback data and then archive it.
[0039] Specifically, in step S301, the report template is customized based on the needs of gastric cancer pathology analysis, or the specific requirements of different hospitals and departments. The data entry specifications and semantic requirements of each template are clearly defined to better meet actual work needs. In step S302, the acquired image analysis data and pathology information, such as various pathological morphological data (tumor type: tubular adenocarcinoma, tissue differentiation: moderately differentiated, no residual cancer cells at the resection margin, lymph node metastasis, etc.), are integrated according to the customized report template format and filled in the corresponding positions of the report template.
[0040] In other embodiments, the completed report content can be further verified using pre-set logical rules, such as checking the logical consistency between tumor invasion depth and lymph node metastasis status. Natural language generation technology can also be used to convert the integrated diagnostic information into a fluent, accurate, and standardized text description to generate a complete pathology report. The report content covers various aspects such as basic patient information, morphological description, and pathological diagnosis opinion.
[0041] In step S303, the preliminary pathology report is submitted to the attending physician for manual review and feedback to supplement and modify the report content; after the review is passed, the final pathology report is confirmed and generated for clinical diagnosis and treatment reference and archived.
[0042] It can be seen that the application provides a full-process pathology image analysis method, which realizes the automatic classification and prediction of multiple pathological features by diagnosing and identifying tumors, resection margins, and lymph nodes in WSI slice images, and adopts distillation technology and data imbalance solutions to improve model performance; through interactive image and text retrieval and automatic generation of pathology reports, while realizing the intelligence of the entire process of gastric cancer pathology analysis, it improves the efficiency and accuracy of gastric cancer pathology analysis.
[0043] It should be noted that the full-process pathological image analysis method provided in the application is not limited to gastric cancer pathological analysis. Based on the same technical concept, it can also be used for pan-tumor types, including full-process pathological image analysis of malignant tumors such as colorectal cancer, esophageal cancer, liver cancer, and pancreatic cancer. Taking into account that there are certain similarities and commonalities between different diseases in the field of pathological analysis, as well as some differences, when extending this method to the pathological analysis of other diseases, data adaptation, model structure adjustment, and knowledge base expansion are required. For example, when performing data adaptation, the data preprocessing process may be adjusted for WSI slice images of different diseases, and a large amount of pathological data of the target disease needs to be collected and organized to meet the analysis needs of the new disease. When adjusting the model structure, the pathological characteristics and diagnostic tasks of different diseases may be different, so the network structure of the multi-task deep learning model needs to be appropriately adjusted. When expanding the knowledge base, the data scope in the knowledge base is extended to pathological images, text reports, clinical information and other content related to the target disease. By integrating and associating new and old disease data, the semantic information of the knowledge base is enriched, so that the interactive image and text retrieval function can provide more comprehensive and accurate reference information for the diagnosis of different diseases.
[0044] Based on the same inventive concept, a full-process pathology image analysis device is also provided in an embodiment of the present application. Since the principle of solving the problem by the device in the embodiment of the present application is similar to the above-mentioned full-process pathology image analysis method in the embodiment of the present application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.
[0045] As the instruction manual Figure 5 As shown, the embodiment of the present application also provides a full-process pathological image analysis device, the device comprising: An image analysis module 501 is configured to construct a multimodal feature library containing image-text pairs based on image features of a large number of WSI slice images and semantic features of a large number of pathology texts, and perform multimodal retrieval of pathology information on the text to be retrieved or the image to be retrieved based on the multimodal feature library; Pathology information retrieval module 502, for performing pathology information retrieval by image search, image search, image search, and text search based on the constructed multimodal feature library; The pathology report generating module 503 is used to integrate the obtained pathological morphology data and the pathological information, and fill in the content according to a preset template to generate a pathology report.
[0046] In some embodiments, the image analysis module 501 performs tumor diagnosis and identification, margin diagnosis and identification, and lymph node diagnosis and identification on the acquired WSI slice image, including: first locating the tumor area from the input WSI slice image based on the anchor frame mechanism through the tumor task branch model, and then using the feature pyramid network FPN to extract multi-scale features of the tumor area to identify different tumor histological types, and extracting morphological features based on the convolutional neural network CNN to evaluate the degree of tissue differentiation; first locating the surgical margin line from the input WSI slice image based on the edge detection algorithm through the margin task branch model, and then identifying cancer cells within the set range of the surgical margin line based on pixel-level semantic segmentation to determine the residual status of cancer cells at the margin; first identifying the lymph node contour from the input WSI slice image based on pixel-level semantic segmentation through the lymph node task branch model, and then performing micro-metastasis detection on the lymph node contour through the target detection network combined with the focal loss function to obtain metastasis parameters.
[0047] In some embodiments, the apparatus further comprises: A training module is used to train the multi-task deep learning model by using knowledge distillation, including: obtaining a WSI slice image sample dataset; the WSI slice image sample dataset includes multiple WSI slice image samples and pathological morphological data annotated for each WSI slice image sample; the pathological morphological data includes the tumor histological type and its tissue differentiation degree, the position of the surgical margin line and the residual status of the surgical margin cancer cells, the lymph node contour and its metastasis parameters; firstly training the constructed teacher model based on the WSI slice image sample dataset to obtain a trained teacher model, and then transferring the knowledge to the constructed student model for training through the set knowledge distillation parameters to obtain the trained student model, and using the trained student model as the final multi-task deep learning model to perform WSI slice image analysis.
[0048] In some embodiments, the training module trains the constructed teacher model based on the WSI slice image sample dataset, including: counting the number of samples of different categories of pathological morphological data in the WSI slice image sample dataset, and determining minority category samples and majority category samples; processing the minority category samples and the majority category samples according to a set balancing strategy; wherein, data enhancement and oversampling are performed on the minority category samples, and undersampling is performed on the majority category samples, and the loss function during model training is set according to the ratio of the number of samples of different categories.
[0049] In some embodiments, the pathological information retrieval module 502 performs multi-modal retrieval of pathological information based on the multi-modal feature library for the text to be retrieved or the image to be retrieved, including: extracting image features of the image to be retrieved and performing similarity calculation with image features in the multi-modal feature library to obtain an image list related to the image to be retrieved and corresponding pathological information; extracting semantic features of the text to be retrieved and mapping them to the same space as the image features, and calculating the similarity between the mapped semantic features and the image features in the multi-modal feature library to obtain an image list related to the text to be retrieved and corresponding pathological information; extracting image features of the image to be retrieved and converting them into text description, and matching the converted text with pathological texts in the multi-modal feature library to obtain pathological information related to the image to be retrieved; extracting semantic features of the text to be retrieved and performing similarity calculation with semantic features in the multi-modal feature library to obtain pathological information related to the text to be retrieved.
[0050] In some embodiments, the pathological report generation module 503 integrates the obtained pathological morphology data and pathological information, and fills in the content according to a preset template to generate a pathological report, including: setting different report templates and their corresponding formats according to the pathological analysis requirements of gastric cancer; structurally processing the obtained pathological morphology data and pathological information, and filling in the content of the selected report template according to the set format to obtain a pathological report; obtaining feedback data of the attending physician, and modifying and archiving the pathological report according to the feedback data.
[0051] In some embodiments, the device further comprises: An optimization module for optimizing the multi-task deep learning model based on the pathological information obtained from the multi-modal feature library.
[0052] The full-process pathology image analysis device described in this application acquires full-field digital WSI slice images from gastric cancer patients through an image analysis module, and analyzes the acquired WSI slice images using a multi-task deep learning model to obtain pathological morphological data; wherein the multi-task deep learning model adopts a multi-task architecture including a tumor task branch model, a resection margin task branch model, and a lymph node task branch model, for performing tumor diagnosis and identification, resection margin diagnosis and identification, and lymph node diagnosis and identification on the acquired WSI slice images; the pathology information retrieval module constructs a multimodal feature library containing image-text pairs based on the image features of a large number of WSI slice images and the semantic features of a large number of pathology texts, and performs multimodal retrieval of pathological information on the text to be retrieved or the image to be retrieved based on the multimodal feature library; the pathology report generation module integrates the obtained pathological morphological data and the pathological information, and fills the content according to a preset template to generate a pathology report. Thus, through pathology image analysis based on multi-task deep learning, cross-modal interactive image-text retrieval, and automatic generation of pathology reports, the full process of pathology analysis is intelligentized, improving the efficiency and accuracy of gastric cancer pathology analysis.
[0053] Based on the same concept of the present invention, as shown in the attached specification Figure 6 As shown, an embodiment of the present application provides a structure of an electronic device 600, which includes: at least one processor 601, at least one network interface 604 or other user interface 603, a memory 605, and at least one communication bus 602. The communication bus 602 is used to implement connection and communication between these components. The electronic device 600 optionally includes a user interface 603, including a display (e.g., a touch screen, LCD, CRT, holographic imaging (Holographic) or projector (Projector), etc.), a keyboard or a pointing device (e.g., a mouse, trackball (trackball), touchpad or touch screen, etc.).
[0054] The memory 605 may include a read-only memory and a random access memory, and provides instructions and data to the processor 601. A portion of the memory 605 may also include a non-volatile random access memory (NVRAM).
[0055] In some embodiments, the memory 605 stores the following elements, executable modules, or data structures, or a subset or extended set thereof: Operating system 6051, including various system programs used to implement various basic services and handle hardware-based tasks; The application module 6052 includes various application programs, such as a launcher, a media player, and a browser, and is used to implement various application services.
[0056] In an embodiment of the present application, by calling a program or instruction stored in the memory 605, the processor 601 is used to execute steps of a full-process pathology image analysis method.
[0057] The present application also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program executes steps in a full-process pathology image analysis method.
[0058] Specifically, the storage medium can be a general storage medium, such as a mobile disk, a hard disk, etc. When the computer program on the storage medium is run, the entire process of pathological analysis can be intelligentized, thereby improving analysis efficiency and accuracy.
[0059] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interface, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0060] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0061] In addition, each functional unit in the embodiments provided in the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0062] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the existing technology, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.
[0063] Finally, it should be noted that the above embodiments are only specific implementation methods of the present application, which are used to illustrate the technical solutions of the present application, rather than to limit them. The scope of protection of the present application is not limited thereto. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above embodiments within the technical scope disclosed in the present application, or replace some of the technical features therein with equivalents. However, these modifications, changes, or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application. They should all be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A full-process pathological image analysis method, characterized in that: The method comprises the following steps: Acquire full-field digital WSI slice images from gastric cancer patients, and analyze the acquired WSI slice images using a multi-task deep learning model to obtain pathological morphological data; wherein the multi-task deep learning model adopts a multi-task architecture including a tumor task branch model, a resection margin task branch model, and a lymph node task branch model, and is used to perform tumor diagnosis and identification, resection margin diagnosis and identification, and lymph node diagnosis and identification on the acquired WSI slice images; Based on the image features of a large number of WSI slice images and the semantic features of a large number of pathology texts, a multimodal feature library containing image-text pairs is constructed, and multimodal retrieval of pathology information is performed on the text to be retrieved or the image to be retrieved based on the multimodal feature library; The obtained pathological morphological data and the pathological information are integrated, and the content is filled in according to a preset template to generate a pathology report.
2. A full-process pathological image analysis method according to claim 1, characterized in that: The step of performing tumor diagnosis and identification, margin diagnosis and identification, and lymph node diagnosis and identification on the acquired WSI slice image comprises the following steps: The tumor task branch model first locates the tumor region from the input WSI slice image based on the anchor frame mechanism, then uses the feature pyramid network (FPN) to extract multi-scale features of the tumor region to identify different tumor histological types, and uses the convolutional neural network (CNN) to extract morphological features to evaluate the degree of tissue differentiation; The surgical margin task branch model first locates the surgical margin line from the input WSI slice image based on the edge detection algorithm, and then identifies cancer cells within the set range of the surgical margin line based on pixel-level semantic segmentation to determine the residual state of cancer cells at the surgical margin; The lymph node task branch model first identifies the lymph node contour from the input WSI slice image based on pixel-level semantic segmentation, and then uses the target detection network combined with the focal loss function to detect micrometastasis lesions on the lymph node contour to obtain metastasis lesion parameters.
3. A full-process pathological image analysis method according to claim 2, characterized in that: in, The multi-task deep learning model is trained using knowledge distillation, including the following steps: Acquire a WSI slice image sample dataset; the WSI slice image sample dataset includes multiple WSI slice image samples and pathological morphological data annotated for each WSI slice image sample; the pathological morphological data includes tumor histological type and degree of tissue differentiation, surgical margin position and residual cancer cell status at the margin, lymph node contour and metastasis parameters; First, the constructed teacher model is trained based on the WSI slice image sample dataset to obtain a trained teacher model. Then, the knowledge is transferred to the constructed student model for training through the set knowledge distillation parameters to obtain a trained student model, which is used as the final multi-task deep learning model for WSI slice image analysis.
4. A full-process pathological image analysis method according to claim 3, characterized in that: The training of the constructed teacher model based on the WSI slice image sample dataset comprises the following steps: Counting the number of samples of different categories of pathological morphology data in the WSI slice image sample dataset, and determining minority category samples and majority category samples; The minority class samples and the majority class samples are processed according to a set balancing strategy; wherein, data enhancement and oversampling are performed on the minority class samples, undersampling is performed on the majority class samples, and the loss function during model training is set according to the ratio of the number of samples of different classes.
5. A full-process pathological image analysis method according to claim 1, characterized in that: The multimodal retrieval of pathological information on the text to be retrieved or the image to be retrieved based on the multimodal feature library includes the following steps: Extracting image features of the image to be retrieved and performing similarity calculation on the features with the image features in the multimodal feature library to obtain an image list related to the image to be retrieved and corresponding pathological information; Extracting semantic features of the text to be retrieved and mapping them to the same space as the image features, and calculating the similarity between the mapped semantic features and the image features in the multimodal feature library to obtain a list of images related to the text to be retrieved and the corresponding pathological information; Extracting image features of the image to be retrieved and converting them into text descriptions, and matching the converted text with pathology text in the multimodal feature library to obtain pathology information related to the image to be retrieved; The semantic features of the text to be retrieved are extracted and similarity calculation is performed between the semantic features in the multimodal feature library to obtain pathological information related to the text to be retrieved.
6. A full-process pathological image analysis method according to claim 1, characterized in that: The step of integrating the obtained pathological morphological data and the pathological information and filling the content according to a preset template to generate a pathology report includes the following steps: Set different report templates and their corresponding formats according to the needs of gastric cancer pathology analysis; Structuring the acquired pathological morphological data and pathological information, and filling the selected report template with content according to a set format to obtain a pathology report; Obtain feedback data from the attending physician, and modify the pathology report based on the feedback data and then archive it.
7. A full-process pathological image analysis method according to claim 3, characterized in that: The method further comprises the following steps: The multi-task deep learning model is optimized based on the pathological information obtained from the multimodal feature library.
8. A full-process pathological image analysis device, characterized in that: The device comprises: An image analysis module is configured to acquire full-field digital WSI slice images from gastric cancer patients and analyze the acquired WSI slice images using a multi-task deep learning model to obtain pathological morphological data; wherein the multi-task deep learning model adopts a multi-task architecture including a tumor task branch model, a resection margin task branch model, and a lymph node task branch model, and is configured to perform tumor diagnosis and identification, resection margin diagnosis and identification, and lymph node diagnosis and identification on the acquired WSI slice images; A pathology information retrieval module is used to construct a multimodal feature library containing image-text pairs based on the image features of a large number of WSI slice images and the semantic features of a large number of pathology texts, and to perform multimodal retrieval of pathology information on the text to be retrieved or the image to be retrieved based on the multimodal feature library; The pathology report generating module is used to integrate the obtained pathological morphology data and the pathological information, and fill in the content according to the preset template to generate a pathology report.
9. An electronic device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor and the memory communicate via the bus. When the machine-readable instructions are executed by the processor, the steps of a full-process pathology image analysis method as described in any one of claims 1 to 7 are performed.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, executes the steps of a full-process pathological image analysis method as described in any one of claims 1 to 7.
Citation Information
Cited By
Pathological image cervical cancer screening model construction method based on liquid-based cytology
CN121437462A