Deep Learning-Based Bladder Cancer Gene Mutation Prediction System

Through the deep learning-based bladder oncogene mutation prediction system, the inconsistency problem of bladder cancer pathological diagnosis is solved, and precise gene mutation diagnosis and individualized treatment guidance are achieved.

CN119580836BActive Publication Date: 2025-07-04BEIJING THOROUGH FUTURE INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411457307.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-18
Publication Date
2025-07-04
Estimated Expiration
2044-10-18

AI Technical Summary

Technical Problem

The existing pathological diagnosis of bladder cancer depends on the professional knowledge and experience of pathologists, resulting in inconsistent diagnosis and cannot meet the needs of patients for precise and individualized treatment.

Method used

A deep learning-based bladder oncogene mutation prediction system is used to predict bladder oncogene expression, methylation and mutation status through histopathological image acquisition, preprocessing and deep learning models to generate auxiliary diagnostic reports.

Benefits of technology

It improves the accuracy and efficiency of bladder oncogene mutation diagnosis, helps pathologists to pre-judgment the genetic changes represented by pathological morphology, and guides individualized treatment strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119580836B_ABST
    Figure CN119580836B_ABST
Patent Text Reader

Abstract

The present invention provides a bladder cancer gene mutation prediction system based on deep learning, comprising: a histopathological image acquisition module for acquiring original tissue histopathological images to be measured; a preprocessing module for preprocessing the original tissue histopathological images to be measured by using the HE staining method and a whole slide imaging scanner to obtain a WSI of a bladder tissue with HE staining; a cloud server module for calling a deep learning prediction model for bladder cancer to respectively predict the gene expression status, gene methylation status and gene mutation status of bladder cancer for the WSI, and generating corresponding prediction results; and a report generation module for automatically generating an auxiliary diagnosis report for bladder cancer gene mutation prediction according to the corresponding prediction results and feeding it back to a specified terminal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of bladder cancer auxiliary diagnosis, and particularly relates to a bladder cancer gene mutation prediction system based on deep learning. Background Art

[0002] Bladder cancer is the most common malignant tumor of the urinary system. Bladder cancer ranks ninth among the most common tumors globally [PMID: 36633525, PMID: 33433946]. It ranks first in the incidence of urogenital tumors in China, and in the West, its incidence is second only to prostate cancer, ranking second. In 2020, there were an estimated 573,275 new cases and 212,536 deaths globally [PMID: 33538338].

[0003] The pathological types of bladder cancer include urothelial carcinoma, squamous cell carcinoma, adenocarcinoma. Other rare types include bladder clear cell carcinoma, bladder small cell carcinoma, and bladder carcinoid. Among them, the most common is bladder urothelial carcinoma, accounting for more than 90% of the total number of bladder cancer patients. Pathological diagnosis is still the gold standard for the diagnosis of bladder cancer at present. However, traditional pathological diagnosis severely relies on the professional knowledge and diagnostic experience of pathologists, and the subjectivity of pathologists leads to diagnostic inconsistencies. At the same time, a simple qualitative diagnosis based on tissue morphology cannot meet the needs of precise and individualized treatment for patients. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, the present invention provides a bladder cancer gene mutation prediction system based on deep learning to solve the above problems.

[0005] A bladder cancer gene mutation prediction system based on deep learning includes: a histopathological image acquisition module for acquiring original tissue pathological images to be measured; a preprocessing module for preprocessing the original tissue pathological images to be measured by using HE staining method and a whole slide imaging scanner to obtain a WSI of bladder tissue with HE staining; a cloud server module for calling a deep learning prediction model of bladder cancer to respectively predict the gene expression state, gene methylation state, and gene mutation state of bladder cancer for the WSI, and generating corresponding prediction results; a report generation module for automatically generating an auxiliary diagnosis report for bladder cancer gene mutation prediction according to the corresponding prediction results and feeding it back to a specified terminal.

[0006] As an embodiment of the present invention, the histopathological image acquisition module is connected to a preset urinary system image acquisition device for obtaining the original tissue pathological images to be measured collected by the preset urinary system image acquisition device.

[0007] As an embodiment of the present invention, the preprocessing module performs the following operations: stain the bladder tissue to be tested according to the HE staining method to obtain a HE-stained bladder tissue; and obtain a WSI of the HE-stained bladder tissue by using a whole slide imaging scanner based on the original histopathological image of the tissue to be tested stained by the HE staining method.

[0008] As an embodiment of the present invention, the training process of the deep learning prediction model for bladder cancer includes the following steps: obtain training samples; construct a deep learning framework based on a convolutional neural network, and train the deep learning framework according to the training samples until the preset conditions are met to end the training and generate a deep learning prediction model for bladder cancer.

[0009] As an embodiment of the present invention, obtaining training samples includes: obtaining training WSIs of a plurality of HE-stained bladder tissues in the public TCGA; performing label marking on the training WSIs, and filtering the training WSIs to obtain a plurality of WSIs that only retain the WSI images with gene expression, gene methylation, and gene mutation labels; dividing the plurality of WSI images to obtain a validation WSI set, a training WSI set, and a test WSI set; and based on a magnification of 400 times, dividing the images of all WSI sets into image patches of 320×320 pixels and filtering the background to obtain a second validation WSI set, a second training WSI set, and a second test WSI set.

[0010] As an embodiment of the present invention, a deep learning-based gene mutation prediction system for bladder cancer further includes using a two-stage pre-analysis method to accelerate the training process. The first-stage acceleration process is: based on a magnification of 100 times, dividing the images of all WSI sets into image patches of 320×320 pixels and filtering the background to obtain a third validation WSI set, a third training WSI set, and a third test WSI set for constructing an analysis data set; and using a pre-trained RestNet-50 model on ImageNet to convert each image patch in the analysis data set into a 2048-dimensional vector as a feature extractor.

[0011] As an embodiment of the present invention, the second-stage acceleration process is: using the 3-fold cross-validation method in the k-fold cross-validation method to evaluate the performance of Xgboost trained on the feature vectors to obtain an evaluation result; wherein, in each fold, randomly select 200 WSIs in the analysis data set for training, and at the same time select 100 WSIs for testing; and based on the RestNet-50 model, re-construct a classification model on the second validation WSI set, the second training WSI set, and the second test WSI set, and combine the evaluation results to provide detailed classification information for each selected gene for marking.

[0012] As an embodiment of the present invention, the report generation module performs the following operations: obtaining the corresponding prediction result, extracting the corresponding original case data from a preset database according to the corresponding prediction result, where the original case data is the case data participated in the prediction by the patient corresponding to the current prediction result during this prediction; extracting the effective initial text corpus in the original case data based on a regular expression, where each initial text corpus uniquely corresponds to one prediction result; sorting all the initial text corpora according to the text length and counting the text length to construct a text length relationship table; performing entity extraction of medical entity types on each initial text corpus to determine the entity type of each initial text corpus, and associating the entity types of all the initial text corpora under each prediction result according to a preset entity association relationship to obtain the entity association relationship of each initial text corpus; determining an initial report template according to the longest text length in the text length relationship table, and performing a merging and filling process on any number of associated initial text corpora based on the entity association relationship of the initial text corpora under each prediction result to obtain a report text to be generated; and filling the report text to be generated and the prediction result into the initial report template according to the initial report text filling logic corresponding to the initial report template to generate an auxiliary diagnosis report for bladder cancer gene mutation prediction.

[0013] As an embodiment of the present invention, a bladder cancer gene mutation prediction system based on deep learning further includes: obtaining the user's diagnosis intention uploaded by a specified terminal, where the user's diagnosis intention includes various medical treatment subjects, and there are multiple user's diagnosis intentions; analyzing all the user's diagnosis intentions uploaded this time to determine the diagnosis field that the user needs auxiliary diagnosis for this time, and removing the irrelevant content in the auxiliary diagnosis report for bladder cancer gene mutation prediction based on the diagnosis field to generate a corresponding precise auxiliary diagnosis report and feedback it to the specified terminal.

[0014] As an embodiment of the present invention, analyzing all the user's diagnosis intentions uploaded this time to determine the diagnosis field that the user needs auxiliary diagnosis for this time includes: establishing a diagnosis intention database, where the diagnosis intention database stores several diagnosis intentions with marked diagnosis fields, each diagnosis intention is marked with an association value with different diagnosis fields, and each diagnosis intention is related to several diagnosis fields at the same time; generating a first association value m between each user's diagnosis intention and different diagnosis fields based on the diagnosis intention database, where there is a first association value m between each user's diagnosis intention and each diagnosis field, and based on all the first association values m, counting several second association values of the user's diagnosis intentions corresponding to any diagnosis field, and the second association values include m1, m2...m n, where n is the total number of user diagnosis intents, and the user diagnosis intents corresponding to all the second association values exceeding the preset statistical threshold p are screened out as the target diagnosis intents, and the number k of target diagnosis intents in each diagnosis field is calculated. i , where i is the number of diagnosis fields; the number k of target diagnosis intents in each diagnosis field is counted. i The proportion j of the total number of user diagnosis intents i ; a condition satisfaction value g is preset in advance. When k i and j i The product of is not less than g, take the product maximum value of k i and j i The corresponding diagnosis field is output as the target diagnosis field to determine the diagnosis field that the user needs assisted diagnosis this time.

[0015] The beneficial effects of the present invention are:

[0016] The present invention provides a deep learning-based bladder cancer gene mutation prediction system, which uses deep learning (DL) to extract pathological morphological features from hematoxylin and eosin (H&E) images that can accurately predict gene abnormalities in bladder cancer patients, establish the connection between bladder cancer gene changes and pathological morphology, help pathologists pre-judge and efficiently the gene changes represented by specific pathological morphology, and guide the next diagnosis and treatment strategy, so as to solve the problem that the existing simple tissue morphology-based qualitative diagnosis cannot meet the needs of patients' precise and individualized treatment.

[0017] Other features and advantages of the present invention will be described in the following specification, and, in part, will become apparent from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained by the structure specifically pointed out in the written specification and the drawings.

[0018] The technical solution of the present invention will be further described in detail below through the drawings and embodiments. Description of the Drawings

[0019] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the drawings:

[0020] Figure 1 is a schematic diagram of the system module of a deep learning-based bladder cancer gene mutation prediction system in an embodiment of the present invention;

[0021] Figure 2 is a flowchart for generating an accurate assisted diagnosis report in a deep learning-based bladder cancer gene mutation prediction system in an embodiment of the present invention;

[0022] Figure 3 This is the diagnostic field determination flowchart for generating a flowchart in a bladder cancer gene mutation prediction system based on deep learning in an embodiment of the present invention. Specific Embodiments

[0023] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only for the purpose of illustrating and explaining the present invention, and are not intended to limit the present invention.

[0024] Please refer to Figure 1 , a bladder cancer gene mutation prediction system based on deep learning, comprising: a histopathological image acquisition module 1 for acquiring original histopathological images to be measured; a preprocessing module 2 for preprocessing the original histopathological images to be measured by using the HE staining method and a whole slide imaging scanner to obtain WSI (whole slide images) of bladder tissue with HE staining; a cloud server module 3 for calling a deep learning prediction model for bladder cancer to respectively perform predictions on the gene expression status, gene methylation status, and gene mutation status of bladder cancer on the WSI, and generating corresponding prediction results; a report generation module 4 for automatically generating an auxiliary diagnostic report for bladder cancer gene mutation prediction according to the corresponding prediction results and feeding it back to a specified terminal;

[0025] The working principle of the above technical solution is as follows: The present invention provides a bladder cancer gene mutation prediction system based on deep learning. The system includes a histopathological image acquisition module, a preprocessing module, a cloud server module, and a report generation module. Among them, the histopathological image acquisition module acquires original histopathological images to be measured; the preprocessing module preprocesses the original histopathological images to be measured by using the HE staining method and a whole slide imaging scanner to obtain WSI of bladder tissue with HE staining; the cloud server module calls a deep learning prediction model for bladder cancer to respectively perform predictions on the gene expression status, gene methylation status, and gene mutation status of bladder cancer on the WSI, and generates corresponding prediction results; the report generation module automatically generates an auxiliary diagnostic report for bladder cancer gene mutation prediction according to the corresponding prediction results and feeds it back to a specified terminal;

[0026] The beneficial effects of the above technical solution are as follows: Through the above technical solution, deep learning (DL) is used to extract pathological morphological features from hematoxylin and eosin (H&E) images that can accurately predict gene abnormalities in bladder cancer patients, establish the connection between gene changes in bladder cancer and pathological morphology, help pathologists pre-judge and efficiently judge the gene changes represented by specific pathological morphology in advance, guide the next diagnosis and treatment strategy, and solve the problem that the existing simple qualitative diagnosis based on tissue morphology cannot meet the needs of precise and individualized treatment of patients.

[0027] In one embodiment, the histopathology image acquisition module is connected to a preset urinary system image acquisition device and is configured to acquire an original histopathology image of a tissue to be tested collected by the preset urinary system image acquisition device.

[0028] In one embodiment, the preprocessing module performs the following operations: staining the bladder tissue to be tested according to the HE staining method to obtain an HE-stained bladder tissue; and obtaining a WSI of the HE-stained bladder tissue by using a whole slide imaging scanner based on the original histopathology image stained by the HE staining method.

[0029] In one embodiment, the training process of the bladder cancer deep learning prediction model includes the following steps: obtaining training samples; constructing a deep learning framework based on a convolutional neural network, training the deep learning framework according to the training samples until a preset condition is met to end the training, and generating a bladder cancer deep learning prediction model.

[0030] It should be noted that the outputs in the bladder cancer deep learning prediction model include a deep learning prediction of the gene expression status of bladder cancer, a deep learning prediction of the gene methylation status of bladder cancer, and a deep learning prediction of the gene mutation status of bladder cancer.

[0031] In a specific embodiment, we first used a deep learning framework to predict the gene expression status of bladder cancer through histopathology images; we collected 14,231 gene expression data points matched to each WSI patient; to reduce the scope of analysis of predictable genes, we performed a two-stage pre-analysis method on all genes; in the cross-validation stage, the performance of the genes was measured by the area under the curve (AUC) of the average sliding level of all three folds; the slope prediction was calculated by averaging the prediction probabilities of all tiles of a WSI. We trained a model from scratch for each selected gene using ResNet-50 on a 400x magnification dataset to obtain the final prediction results; finally, in the test group, the top ten genes - CALD1, FGF7, ITGA5, PLN, SFRP2, VSIG4, PDLIM3, CYBRD1, SHISAL1, and CPXM2 were predicted from the pathological images, with AUCs ranging from 0.854 to 0.951.

[0032] Then, based on the same method, the gene methylation status of bladder cancer was predicted. In a specific embodiment, 14,231 gene methylation data points matched to each WSI patient were collected and processed by binning at the median; using a processing pipeline, the top 10 genes - EPP, PRICKLE4, ANK3, CABLES2, ZSCAN2, KDM4B, ZNF556, ZNRF1, MMEL1, and PLGLB2 were identified, with AUCs ranging from 0.855 to 0.903.

[0033] The same method was used to predict the gene mutation status of bladder cancer by deep learning. In a specific embodiment, to ensure there are sufficient mutant samples in training, we only used genes with a mutation rate exceeding 10%, thus reducing the number of genes to be analyzed to an acceptable range. We directly trained a model from scratch for each selected gene using ResNet-50 on a 400-fold magnified dataset. The top five mutant genes predicted by deep learning were COL11A1, FGFR3, SRRM2, LRP1B, and MUC17, with AUC values ranging from 0.664 to 0.775.

[0034] In one embodiment, obtaining the training samples includes: obtaining a number of training whole slide images (WSIs) of H&E-stained bladder tissues in the publicly available TCGA; performing label marking on the training WSIs and filtering the training WSIs to obtain a number of WSIs that only retain the images with gene expression, gene methylation, and gene mutation labels; dividing the number of WSIs to obtain a validation WSIs set, a training WSIs set, and a test WSIs set; based on a 400-fold magnification, dividing all WSIs sets into image patches of 320×320 pixels and filtering the background to obtain a second validation WSIs set, a second training WSIs set, and a second test WSIs set; where all WSIs sets include the validation WSIs set, the training WSIs set, and the test WSIs set.

[0035] The working principle and beneficial effects of the above technical solution are as follows: We used the WSIs of H&E-stained bladder tissues in the publicly available TCGA (The Cancer Genome Atlas) to develop and evaluate our model. We filtered out the WSIs with poor quality and only retained the images with gene expression, gene methylation, and gene mutation labels. After screening, we obtained 400 WSIs, of which 250 were used for training, 50 for validation, and 103 as the test set. Since the size of the WSIs was too large to be directly used as the input of the neural network, we divided the WSIs into image patches of 320×320 pixels at a 400-fold magnification and filtered out the background. Finally, we obtained a training set (second training WSIs set) containing 150,000 image patches, a validation set (second validation WSIs set) containing 50,000 image patches, and a test set (second test WSIs set) containing 100,000 image patches.

[0036] In one embodiment, a deep learning-based bladder cancer gene mutation prediction system further includes using a two-stage pre-analysis method to accelerate the training process. The specific acceleration process is as follows: In the first stage, based on a 100-fold magnification, all WSI sets are divided into image patches of 320×320 pixels, and the background is filtered to obtain a third validation WSI set, a third training WSI set, and a third test WSI set for constructing an analysis data set. Among them, all WSI sets include a validation WSI set, a training WSI set, and a test WSI set; the RestNet-50 model pre-trained on ImageNet is used to convert each image patch in the analysis data set into a 2048-dimensional vector as a feature extractor; in the second stage, the 3-fold cross-validation method in the k-fold cross-validation method is used to evaluate the performance of Xgboost trained on the feature vectors to obtain an evaluation result; among them, in each fold, 200 WSIs in the analysis data set are randomly selected for training, and 100 WSIs are selected for testing; a classification model is re-constructed on the second validation WSI set, the second training WSI set, and the second test WSI set based on the RestNet-50 model, and detailed classification information is provided and marked for each selected gene in combination with the evaluation result;

[0037] The working principle and beneficial effects of the above technical solution are as follows: To improve the analysis efficiency, we propose a two-stage pre-analysis method to accelerate the process; in the first stage, we use image patches of 320×320 pixels in a similar way to construct a smaller data set at a 100-fold magnification; then, the RestNet-50 model pre-trained on ImageNet is used to convert each image patch into a 2048-dimensional vector as a feature extractor; these feature vectors are then used as the input for the next stage, with the advantage of fast analysis; in the second stage, we use 3-fold cross-validation to evaluate the performance of Xgboost trained on the feature vectors; in each fold, 200 WSIs are randomly selected for training, while 100 WSIs are selected for testing to ensure that there are no overlapping WSIs between the data sets; the first few genes selected by the two-stage pre-analysis method are considered predictable; finally, we use RestNet-50 to construct a classification model from scratch on a 400-fold magnified data set to provide more detailed information for each selected gene; based on this framework, we provide the results of three gene-related prediction tasks, including gene expression, gene methylation, and gene mutation status.

[0038] In one embodiment, the report generation module performs the following operations: obtaining corresponding prediction results, extracting corresponding original case data from a preset database according to the corresponding prediction results, where the original case data is the case data participated in the prediction by the patient corresponding to the current prediction result during this prediction; extracting valid initial text corpora from the original case data based on regular expressions, where each initial text corpus uniquely corresponds to one prediction result; sorting all the initial text corpora according to the text length and counting the text length to construct a text length relationship table; performing entity extraction of medical entity types on each initial text corpus to determine the entity type of each initial text corpus, and associating the entity types of all the initial text corpora under each prediction result according to a preset entity association relationship to obtain the entity association relationship of each initial text corpus; determining an initial report template according to the longest text length in the text length relationship table, and performing a merging and filling process on any number of associated initial text corpora based on the entity association relationship of the initial text corpora under each prediction result to obtain a report text to be generated; based on the initial report text filling logic corresponding to the initial report template, corresponding the report text to be generated and the prediction results and filling them into the initial report template to generate an auxiliary diagnosis report for bladder cancer gene mutation prediction;

[0039] The working principle of the above technical solution is as follows: Obtain the prediction results of the gene expression status, gene methylation status, and gene mutation status of bladder cancer. Based on the above multiple prediction results, extract the corresponding original case data from a preset database. Here, the original case data is the case data of the patient corresponding to the current prediction result when participating in the prediction this time. Preferably, standard data for comparison is attached to each piece of case data. Then, extract effective initial text corpora from the case data participating in the prediction this time through regular expressions. Each initial text corpus uniquely corresponds to one prediction result. Then, sort the initial text corpora according to the text length and count the text length to obtain the sorting result and the counting result, and construct a text length relationship table. Perform entity extraction of medical entity types for each initial text corpus to determine the medical entity type of each initial text corpus. There are multiple medical entity types, such as diseases, clinical manifestations, drugs, medical devices, medical procedures, the body, medical test items, microorganisms, departments, etc. According to the preset entity association relationship, associate the entity types of all initial text corpora under each prediction result to obtain the entity association relationship of each initial text corpus. Determine the initial report template according to the longest text length in the text length relationship table. Based on the entity association relationship of the initial text corpora under each prediction result, perform a merging and filling process on any number of associated initial text corpora to obtain the text of the report to be generated. Here, any number is at least 2, and the entity types corresponding to more than 2 initial text corpora must be pairwise associated. For example, if there are three entity types A, B, and C, then A and B must be able to form an entity pair, A and C must be able to form an entity pair, and B and C must be able to form an entity pair. And the text length of the merged text corpus after merging and filling is less than the longest text length. Based on the filling logic of the initial report text corresponding to the initial report template, fill the text of the report to be generated and the prediction results into the initial report template to generate an auxiliary diagnostic report for bladder cancer gene mutation prediction;

[0040] The beneficial effects of the above technical solution are as follows: Through the above technical solution, an auxiliary diagnostic report is generated by means of entity association and entity extraction, which is beneficial to improving the report generation speed. At the same time, through the entity association method, the user's understanding speed of the report is accelerated.

[0041] Please refer to Figure 2 、 Figure 3, in one embodiment, a deep learning-based bladder cancer gene mutation prediction system further includes: S101. Obtain the user's diagnosis intention uploaded by a specified terminal, where the user's diagnosis intention includes various medical diagnosis subjects, and among them, there are multiple user's diagnosis intentions; S102. Analyze all the user's diagnosis intentions uploaded this time to determine the diagnosis field that the user needs auxiliary diagnosis for this time; S103. Remove the irrelevant content in the bladder cancer gene mutation prediction auxiliary diagnosis report based on the diagnosis field, and generate a corresponding precise auxiliary diagnosis report and feedback it to the specified terminal; among them, analyzing all the user's diagnosis intentions uploaded this time to determine the diagnosis field that the user needs auxiliary diagnosis for this time includes: S201. Establish a diagnosis intention database, where the diagnosis intention database stores several diagnosis intentions with marked diagnosis fields, each diagnosis intention is marked with an association value with different diagnosis fields, and each diagnosis intention is related to several diagnosis fields at the same time; S202. Based on the diagnosis intention database, generate the first association value m between each user's diagnosis intention and different diagnosis fields, where there is a first association value m between each user's diagnosis intention and each diagnosis field; S203. Based on all the first association values m, count several second association values of the user's diagnosis intentions corresponding to any diagnosis field; The second association values include m1, m2...m n , where n is the total number of user's diagnosis intentions; S204. Screen out the user's diagnosis intentions corresponding to all the second association values exceeding the preset statistical threshold p as the target diagnosis intentions, and calculate the number k of the target diagnosis intentions in each diagnosis field i , where i is the number of diagnosis fields; S205. Count the number k of the target diagnosis intentions in each diagnosis field i The proportion j of the total number of user's diagnosis intentions i ; S206. Preset a condition satisfaction value g. When the product of k i and j i is not less than g, take the product maximum value of k i and j i The corresponding diagnosis field is output as the target diagnosis field to determine the diagnosis field that the user needs auxiliary diagnosis for this time;

[0042] The working principle of the above technical solution is as follows: Obtain the user's diagnostic intention uploaded by the specified terminal. The user's diagnostic intention includes various medical treatment subjects, such as surgery, internal medicine, etc. in the first-level subjects, and respiratory medicine, digestive medicine, etc. in the second-level subjects. Among them, the user's diagnostic intention can include multiple ones. Analyze all the user's diagnostic intentions uploaded this time to determine the diagnostic field that the user needs for auxiliary diagnosis this time, and based on the diagnostic field, remove the irrelevant content in the auxiliary diagnosis report for predicting bladder cancer gene mutations, and generate a corresponding precise auxiliary diagnosis report and feedback it to the specified terminal. Among them, the process of analyzing all the user's diagnostic intentions uploaded this time to determine the diagnostic field that the user needs for auxiliary diagnosis this time includes: Establish a diagnostic intention database. Among them, the diagnostic intention database stores several diagnostic intentions that have been marked with diagnostic fields. Each diagnostic intention is marked with an association value with different diagnostic fields, and each diagnostic intention can be related to several diagnostic fields at the same time. The association value is output and marked in advance according to the association value scoring model; Generate the first association value m between each user's diagnostic intention and different diagnostic fields. Among them, there is a first association value m between each user's diagnostic intention and each different diagnostic field. If there is 1 user's diagnostic intention that has no association with the current diagnostic field, its first association value is 0. For example, there are 10 user's diagnostic intentions and 10 diagnostic fields in the database, then the number of existing first association values m is 100, and 10 first association values will be generated in each diagnostic field; Or, there are 9 user's diagnostic intentions and 10 diagnostic fields in the database, then the number of existing first association values m is 90, and 9 first association values will be generated in each diagnostic field; Based on all the first association values m, count several second association values of the user's diagnostic intentions corresponding to any diagnostic field. The second association values include m1, m2...m n , where n is the total number of user's diagnostic intentions, and then screen out the user's diagnostic intentions corresponding to all the second association values that exceed the preset statistical threshold p as the target diagnostic intentions, and calculate the number k of the target diagnostic intentions in each diagnostic field i , where i is the number of diagnostic fields; Count the number k of the target diagnostic intentions in each diagnostic field i The proportion j of the total number of user's diagnostic intentions i , that is, j i =k i / n; Preset the condition satisfaction value g. When the product of k i and j i is not less than g, take the product maximum value of k i and j i The corresponding diagnostic field is output as the target diagnostic field; It should be noted that when k i and j iWhen the product is less than g, an associated error is output and fed back to the specified terminal, and at the same time, an auxiliary diagnostic report for predicting bladder cancer gene mutations is output to the instruction terminal; among them, the evaluation process of the association value between each user diagnosis intention and different diagnosis fields is preferably carried out by extracting the diagnosis and treatment text features from each user diagnosis intention, and the diagnosis and treatment text features include, but are not limited to, feature information such as text length, diagnosis and treatment entity naming, etc. A deep learning model is used to establish an association value scoring model, and each diagnosis and treatment text feature is labeled in advance through a tag system with high accuracy. The labeled data includes different diagnosis field data, and then the association value scoring model is optimized and verified through the labeled feature data to improve the association value scoring model; in actual use, the user diagnosis intention is input into the association value scoring model, and the association value between each user diagnosis intention and different diagnosis fields is obtained as output;

[0043] The beneficial effects of the above technical solution are as follows: Through the above technical solution, the user's diagnosis intention is analyzed and the diagnosis intention is matched with the actual diagnosis field, thereby streamlining the content of the auxiliary report, so that there will be no useless data on the auxiliary report, providing personalized services for users. At the same time, the streamlined auxiliary report can reduce the probability of the user's attention being distracted, so that the user will not be distracted by irrelevant data.

[0044] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include these modifications and variations.

Claims

1. A bladder cancer gene mutation prediction system based on deep learning, characterized in that, Including: A histopathological image acquisition module for acquiring original histopathological images to be measured; A preprocessing module for preprocessing the original histopathological images to be measured by using the HE staining method and a whole-slide imaging scanner to obtain a WSI of bladder tissue with HE staining; A cloud server module for calling a deep learning prediction model for bladder cancer to respectively predict the gene expression status, gene methylation status, and gene mutation status of bladder cancer for the WSI, and generating corresponding prediction results; A report generation module for automatically generating an auxiliary diagnostic report for bladder cancer gene mutation prediction based on the corresponding prediction results and feeding it back to a specified terminal; The report generation module performs the following operations: Obtain the corresponding prediction results, and extract the corresponding original case data from a preset database according to the corresponding prediction results, where the original case data is the case data participated in the prediction by the patient corresponding to the current prediction result during this prediction; Extract the effective initial text corpus in the original case data based on regular expressions, where each initial text corpus uniquely corresponds to one prediction result; Sort all the initial text corpora according to the text length, count the text length, and construct a text length relationship table; Perform entity extraction of medical entity types for each initial text corpus, determine the entity type of each initial text corpus, and associate the entity types of all the initial text corpora under each prediction result according to the preset entity association relationship to obtain the entity association relationship of each initial text corpus; Determine the initial report template according to the longest text length in the text length relationship table, and perform merging and filling processing on any number of associated initial text corpora based on the entity association relationship of the initial text corpora under each prediction result to obtain the report text to be generated; Based on the initial report text filling logic corresponding to the initial report template, fill the report text to be generated and the prediction results into the initial report template after corresponding to generate an auxiliary diagnostic report for bladder cancer gene mutation prediction.

2. The bladder cancer gene mutation prediction system based on deep learning according to claim 1, wherein The histopathological image acquisition module is connected to a preset urinary system image acquisition device for acquiring the original histopathological images to be measured acquired by the preset urinary system image acquisition device.

3. The bladder cancer gene mutation prediction system based on deep learning according to claim 1, characterized in that The preprocessing module performs the following operations: Stain the bladder tissue to be measured according to the HE staining method to obtain bladder tissue with HE staining; According to the original histopathological image stained by the HE staining method, use a whole-slide imaging scanner to obtain a WSI of bladder tissue with HE staining.

4. A bladder cancer gene mutation prediction system based on deep learning according to claim 1, characterized in that The training process of the deep learning prediction model for bladder cancer includes the following steps: Obtain training samples; Construct a deep learning framework based on a convolutional neural network, and train the deep learning framework according to the training samples until the preset conditions are met to end the training and generate a deep learning prediction model for bladder cancer.

5. The bladder cancer gene mutation prediction system based on deep learning according to claim 4, wherein Obtaining training samples includes: obtaining training WSI of several HE-stained bladder tissues in the public TCGA; performing label marking on the training WSI and filtering the training WSI to obtain several WSI images that only retain the labels of gene expression, gene methylation, and gene mutation; dividing the several WSI images to obtain a validation WSI set, a training WSI set, and a test WSI set; based on a magnification of 400 times, dividing the images of all WSI sets into image patches of 320×320 pixels and filtering the background to obtain a second validation WSI set, a second training WSI set, and a second test WSI set; where all WSI sets include the validation WSI set, the training WSI set, and the test WSI set.

6. The system for predicting bladder cancer gene mutations based on deep learning according to claim 5, characterized in that It also includes accelerating the training process using a two-stage pre-analysis method. The first-stage acceleration process includes: based on a magnification of 100 times, dividing the images of all WSI sets into image patches of 320×320 pixels and filtering the background to obtain a third validation WSI set, a third training WSI set, and a third test WSI set for constructing an analysis data set; using the pre-trained RestNet-50 model on ImageNet to convert each image patch in the analysis data set into a 2048-dimensional vector as a feature extractor; where all WSI sets include the validation WSI set, the training WSI set, and the test WSI set.

7. A bladder cancer gene mutation prediction system based on deep learning according to claim 6, characterized in that, The second-stage acceleration process includes: using the 3-fold cross-validation method in the k-fold cross-validation method to evaluate the performance of Xgboost trained on the feature vectors to obtain an evaluation result; where in each fold, 200 WSI in the analysis data set are randomly selected for training, and 100 WSI are selected for testing at the same time; based on the RestNet-50 model, reconstruct a classification model on the second validation WSI set, the second training WSI set, and the second test WSI set, and combine the evaluation results to provide detailed classification information for each selected gene for marking.

8. A bladder cancer gene mutation prediction system based on deep learning according to claim 1, characterized in that, It also includes: Obtaining the user's diagnostic intention uploaded by the specified terminal, the user's diagnostic intention includes various medical treatment subjects, where there are multiple user's diagnostic intentions; analyzing all the user's diagnostic intentions uploaded this time to determine the diagnostic field that the user needs auxiliary diagnosis for this time, and removing the irrelevant content in the bladder cancer gene mutation prediction auxiliary diagnosis report based on the diagnostic field, and generating a corresponding precise auxiliary diagnosis report and feedbacking it to the specified terminal.

9. The bladder cancer gene mutation prediction system based on deep learning according to claim 8, characterized in that Analyze all user diagnostic intents uploaded this time to determine the diagnostic fields that the user needs assisted diagnosis for this time, including: establishing a diagnostic intent database, where the diagnostic intent database stores a number of diagnostic intents with marked diagnostic fields, each diagnostic intent is marked with an association value with different diagnostic fields, and each diagnostic intent is related to several diagnostic fields at the same time; based on the diagnostic intent database, generate a first association value m between each user diagnostic intent and the diagnostic field, where there is a first association value m between each user diagnostic intent and each diagnostic field; based on all the first association values m, count several second association values of the user diagnostic intents corresponding to any diagnostic field, and the second association values include m1, m2...m n , where n is the total number of user diagnostic intents; screen out the user diagnostic intents corresponding to all second association values exceeding the preset statistical threshold p as the target diagnostic intents, and calculate the number k of target diagnostic intents in each diagnostic field i , where i is the number of diagnostic fields; count the number k of target diagnostic intents in each diagnostic field i The proportion j of the total number of user diagnostic intents i ; preset a condition satisfaction value g, when the product of k i and j i is not less than g, take the diagnostic field corresponding to the maximum value of the product of k i and j i as the target diagnostic field for output, and determine the diagnostic field that the user needs assisted diagnosis for this time.

Citation Information

Patent Citations

  • Intelligent cognitive disease system based on human excrement

    CN112002415A

  • Bladder cancer pathomics intelligent diagnosis method based on machine learning and prognosis model thereof

    CN112435743A