Microsatellite instability colorectal cancer mechanism information analysis system based on deep learning
By establishing a microsatellite instable colorectal cancer mechanism information analysis system based on deep learning, the problem of cumbersome and missed diagnosis of Lynch syndrome is solved, and the rapid and accurate identification and personalized diagnosis of Lynch syndrome patients is achieved, and the intelligent development of pathology is promoted.
Patent Information
- Application Number
- CN202311047223.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-18
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2043-08-18
AI Technical Summary
The existing technology is complicated, expensive and easy to miss diagnosis during the diagnosis of Lynch syndrome. The existing deep learning models are complicated and easy to miss diagnosis in the diagnosis of microsatellite instability colorectal cancer, and cannot effectively distinguish the molecular mechanism of microsatellite instability colorectal cancer from the pathogenesis and prognosis prediction of patients with Lynch syndrome.
Establish a microsatellite instability colorectal cancer mechanism information analysis system based on deep learning, collect and screen case document data through data set establishment module, use the gene screening module to screen mutant genes, train the model to build a digital slice model, and learn the model to build a module to build an artificial intelligence deep learning model, including microsatellite instability colorectal cancer diagnosis, Lynch syndrome diagnosis and prognosis prediction models to achieve automatic identification of mutant genes.
Reduce detection steps, reduce costs, reduce missed diagnosis, provide fast and accurate identification and personalized diagnosis solutions for patients with Lynch syndrome, realize molecular diagnosis, prognosis prediction and efficacy evaluation, and promote the development of pathology towards intelligence.
Smart Images

Figure CN117174179B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence and deep learning model technology, and in particular to a microsatellite instability colorectal cancer mechanism information analysis system based on deep learning. Background Art
[0002] Colorectal cancer is one of the most common malignancies worldwide. Its development and progression are closely linked to genomic instability and can be categorized as hereditary or sporadic. Hereditary forms of colorectal cancer account for approximately 15% of all colorectal cancers. Lynch syndrome, formerly known as hereditary non-polyposis colorectal cancer, is an autosomal dominant genetic disorder that accounts for approximately 3%-5% of all colorectal cancers. Studies have shown that Lynch syndrome is caused by germline mutations in DNA mismatch repair genes (primarily hMLH1, hMSH2, hMSH6, hPMS2, or EPCAM), leading to a genetic susceptibility to various cancers. The DNA mismatch repair system is a crucial mechanism for maintaining DNA replication accuracy and genomic stability, primarily repairing base mismatches and small insertion-deletion loops in newly synthesized DNA strands during DNA replication. Deficient mismatch repair (dMMR) increases the rate of spontaneous genomic mutations, leading to microsatellite instability (MSI) and potentially contributing to a range of cancers. In addition to Lynch syndrome, sporadic microsatellite instability colorectal cancer can also occur due to factors such as somatic MLH1 gene promoter hypermethylation and BRAF gene mutation. Among microsatellite instability colorectal cancers, Lynch syndrome accounts for approximately 30% and sporadic cases account for approximately 70%.
[0003] In the diagnosis and treatment of colorectal cancer, pathological diagnosis, as the gold standard, provides crucial support for clinicians' treatment decisions. However, precision medicine for colorectal cancer presents higher demands and challenges for pathological diagnosis. Beyond routine diagnosis, in-depth understanding of the molecular pathological changes and pathogenesis of colorectal cancer, including gene mutation status associated with drug sensitivity, and immunohistochemical analysis of mismatch repair protein expression and microsatellite instability have become routine diagnostic components. Identifying familial genetic alterations in tumors is also essential for personalized colorectal cancer diagnosis. Colorectal cancer diagnosis accounts for a significant portion of routine pathology work, and the substantial increase in diagnostic demand has had a significant impact on pathology departments. Pathology has undergone centuries of development, achieving significant progress from morphology to molecular changes, but also facing bottlenecks. The cross-disciplinary integration and application of new technologies can bring opportunities and breakthroughs. Digitalization of pathology and artificial intelligence-assisted diagnosis are now on the agenda.
[0004] In recent years, deep learning technology has been gradually applied to the field of medical image analysis, particularly in imaging, where it has been successfully applied to the segmentation, recognition, and diagnosis of brain, abdominal, and lung images. With the public release of the TCGA and CAMELYON datasets, deep learning has also been introduced into histopathology research based on HE staining of pathological sections, achieving promising results in the diagnosis of lymph node, breast, and lung tissue sections. In the field of colorectal cancer, deep learning research is being used to automatically detect gross abnormalities such as polyps in endoscopic video images. For diagnostic applications of histopathology images, reported deep learning model studies have primarily focused on the pathological diagnosis of a single pathological type, such as the identification of intestinal cancer and intestinal polyps.
[0005] The current gold standard for the diagnosis of Lynch syndrome is to identify germline pathogenic mutations through multi-step molecular pathology testing, and to exclude other causes of microsatellite instability such as MLH1 gene promoter hypermethylation and BRAF gene mutations. In addition, previous studies have also found that MLH1 and MSH2 genes are prone to large-scale gene rearrangements (Largegenomic rearrangements, LGRs). Conventional second-generation gene sequencing alone cannot accurately identify large-scale rearrangements, resulting in approximately 17% of Lynch syndrome patients being missed. Therefore, the diagnosis and molecular typing of Lynch syndrome patients require multiple procedures such as immunohistochemistry staining, second-generation gene sequencing, promoter region methylation testing, microsatellite instability testing, and large-scale gene rearrangement testing. The testing process is cumbersome, the testing cost is high, there are many quality control links, and missed diagnoses may still occur. At present, the pathology departments of most hospitals in my country do not have all the testing technologies and platforms, which seriously limits their feasibility in clinical practice.
[0006] Prior art 1, application number CN202210277503.8, discloses a microsatellite instability detection system based on colorectal pathology images. The image acquisition module is used to obtain colorectal pathology images of the user being tested; the tumor tissue detection module is used to crop the image into several image blocks and detect whether each image block is a tumor image; the microsatellite instability detection module uses a microsatellite instability classification model to identify whether each tumor image has high microsatellite instability; when training the tumor tissue detection model, it is initialized using pre-trained weights, and the top classifier and auxiliary classifier are trained at a first training rate. The parameters of the other layers of the tumor tissue detection model are frozen until the tumor tissue detection model converges. All parameters of the tumor tissue detection model are then unfrozen and the tumor tissue detection model is fitted at a second training rate. Although reliable, widely applicable, and cost-effective, the detection process is cumbersome and involves multiple quality control steps.
[0007] Prior art 2, application number CN202211274443.0, discloses a system and image segmentation method for predicting microsatellite instability in colorectal cancer. This method reads multi-resolution images from the original full-field slice image data, uses the low-resolution image to remove the background from the high-resolution image, and then segments the image into several blocks. The image entropy of all blocks is calculated. Blocks are screened based on the image entropy, and blocks with too low an entropy are discarded to obtain an entropy mask consisting of the remaining blocks. The entropy mask is multiplied by the high-resolution image to obtain the segmented valid area. While this method is efficient, reliable, widely applicable, requires no manual annotation, and is cost-effective, it holds great promise for clinical application. However, the process is cumbersome and can easily lead to missed diagnoses.
[0008] Existing technology three, deep learning can predict microsatellite instability in gastrointestinal cancer directly from histology [Kather JN, Pearson AT, Halama N, et al. Deep learning can predict microsatellite instability directly from histology in gastrointestinal cancer. Nature Medicine. 2019; 25: 1054-1056.], what is the relationship between MSI and dMMR? Because of the relationship, because of the functional defect of the mismatch repair gene (dMMR) leads to the abnormal function of the mismatch repair protein, the random errors generated during DNA replication cannot be repaired normally, and then the phenomenon of high microsatellite instability (MSI-H) occurs. Clinically, it is generally assumed that dMMR = MSI-H. Mutations in certain parts of the DNA in cancer cells (the so-called microsatellites, MS) can affect the response of patients with gastrointestinal tumors to immunotherapy. Microsatellite instability (MSI) / dMMR is a frequently mentioned metric in patients with gastrointestinal tumors. This metric can determine whether patients with gastrointestinal tumors respond well to immunotherapy. Although the discovery of an indicator that determines the response of gastrointestinal tumor patients to immunotherapy has been achieved, further research on deep learning models for identifying the molecular mechanisms of microsatellite instability in colorectal cancer and the pathogenesis and prognosis prediction of Lynch syndrome patients has not been reported.
[0009] Currently, existing technologies 1, 2, and 3 have the problems of cumbersome detection processes and inaccurate results. Therefore, the present invention provides a microsatellite instability colorectal cancer mechanism information analysis system based on deep learning, establishes a complete molecular pathology detection platform, clarifies the mutant genes, uses tissue pathology digital sections as training data sets, and groups them according to different mutant genes. A computer-aided diagnosis system based on deep learning is obtained for the molecular mechanism of microsatellite instability colorectal cancer and the pathogenesis and prognosis prediction of Lynch syndrome patients. Summary of the Invention
[0010] In order to solve the above technical problems, the present invention provides a deep learning-based microsatellite instability colorectal cancer mechanism information analysis system, comprising:
[0011] The data set establishment module is responsible for collecting the first case document data of surgically resected colorectal cancer, screening the second case document data of microsatellite unstable colorectal cancer from the first case document data, wherein the second case document data includes: the third case document data of colorectal cancer in a family with Lynch syndrome, and selecting data from the third case document data as the training data set and test data set of the artificial intelligence deep learning model;
[0012] The gene screening module is responsible for screening the mutant genes of colorectal cancer in Lynch syndrome families from the second case file data and the third case file data;
[0013] The training model building module is responsible for constructing a digital slice training model based on the training data set. It scans stained pathological tissue sections of surgically removed colorectal cancer to obtain digital slices. It then uses the colorectal cancer diagnostic model to pre-annotate the digital slices and assign labels for different mutated genes to the entire digital slice.
[0014] The learning model establishment module is responsible for establishing an artificial intelligence deep learning model, testing the artificial intelligence deep learning model using a test data set, inputting labels into the artificial intelligence deep learning model, and obtaining microsatellite unstable colorectal cancer and Lynch syndrome family colorectal cancer, which cause inherited germline mutation genes.
[0015] Optional artificial intelligence deep learning models include: microsatellite instability colorectal cancer diagnosis model, Lynch syndrome diagnosis model, Lynch syndrome gene mutation mechanism model, microsatellite instability colorectal cancer and Lynch syndrome prognosis prediction model.
[0016] Optional, dataset creation module, including:
[0017] The first data screening submodule is responsible for acquiring first case document data stored in the server through network crawling, and establishing a database to be screened based on the first case document data using the gene characteristics of colorectal cancer removed by surgical operation as the first gene attribute characteristics and the first gene attribute characteristics as the feature identifier;
[0018] The second data screening submodule is responsible for using the microsatellite instability colorectal cancer gene feature as the second gene attribute feature and as the screening condition to screen from the database to be screened to obtain the second case document data;
[0019] The third data screening submodule is responsible for receiving the second case document data, evaluating the importance of the second case document data according to the third gene attribute feature, obtaining an importance evaluation score, and establishing the third case document data with the third gene attribute feature as a feature identifier; the colorectal cancer gene feature of the Lynch syndrome family is used as the third gene attribute feature;
[0020] The dataset requirement submodule is responsible for selecting case document data from the third case document data according to the dataset requirements of the artificial intelligence deep learning model, and constructing the training dataset and test dataset.
[0021] Optional, genetic screening module, including:
[0022] The first screening submodule is responsible for obtaining the genetic characteristics of microsatellite instability colorectal cancer and the genetic characteristics of colorectal cancer in Lynch syndrome families, and screening out important genes from the genetic characteristics;
[0023] The second screening submodule is responsible for functional analysis of important genes and screening the gene loci in important genes;
[0024] The contribution calculation submodule is responsible for performing principal component analysis on the gene loci in important genes using the principal component analysis method to obtain the functional contribution of each important gene in the gene feature, and querying the functional contribution of non-important genes from the data contribution of each important gene in the gene feature;
[0025] The third screening module is responsible for screening the mutated genes of colorectal cancer in Lynch syndrome families from important genes based on the functional contribution of each important gene.
[0026] Optional, training model building module, including:
[0027] The image acquisition submodule is responsible for acquiring digital images in the labeled training set and training the labeled digital images through deep learning;
[0028] The slice acquisition submodule is responsible for staining and slicing the pathological tissue of surgically removed colorectal cancer. It uses a digital pathology scanner to scan the digital image at 40 times the resolution to obtain digital slices containing pathological information.
[0029] The auxiliary annotation submodule is responsible for detecting the bounding box of the digital slice and capturing only the slice image within the bounding box. The module performs peripheral contour detection on the slice image to obtain at least one bounding box, and determines the bounding box with the largest area as the corresponding bounding box contained in the digital image. The digital slice and its peripheral contour are then pre-annotated using the colorectal cancer diagnostic model.
[0030] The gene label submodule is responsible for finding mutant genes from digital slices through convolutional neural networks, extracting the numerical features of mutant genes, and assigning labels to different mutant genes according to the numerical features.
[0031] Optionally, among the multiple bounding boxes obtained by performing peripheral contour detection on the digital image, the bounding boxes whose areas are smaller than an area threshold, whose lengths are smaller than a length threshold, and whose widths are smaller than a width threshold are eliminated.
[0032] Optional learning model building module, including:
[0033] The data acquisition submodule is responsible for acquiring multiple digital slices that are assigned different mutation gene labels;
[0034] The data annotation submodule is responsible for obtaining the information of digital slices and auxiliary annotation of their peripheral contours;
[0035] The model training submodule is responsible for training the microsatellite instability colorectal cancer diagnostic model and the Lynch syndrome diagnostic model using a convolutional neural network based on residual networks and classical networks. The convolutional neural network automatically discovers the characteristic information of digital slices from various angles. By extracting features at various levels, it obtains the most essential features related to microsatellite instability and Lynch syndrome, completing automatic identification.
[0036] The service launch sub-module is responsible for accumulating digital slices. When the accumulation reaches a set threshold, it uses automated machine learning to build a convolutional neural network to realize the construction of microsatellite instability colorectal cancer diagnosis model, Lynch syndrome diagnosis model and Lynch syndrome gene mutation mechanism model.
[0037] Optional service launch submodule, including:
[0038] The data preparation unit is responsible for selecting multiple digital slices with different mutation gene labels as experimental data sets, and performing preprocessing and normalization preprocessing to obtain the preprocessed experimental data sets;
[0039] Search space unit, responsible for defining the search space, including basic units of residual networks and classical networks, convolution kernel sizes of feature maps, numbers of convolution kernels, activation functions, padding methods, downsampling methods, etc.;
[0040] Objective function definition unit, responsible for evaluating the performance of different convolutional neural network models using the method of cross-validation, and defining the objective function as an index for optimizing the network model;
[0041] Running automatic search unit, responsible for searching for convolutional network model parameters and structures using search algorithms and deep learning techniques;
[0042] Model evaluation unit, responsible for verifying the accuracy and generalization ability of the convolutional network model using the preprocessed experimental data set.
[0043] Optionally, the learning model establishment module includes:
[0044] Image slicing sub-module, responsible for scanning the pathological tissue stained sections of colorectal cancer to obtain the images of full-scanned digital slices, and using the weakly supervised learning framework to equally slice the images into image patches. Each image of the full-scanned digital slice is divided into N×N patches of equal size. M and m represent the sizes of the image of the full-scanned digital slice and the image patch respectively, and N = M / m is the scale factor;
[0045] Classification label sub-module, responsible for finding the representative image patches in the image through the training of multi-instance learning, and the prediction result of the training is the classification label of the image of the whole digital slice;
[0046] Model establishment sub-module, responsible for implementing the coarse-grained label at the image level of the whole digital slice, and establishing a deep learning model through the classification label at the slice level to complete the establishment of the prognostic prediction models for microsatellite instability colorectal cancer and Lynch syndrome respectively.
[0047] Optionally, the image slicing sub-module includes:
[0048] [[ID=二十七]]Matrix establishment unit, responsible for establishing an attention matrix for each image of the full-scanned digital slice, and storing the attention weights corresponding to all image patches in the matrix;
[0049] Probability prediction unit, responsible for storing the prediction probabilities of the classification model for the image patches using weights. If not predicted, the weight is defaulted to N / A, where A represents the image patch; defining the sampling number as n, n << N, and sampling the image patches according to the attention weights with probability p in each iteration. p ranges from 0.0 to 1.0, that is, on average, sampling p×n image patches with non-N / A attention weights;
[0050] The image sampling unit is responsible for randomly sampling from image blocks with a weight of N / A with a probability of 1-p, that is, sampling an average of (1-p)×n image blocks with non-N / A attention weights. In each iteration, a single fully scanned digital slice image is sampled with a total of n image blocks for classification model training.
[0051] The data set establishment module of the present invention is responsible for collecting the first case document data of colorectal cancer removed by surgery, screening out the second case document data of microsatellite unstable colorectal cancer from the first case document data, wherein the second case document data includes: the third case document data of colorectal cancer in a family with Lynch syndrome, and selecting data from the third case document data as the training data set and test data set of the artificial intelligence deep learning model; the gene screening module is responsible for screening out the mutant genes of colorectal cancer in a family with Lynch syndrome from the second case document data and the third case document data; the training model establishment module is responsible for constructing a digital slice training model from the training data set, scanning the pathological tissue stained slices of colorectal cancer removed by surgery to obtain digital slices, using the colorectal cancer diagnosis model to pre-annotate the digital slices, and assigning labels of different mutant genes to the entire digital slices; the learning model establishment module is responsible for establishing an artificial intelligence deep learning model, using the test The artificial intelligence deep learning model was tested on a trial data set, and the labels were input into the artificial intelligence deep learning model to obtain microsatellite instability colorectal cancer and Lynch syndrome family colorectal cancer, as well as the inherited germline mutation genes that cause them; among them, the artificial intelligence deep learning model includes: microsatellite instability colorectal cancer diagnosis model, Lynch syndrome diagnosis model, Lynch syndrome gene mutation mechanism model, microsatellite instability colorectal cancer and Lynch syndrome prognosis prediction model; the above scheme uses a large number of colorectal cancer surgical specimens to establish a complete molecular pathology detection platform, clarifies the mutant genes of each case, and will combine the artificial intelligence deep learning model, use surgical specimen tissue pathology digital slices as training data sets, group according to different mutant genes, and combine family medical history and clinical follow-up information to develop a computer-aided diagnosis system based on deep learning for the molecular mechanism of microsatellite instability colorectal cancer and the pathogenesis and prognosis prediction of Lynch syndrome patients. The embodiments of the present invention not only reduce the complex multi-step testing process in the Lynch syndrome screening process, saving time, manpower and testing costs, and reducing missed diagnoses, but also provide remote diagnosis capabilities for institutions and patients that do not have the necessary testing conditions using scanned pathology images, quickly identifying patients with Lynch syndrome and developing personalized diagnostic plans. At the same time, as an exploratory work in the field of artificial intelligence in pathology, its results will not only assist pathologists in their existing workflow, but will also go beyond traditional pathology diagnosis to achieve diversified and in-depth diagnosis such as molecular diagnosis, prognosis prediction, and efficacy evaluation, reflecting the application value of artificial intelligence, providing clinicians and patients with more meaningful diagnostic information, and promoting the development of pathology towards intelligence. It has important clinical, economic, and social significance for the definitive diagnosis of Lynch syndrome.
[0052] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present invention. The purposes and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description, claims, and drawings.
[0053] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0055] Figure 1 This is a block diagram of the microsatellite instability colorectal cancer mechanism information analysis system based on deep learning in Example 1 of the present invention;
[0056] Figure 2 A module block diagram for establishing a data set in Example 2 of the present invention;
[0057] Figure 3 This is a block diagram of the gene screening module in Example 3 of the present invention;
[0058] Figure 4 A block diagram of a module for establishing a training model in Example 4 of the present invention;
[0059] Figure 5 Establishing a module framework for the learning model in Example 5 of the present invention Figure 1 ;
[0060] Figure 6 This is a block diagram of the service online submodule in Example 6 of the present invention;
[0061] Figure 7 Establishing a module framework for the learning model in Example 7 of the present invention Figure 2 ;
[0062] Figure 8 This is a block diagram of the image segmentation submodule in Example 8 of the present invention;
[0063] Figure 9 This is the classification label submodule frame in Example 9 of the present invention. Figure 1 ;
[0064] Figure 10 The classification label submodule frame in embodiment 10 of the present invention Figure 2 . DETAILED DESCRIPTION
[0065] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0066] The terms used in the embodiments of the present application are for the purpose of describing specific embodiments only and are not intended to limit the embodiments of the present application. The singular forms "a," "the," and "the" used in the embodiments of the present application and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.
[0067] When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims. In the description of the present application, it should be understood that the terms "first", "second", "third", etc. are only used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence, nor can they be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to the specific circumstances.
[0068] Example 1: Figure 1 As shown, an embodiment of the present invention provides a microsatellite instability colorectal cancer mechanism information analysis system based on deep learning, comprising:
[0069] The data set establishment module is responsible for collecting the first case document data of surgically resected colorectal cancer, screening the second case document data of microsatellite unstable colorectal cancer from the first case document data, wherein the second case document data includes: the third case document data of colorectal cancer in a family with Lynch syndrome, and selecting data from the third case document data as the training data set and test data set of the artificial intelligence deep learning model;
[0070] The gene screening module is responsible for screening the mutant genes of colorectal cancer in Lynch syndrome families from the second case file data and the third case file data;
[0071] The training model building module is responsible for constructing a digital slice training model based on the training data set. It scans stained pathological tissue sections of surgically removed colorectal cancer to obtain digital slices. It then uses the colorectal cancer diagnostic model to pre-annotate the digital slices and assign labels for different mutated genes to the entire digital slice.
[0072] The learning model building module is responsible for establishing an artificial intelligence deep learning model, testing the artificial intelligence deep learning model using a test data set, and inputting labels into the artificial intelligence deep learning model to obtain the germline mutation genes that cause microsatellite instability colorectal cancer and Lynch syndrome familial colorectal cancer. The artificial intelligence deep learning model includes: a microsatellite instability colorectal cancer diagnosis model, a Lynch syndrome diagnosis model, a Lynch syndrome gene mutation mechanism model, and a microsatellite instability colorectal cancer and Lynch syndrome prognosis prediction model.
[0073] The working principle and beneficial effects of the above technical solution are as follows: the data set establishment module of this embodiment is responsible for collecting the first case document data of colorectal cancer removed by surgery, and screening out the second case document data of microsatellite unstable colorectal cancer from the first case document data, wherein the second case document data includes: the third case document data of colorectal cancer in a family with Lynch syndrome, and selecting data from the third case document data as the training data set and test data set of the artificial intelligence deep learning model; the gene screening module is responsible for screening out the mutant genes of colorectal cancer in a family with Lynch syndrome from the second case document data and the third case document data; the training model establishment module is responsible for constructing a digital slice training model from the training data set, scanning the pathological tissue stained slices of colorectal cancer removed by surgery to obtain digital slices, using the colorectal cancer diagnostic model to pre-annotate the digital slices, and assigning labels of different mutant genes to the entire digital slices; the learning model establishment module is responsible for establishing the artificial intelligence deep learning model. The artificial intelligence deep learning model was tested using a test data set, and the labels were input into the artificial intelligence deep learning model to obtain the inherited germline mutation genes that cause microsatellite instability colorectal cancer and Lynch syndrome family colorectal cancer; among them, the artificial intelligence deep learning model includes: microsatellite instability colorectal cancer diagnosis model, Lynch syndrome diagnosis model, Lynch syndrome gene mutation mechanism model, microsatellite instability colorectal cancer and Lynch syndrome prognosis prediction model; the above scheme uses a large number of colorectal cancer surgical specimens to establish a complete molecular pathology detection platform, clarifies the mutant genes of each case, and will combine the artificial intelligence deep learning model, use surgical specimen tissue pathology digital slices as training data sets, group according to different mutant genes, and combine family medical history and clinical follow-up information to develop a computer-aided diagnosis system based on deep learning for the molecular mechanism of microsatellite instability colorectal cancer and the pathogenesis and prognosis prediction of Lynch syndrome patients. The embodiments of the present invention not only reduce the complex multi-step testing process in the Lynch syndrome screening process, saving time, manpower and testing costs, and reducing missed diagnoses, but also provide remote diagnosis capabilities for institutions and patients that do not have the necessary testing conditions using scanned pathology images, quickly identifying patients with Lynch syndrome and developing personalized diagnostic plans. At the same time, as an exploratory work in the field of artificial intelligence in pathology, its results will not only assist pathologists in their existing workflow, but will also go beyond traditional pathology diagnosis to achieve diversified and in-depth diagnosis such as molecular diagnosis, prognosis prediction, and efficacy evaluation, reflecting the application value of artificial intelligence, providing clinicians and patients with more meaningful diagnostic information, and promoting the development of pathology towards intelligence. It has important clinical, economic, and social significance for the definitive diagnosis of Lynch syndrome.
[0074] Example 2: Figure 2 As shown, based on Example 1, the dataset establishment module provided in this embodiment of the present invention includes:
[0075] The first data screening submodule is responsible for acquiring first case document data stored in the server through network crawling, and establishing a database to be screened based on the first case document data using the gene characteristics of colorectal cancer removed by surgical operation as the first gene attribute characteristics and the first gene attribute characteristics as the feature identifier;
[0076] The second data screening submodule is responsible for using the microsatellite instability colorectal cancer gene feature as the second gene attribute feature and as the screening condition to screen from the database to be screened to obtain the second case document data;
[0077] The third data screening submodule is responsible for receiving the second case document data, evaluating the importance of the second case document data according to the third gene attribute feature, obtaining an importance evaluation score, and establishing the third case document data with the third gene attribute feature as a feature identifier; the colorectal cancer gene feature of the Lynch syndrome family is used as the third gene attribute feature;
[0078] The dataset requirement submodule is responsible for selecting case document data from the third-party case document data and constructing training and test datasets according to the dataset requirements of the artificial intelligence deep learning model;
[0079] The working principle and beneficial effects of the above technical solution are as follows: the first data screening submodule of this embodiment is responsible for acquiring the first case document data stored from the server through network crawling, and for the first case document data, using the gene characteristics of colorectal cancer removed by surgical operation as the first gene attribute characteristics, and using the first gene attribute characteristics as the feature identifier, a database to be screened is established; the second data screening submodule is responsible for using the gene characteristics of microsatellite unstable colorectal cancer as the second gene attribute characteristics, as the screening condition, to screen from the database to be screened, and obtain the second case document data; the third data screening submodule is responsible for receiving the second case document data, and evaluating the importance of the second case document data according to the third gene attribute characteristics, obtaining the importance evaluation score, and establishing the third case document data with the third gene attribute characteristics as the feature identifier; the gene characteristics of colorectal cancer in Lynch syndrome families are used as the screening conditions. is the third gene attribute feature; the dataset requirement submodule is responsible for selecting case document data from the third case document data according to the dataset requirements of the artificial intelligence deep learning model, and constructing a training dataset and a test dataset; the above scheme uses the first case document data with a large sample size as the basic data, and the genetic characteristics of various types of colorectal cancer as characteristic identifiers, to achieve the distinction between microsatellite unstable colorectal cancer and Lynch syndrome family colorectal cancer. Using genetic characteristics as the distinguishing point firstly improves the accuracy of distinguishing different types of colorectal cancer, and also ensures the accuracy of the colorectal cancer mechanism information analysis results. At the same time, the first case document data is screened twice to achieve the effectiveness of data screening, ensure the integrity of the second case document dataset and the third case document data, and also ensure the number of data in the training set and the test set, ensuring the efficient and accurate operation of the artificial intelligence deep learning model.
[0080] Example 3: Figure 3 As shown, based on Example 1, the gene screening module provided by the embodiment of the present invention includes:
[0081] The first screening submodule is responsible for obtaining the genetic characteristics of microsatellite instability colorectal cancer and the genetic characteristics of colorectal cancer in Lynch syndrome families, and screening out important genes from the genetic characteristics;
[0082] The second screening submodule is responsible for functional analysis of important genes and screening the gene loci in important genes;
[0083] The contribution calculation submodule is responsible for performing principal component analysis on the gene loci in important genes using the principal component analysis method to obtain the functional contribution of each important gene in the gene feature, and querying the functional contribution of non-important genes from the data contribution of each important gene in the gene feature;
[0084] The third screening module is responsible for screening the mutated genes of colorectal cancer in Lynch syndrome families from important genes based on the functional contribution of each important gene;
[0085] The working principle and beneficial effects of the above technical solution are as follows: the first screening submodule of this embodiment obtains the microsatellite instability colorectal cancer gene characteristics and the Lynch syndrome family colorectal cancer gene characteristics, and screens out important genes from the gene characteristics; the second screening submodule performs functional analysis on the important genes and screens out the gene loci in the important genes; the contribution calculation submodule uses the principal component analysis method to perform principal component analysis on the gene loci in the important genes to obtain the functional contribution of each important gene in the gene characteristics, and queries the functional contribution of non-important genes in the data contribution of each important gene in the gene characteristics; the third screening module performs functional analysis on the gene loci in the important genes according to the data contribution of each important gene in the gene characteristics. Based on the functional contribution of the gene, the mutant genes of Lynch syndrome family colorectal cancer were screened from the important genes; the above scheme obtained the mutant genes by three screenings of the gene characteristics of microsatellite unstable colorectal cancer and the gene characteristics of Lynch syndrome family colorectal cancer. The order of magnitude of the screening improved the screening accuracy of the mutant genes and ensured the accuracy of the mechanism information analysis of Lynch syndrome family colorectal cancer; in addition, through the calculation of the functional contribution, the proportion of important genes in the gene characteristics was confirmed, and the pathogenesis of microsatellite unstable colorectal cancer and Lynch syndrome family colorectal cancer was further clarified, which provided computer-assisted means for their treatment and even cure.
[0086] Example 4: Figure 4 As shown, based on Example 1, the training model establishment module provided in this embodiment of the present invention includes:
[0087] The image acquisition submodule is responsible for acquiring digital images in the labeled training set and training the labeled digital images through deep learning;
[0088] The slice acquisition submodule is responsible for staining and slicing the pathological tissue of surgically removed colorectal cancer. It uses a digital pathology scanner to scan the digital image at 40 times the resolution to obtain digital slices containing pathological information.
[0089] The auxiliary annotation submodule is responsible for detecting the bounding box of the digital slice and capturing only the slice image within the bounding box from the digital slice; performing peripheral contour detection on the slice image to obtain at least one bounding box, and determining the bounding box with the largest area as the corresponding bounding box contained in the digital image; removing the bounding boxes whose area, length, and width are less than an area threshold, length threshold, and width threshold, from the multiple bounding boxes obtained by performing peripheral contour detection on the digital image, and obtaining the digital slice and its peripheral contour after removal; and performing pre-assisted annotation on the digital slice and its peripheral contour using the colorectal cancer diagnostic model;
[0090] The gene labeling submodule is responsible for finding mutant genes from digital slices through convolutional neural networks, extracting the numerical features of the mutant genes, and assigning labels to different mutant genes based on the numerical features;
[0091] The working principle and beneficial effects of the above technical solution are as follows: the image acquisition submodule of this embodiment acquires digital images in the labeled training set, and trains the labeled digital images through deep learning; the slice acquisition submodule stains the pathological tissue slices of surgically removed colorectal cancer, and uses a digital pathology scanner to scan the digital images at 40 times the resolution to obtain digital slices containing pathological information; the auxiliary annotation submodule detects the bounding box of the digital slice, and cuts out the slice image from the digital slice that only contains the slice image within the bounding box; the slice image is subjected to peripheral contour detection to obtain at least one bounding box, and the bounding box with the largest area is determined to be the corresponding bounding box contained in the digital image, and the bounding boxes with an area less than the area threshold, a length less than the length threshold, and a width less than the width threshold are eliminated from the multiple bounding boxes obtained by performing peripheral contour detection on the digital image. The bounding box is used to obtain the digital slices and their outer contours after the removal, and the colorectal cancer diagnostic model is used to pre-annotate the digital slices and their outer contours; the gene label submodule uses a convolutional neural network to find mutant genes from the digital slices, extract the numerical features of the mutant genes, and assign labels to different mutant genes according to the numerical features; the above scheme realizes the accurate labeling of the outer contours of the digital slices by processing the digital slices, and displays the size and range of the pathological tissue, which helps to provide accurate digital images for clinical diagnosis. The colorectal cancer diagnostic model is used to assist in the labeling of the digital slices, realizes the accurate identification of the information in the digital slices, and at the same time finds mutant genes through the convolutional neural network, improves the recognition rate of colorectal cancer tissue and normal tissue, and further ensures the accuracy of the labeling of mutant genes in colorectal cancer.
[0092] Example 5: Figure 5 As shown, based on Example 1, the learning model establishment module provided by the embodiment of the present invention includes:
[0093] The data acquisition submodule is responsible for acquiring multiple digital slices that are assigned different mutation gene labels;
[0094] The data annotation submodule is responsible for obtaining the information of digital slices and auxiliary annotation of their peripheral contours;
[0095] The model training submodule is responsible for training the microsatellite instability colorectal cancer diagnostic model and the Lynch syndrome diagnostic model using a convolutional neural network based on residual networks and classical networks. The convolutional neural network automatically discovers the characteristic information of digital slices from various angles. By extracting features at various levels, it obtains the most essential features related to microsatellite instability and Lynch syndrome, completing automatic identification.
[0096] The service launch submodule is responsible for accumulating digital slides. When the accumulation reaches a set threshold, it uses automated machine learning to build a convolutional neural network to implement the construction of microsatellite instability colorectal cancer diagnosis models, Lynch syndrome diagnosis models, and Lynch syndrome gene mutation mechanism models;
[0097] The working principle and beneficial effects of the above technical solution are as follows: the data acquisition submodule of this embodiment obtains multiple digital slices with different mutation gene labels; the data annotation submodule obtains the information of auxiliary annotation of the digital slices and their peripheral contours; the model training submodule uses a convolutional neural network with residual network and classical network as basic units to train the microsatellite instability colorectal cancer diagnosis model and the Lynch syndrome diagnosis model; the convolutional neural network automatically discovers the feature information of the digital slices from various angles, and obtains the most essential features related to microsatellite instability and Lynch syndrome through feature extraction at various levels, and completes automatic identification; the service online submodule accumulates digital slices, and when the accumulation reaches the set threshold, it uses automation Machine learning constructs a convolutional neural network to realize the construction of microsatellite instability colorectal cancer diagnostic models, Lynch syndrome diagnostic models and Lynch syndrome gene mutation mechanism models; the above scheme uses convolutional neural networks to construct microsatellite instability colorectal cancer diagnostic models, Lynch syndrome diagnostic models and Lynch syndrome gene mutation mechanism models. The convolutional neural network can automatically identify genetic characteristics related to microsatellite instability and Lynch syndrome, thereby improving the accuracy of the result output of the microsatellite instability colorectal cancer diagnostic model and the Lynch syndrome diagnostic model; at the same time, automated machine learning is used to construct a convolutional neural network to provide accuracy guarantee for the prediction of the microsatellite instability colorectal cancer diagnostic model and the Lynch syndrome diagnostic model.
[0098] Example 6: Figure 5 As shown, based on Example 5, the service online submodule provided in this embodiment of the present invention includes:
[0099] The data preparation unit is responsible for selecting multiple digital slices with different mutation gene labels as experimental data sets, and performing preprocessing and normalization preprocessing to obtain the preprocessed experimental data sets;
[0100] The search space unit is responsible for defining the search space, including the basic units of the residual network and the classical network, the convolution kernel size of the feature map, the number of convolution kernels, the activation function, the padding method, and the downsampling method;
[0101] The objective function definition unit is responsible for evaluating the performance of different convolutional neural network models using the cross-validation method and defining the objective function as an indicator for network model optimization;
[0102] Run the automated search unit, which is responsible for searching for convolutional network model parameters and structures using search algorithms and deep learning techniques;
[0103] The model evaluation unit is responsible for verifying the accuracy and generalization ability of the convolutional network model using the preprocessed experimental dataset;
[0104] The working principle and beneficial effects of the above technical solution are as follows: the data preparation unit of this embodiment selects multiple digital slices assigned with different mutant gene labels as experimental data sets, and performs preprocessing and normalization preprocessing to obtain a preprocessed experimental data set; the search space unit defines the search space, including the basic units of the residual network and the classical network, the convolution kernel size of the feature map, the number of convolution kernels, the activation function, the padding method and the downsampling method, etc.; the objective function definition unit uses the cross-validation method to evaluate the performance of different convolutional neural network models, and defines the objective function as an indicator for network model optimization; the automated search unit uses the search algorithm and deep learning technology to search for the parameters and structure of the convolutional network model; the model evaluation unit uses the preprocessed experimental data set to verify the accuracy and generalization ability of the convolutional network model; the automated machine learning of the above scheme is used to construct a convolutional neural network. The residual network and the classical network are the basic units in the convolutional neural network, which are used to construct convolutional neural networks of different types and complexities. In the automated machine learning, an automated tool based on the search algorithm and deep learning technology is used to construct the convolutional neural network, thereby reducing the workload and time cost of manually designing the convolutional neural network model.
[0105] Example 7: Figure 7 As shown, based on Example 1, the learning model establishment module provided in this embodiment of the present invention further includes:
[0106] The image segmentation submodule is responsible for scanning the stained pathological tissue sections of colorectal cancer to obtain full-scan digital slice images. Using a weakly supervised learning framework, the images are divided into image blocks at equal intervals. Each full-scan digital slice image is divided into N×N blocks of equal size, where M and m represent the size of the full-scan digital slice image and the image block, respectively, and N = M / m is the scaling factor.
[0107] The classification label submodule is responsible for finding representative image patches in the image through combined multi-instance learning training. The training prediction result is the classification label of the entire digital slice image;
[0108] The model building submodule is responsible for achieving coarse-grained labeling at the image level of the entire digital slice. Through slice-level classification labels, a deep learning model is established to complete the establishment of prognostic prediction models for microsatellite instability colorectal cancer and Lynch syndrome respectively.
[0109] The working principle and beneficial effects of the above technical solution are as follows: The image slicing sub-module in the embodiment of the present invention scans the pathological tissue stained section of colorectal cancer to obtain the image of the fully scanned digital section, and equally divides the image into image blocks. Each image of the fully scanned digital section is divided into N×N blocks of equal size. M and m respectively represent the sizes of the image of the fully scanned digital section and the image block, and N = M / m is the scaling factor; the classification label sub-module finds the representative image blocks in the image through the training of the combined multi-instance learning, and the prediction result of the training is the classification label of the entire digital section image; the model establishment sub-module realizes the coarse-grained label at the image level of the entire digital section image. Through the slice-level classification label, a deep learning model is established to complete the establishment of the prognostic prediction models for microsatellite instability colorectal cancer and Lynch syndrome respectively; the above solution uses a weakly supervised learning framework to automatically label each image block during the process of enhancing the supervision information, converting the weakly supervised problem into a supervised problem; the method of combined multi-instance learning is adopted to ensure the effectiveness of the weakly supervised learning framework and the quality of the image block labeling after the enhancement of the supervision information, and improve the accuracy of the labeling.
[0110] Embodiment 8: As Figure 8 shown, based on Embodiment 7, the image slicing sub-module provided by the embodiment of the present invention includes:
[0111] The matrix establishment unit is responsible for establishing an attention matrix for each image of the fully scanned digital section, and the attention weights corresponding to all image blocks are stored in the matrix;
[0112] The probability prediction unit is responsible for using the weight storage classification model to predict the probability of the image block. If it is not predicted, the weight is defaulted to N / A, where A represents the image block; the defined sampling number is n (n << N). In each round of iteration, a certain probability p (ranging from 0.0 to 1.0) is used to sample the image blocks according to the attention weights, that is, an average of p×n image blocks with non-N / A attention weights are sampled;
[0113] The image sampling unit is responsible for randomly sampling in the image blocks with N / A weights with a probability of 1 - p, that is, an average of (1 - p)×n image blocks with non-N / A attention weights are sampled. A total of n image blocks are sampled for each image of the fully scanned digital section in each iteration for the training of the classification model;
[0114] The working principle and beneficial effects of the above technical solution are as follows: The matrix establishment unit in this embodiment establishes an attention matrix for the image of each full-scanned digital slice, and the attention weights corresponding to all image patches are stored in the matrix; the probability prediction unit uses the weight storage classification model to predict the probability of the image patch. If it is not predicted, the weight is defaulted to N / A, where A represents the image patch; define the sampling number as n (n << N). In each round of iteration, with a certain probability p (between 0.0 and 1.0), sample the image patches according to the attention weights, that is, on average, sample p × n image patches with non-N / A attention weights; the image sampling unit samples randomly among the image patches with N / A weights with a probability of 1 - p, that is, on average, sample (1 - p) × n image patches with non-N / A attention weights. Each time an iteration is performed, a total of n image patches of the image of a single full-scanned digital slice are sampled for the classification model training; the above solution adopts an image patch sampling method based on the attention mechanism, realizing accurate sampling of image patches, not only obtaining representative image patches, but also improving the accuracy of the analysis of the colorectal cancer mechanism and ensuring the representativeness of the image of the full-scanned digital slice.
[0115] Example 9: As Figure 9 shown, on the basis of Example 7, the classification label sub-module provided by the embodiment of the present invention includes:
[0116] The hot spot area unit is responsible for regarding all the image patches sampled from the image of a single full-scanned digital slice as an image packet, and defining the area related to the predicted label as the hot spot area;
[0117] The area judgment unit is responsible for inferring that at least one image patch contains a hot spot area if the image of a single full-scanned digital slice contains a hot spot area; if there is no hot spot area in the image of a single full-scanned digital slice, then none of the image patches have a hot spot;
[0118] The selection criterion unit is responsible for using two different image patch selection criteria, Max-Max and Max-Min, through the combination of multi-instance learning. For the images of full-scanned digital slices with and without hot spot areas, Max-Max uses the image patch with the maximum prediction probability as the representative image patch, while Max-Min uses the image patches with the maximum and minimum prediction probabilities as the representatives for the images of full-scanned digital slices with and without hot spot areas respectively;
[0119] The working principle and beneficial effects of the above technical solution are as follows: the hotspot area unit of this embodiment refers to all image blocks sampled from a single fully scanned digital slide image as an image package, and defines the area related to the predicted label as a hotspot area; the area judgment unit is responsible for inferring whether at least one image block in a fully scanned digital slide image contains a hotspot area; if there is no hotspot area in a fully scanned digital slide image, then all image blocks do not have a hotspot; the selection criterion unit uses two different image block selection criteria, Max-Max and Max-Min, through combined multi-instance learning. Max-Max uses the image block with the largest prediction probability as the representative image block for fully scanned digital slide images with and without hotspot areas, while Max-Min uses the image blocks with the maximum and minimum prediction probabilities as representatives for fully scanned digital slide images with and without hotspot areas, respectively; the above solution obtains representative image blocks in the fully scanned digital slide image through sampling, improves the accuracy of the prediction result, provides a reference basis for the classification label of the entire digital slide image, ensures the effectiveness of the deep learning model, and provides a reliable data foundation for the establishment of a prognosis prediction model for microsatellite instability colorectal cancer and Lynch syndrome.
[0120] Example 10: Figure 10 As shown, based on Example 7, the classification label submodule provided by this embodiment of the present invention further includes:
[0121] The image block selection unit is responsible for selecting an image block from each image package using Max-Max or Max-Min, and calculating the cost function based on the image block's prediction results and image labels. Two deep learning models are trained using the Max-Max or Max-Min criteria respectively. Image blocks from all training full-scan digital slide images are fed into the two deep learning models to obtain prediction results for each image block. The most representative image blocks are selected using the corresponding criteria, and image blocks with different prediction results from the two deep learning models are excluded. A convolutional neural network is used as the classifier.
[0122] The image block prediction unit is responsible for training a new classifier using the selected annotated image blocks. The trained classifier is then used to predict image blocks in all fully scanned digital slides in the training set, obtaining rough hotspot prediction results for each fully scanned digital slide. The image-level annotations of the fully scanned digital slides are enhanced to the image block level, obtaining N×N times more supervised information.
[0123] The model training unit is responsible for assigning image block-level labels directly to each pixel and using image segmentation models such as deep learning networks and convolutional neural networks (all types of image segmentation models are applicable) to train pixel-level prediction models in a supervised manner.
[0124] The working principle and beneficial effects of the above technical solution are as follows: the image block selection unit of this embodiment uses Max-Max or Max-Min to select an image block from each image package, and calculates the cost function according to the prediction result of the image block and the image label; the two Max-Max or Max-Min criteria are trained to obtain two deep learning models respectively, and the image blocks in the images of all training full-scan digital slices are sent to the two deep learning models respectively to obtain the prediction results of each image block, and the most representative image blocks are selected by the corresponding criteria, and the image blocks with different prediction results of the two deep learning models are excluded, and the convolutional neural network is used as the classifier; the image block prediction unit uses the selected labeled image blocks to train a new classifier, and uses the trained classifier to obtain the predicted image blocks. The trained classifier predicts the image blocks in all the fully scanned digital slide images in the training set, and obtains a rough prediction result of the hot spot area of each fully scanned digital slide image; the image-level annotation of the fully scanned digital slide is enhanced to the image block level, and N×N times of supervised information is obtained; the model training unit directly assigns the image block level annotation to each pixel point, and uses image segmentation models such as deep learning networks and convolutional neural networks (all kinds of image segmentation models are applicable) to train the pixel-level prediction model in a supervised manner; the above scheme improves the effective prediction of image block recognition, realizes the prediction of image blocks through the establishment of deep learning models, and realizes the judgment of image block representativeness through criteria, effectively realizing the prediction of digital slide images.
[0125] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.
Claims
1. A deep learning-based microsatellite instability colorectal cancer mechanism information analysis system, characterized by: include: The data set establishment module is responsible for collecting the first case document data of surgically resected colorectal cancer, screening the second case document data of microsatellite unstable colorectal cancer from the first case document data, wherein the second case document data includes: the third case document data of colorectal cancer in a family with Lynch syndrome, and selecting data from the third case document data as the training data set and test data set of the artificial intelligence deep learning model; The gene screening module is responsible for screening the mutant genes of colorectal cancer in Lynch syndrome families from the second case file data and the third case file data; The training model building module is responsible for constructing a digital slice training model based on the training data set. It scans stained pathological tissue sections of surgically removed colorectal cancer to obtain digital slices. It then uses the colorectal cancer diagnostic model to pre-annotate the digital slices and assign labels for different mutated genes to the entire digital slice. The learning model building module is responsible for building an artificial intelligence deep learning model, testing the artificial intelligence deep learning model using a test data set, inputting labels into the artificial intelligence deep learning model, and obtaining microsatellite instability colorectal cancer and Lynch syndrome family colorectal cancer, which cause inherited germline mutation genes; Learning model building module, including: The image segmentation submodule is responsible for scanning stained colorectal cancer pathological tissue sections to obtain fully scanned digital slice images. Using a weakly supervised learning framework, the images are divided into image blocks at equal intervals. Each fully scanned digital slice image is divided into N×N blocks of equal size, where M and m represent the size of the fully scanned digital slice image and the image block, respectively, and N = M / m is the scaling factor. The classification label submodule is responsible for finding representative image patches in the image through combined multi-instance learning training. The training prediction result is the classification label of the entire digital slice image; The model building submodule is responsible for achieving coarse-grained labeling at the image level of the entire digital slice. Through slice-level classification labels, a deep learning model is established to complete the establishment of prognostic prediction models for microsatellite instability colorectal cancer and Lynch syndrome respectively. Image segmentation submodule, including: The matrix building unit is responsible for building an attention matrix for each fully scanned digital slice image. The matrix stores the attention weights corresponding to all image blocks. The probability prediction unit is responsible for using weights to store the predicted probability of the classification model for the image block. If no prediction is made, the weight defaults to N / A, where A represents the image block. The sampling number is defined as n, where n << N. In each iteration, the image block is sampled according to the attention weight with probability p, which is between 0.0 and 1.
0. That is, on average, p×n image blocks with attention weights not N / A are sampled. The image sampling unit is responsible for randomly sampling from image blocks with a weight of N / A with a probability of 1-p, that is, sampling an average of (1-p)×n image blocks with non-N / A attention weights. In each iteration, a single fully scanned digital slice image is sampled with a total of n image blocks for classification model training.
2. The deep learning-based microsatellite instability colorectal cancer mechanism information analysis system according to claim 1, characterized in that: Artificial intelligence deep learning models include: microsatellite instability colorectal cancer diagnosis model, Lynch syndrome diagnosis model, Lynch syndrome gene mutation mechanism model, microsatellite instability colorectal cancer and Lynch syndrome prognosis prediction model.
3. The deep learning-based microsatellite instability colorectal cancer mechanism information analysis system according to claim 1, characterized in that: Dataset creation module, including: The first data screening submodule is responsible for acquiring first case document data stored in the server through network crawling, and establishing a database to be screened based on the first case document data using the gene characteristics of colorectal cancer removed by surgical operation as the first gene attribute characteristics and the first gene attribute characteristics as the feature identifier; The second data screening submodule is responsible for using the microsatellite instability colorectal cancer gene feature as the second gene attribute feature and as the screening condition to screen from the database to be screened to obtain the second case document data; The third data screening submodule is responsible for receiving the second case document data, evaluating the importance of the second case document data according to the third gene attribute feature, obtaining an importance evaluation score, and establishing the third case document data with the third gene attribute feature as a feature identifier; the colorectal cancer gene feature of the Lynch syndrome family is used as the third gene attribute feature; The dataset requirement submodule is responsible for selecting case document data from the third case document data according to the dataset requirements of the artificial intelligence deep learning model, and constructing the training dataset and test dataset.
4. The deep learning-based microsatellite instability colorectal cancer mechanism information analysis system according to claim 1, characterized in that: Gene screening module, including: The first screening submodule is responsible for obtaining the genetic characteristics of microsatellite instability colorectal cancer and the genetic characteristics of colorectal cancer in Lynch syndrome families, and screening out important genes from the genetic characteristics; The second screening submodule is responsible for functional analysis of important genes and screening the gene loci in important genes; The contribution calculation submodule is responsible for performing principal component analysis on the gene loci in important genes using the principal component analysis method to obtain the functional contribution of each important gene in the gene feature, and querying the functional contribution of non-important genes from the data contribution of each important gene in the gene feature; The third screening module is responsible for screening the mutated genes of colorectal cancer in Lynch syndrome families from important genes based on the functional contribution of each important gene.
5. The deep learning-based microsatellite instability colorectal cancer mechanism information analysis system according to claim 1, characterized in that: Training model building module, including: The image acquisition submodule is responsible for acquiring digital images in the labeled training set and training the labeled digital images through deep learning; The slice acquisition submodule is responsible for staining and slicing the pathological tissue of surgically removed colorectal cancer. It uses a digital pathology scanner to scan the digital image at 40 times the resolution to obtain digital slices containing pathological information. The auxiliary annotation submodule is responsible for detecting the bounding box of the digital slice and capturing only the slice image within the bounding box. The module performs peripheral contour detection on the slice image to obtain at least one bounding box, and determines the bounding box with the largest area as the corresponding bounding box contained in the digital image. The digital slice and its peripheral contour are then pre-annotated using the colorectal cancer diagnostic model. The gene label submodule is responsible for finding mutant genes from digital slices through convolutional neural networks, extracting the numerical features of mutant genes, and assigning labels to different mutant genes according to the numerical features.
6. The deep learning-based microsatellite instability colorectal cancer mechanism information analysis system according to claim 5, characterized in that: Among the multiple bounding boxes obtained by performing peripheral contour detection on the digital image, the bounding boxes whose areas are smaller than an area threshold, whose lengths are smaller than a length threshold, and whose widths are smaller than a width threshold are eliminated.
7. The deep learning-based microsatellite instability colorectal cancer mechanism information analysis system according to claim 1, characterized in that: Learning model building module, including: The data acquisition submodule is responsible for acquiring multiple digital slices that are assigned different mutation gene labels; The data annotation submodule is responsible for obtaining the information of digital slices and auxiliary annotation of their peripheral contours; The model training submodule is responsible for training the microsatellite instability colorectal cancer diagnostic model and the Lynch syndrome diagnostic model using a convolutional neural network based on residual networks and classical networks. The convolutional neural network automatically discovers the characteristic information of digital slices from various angles. By extracting features at various levels, it obtains the most essential features related to microsatellite instability and Lynch syndrome, completing automatic identification. The service launch sub-module is responsible for accumulating digital slices. When the accumulation reaches a set threshold, it uses automated machine learning to build a convolutional neural network to realize the construction of microsatellite instability colorectal cancer diagnosis model, Lynch syndrome diagnosis model and Lynch syndrome gene mutation mechanism model.
8. The deep learning-based microsatellite instability colorectal cancer mechanism information analysis system according to claim 7, characterized in that: The service launch sub-module includes: The data preparation unit is responsible for selecting multiple digital slices with different mutation gene labels as experimental data sets, and performing preprocessing and normalization preprocessing to obtain the preprocessed experimental data sets; The search space unit is responsible for defining the search space, including the basic units of the residual network and the classical network, the convolution kernel size of the feature map, the number of convolution kernels, the activation function, the padding method, and the downsampling method; The objective function definition unit is responsible for evaluating the performance of different convolutional neural network models using the cross-validation method and defining the objective function as an indicator for network model optimization; Run the automated search unit, which is responsible for searching for convolutional network model parameters and structures using search algorithms and deep learning techniques; The model evaluation unit is responsible for verifying the accuracy and generalization ability of the convolutional network model using the preprocessed experimental dataset.
Citation Information
Patent Citations
Microsatellite instability detection system based on colorectal pathology image
CN114612455A
System and image segmentation method for colorectal cancer microsatellite instability prediction
CN115760694A
Lung cancer pathological diagnosis system based on deep transfer learning
CN116309368A
Colorectal cancer artificial intelligence auxiliary pathological diagnosis system based on deep learning
CN116386902A