Tumor molecular typing recognition system and method based on neural network
Through the neural network-based identification system, the DenSeNet121 model is used to identify tumor molecular typing, which solves the problem of high cost of tumor molecular typing in the prior art, and achieves rapid and accurate diagnosis and popularization.
Patent Information
- Application Number
- CN202510094992.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, tumor molecular typing and identification problems are high in human, financial and time costs.
Using a neural network-based recognition system, a full-field digital pathological section of hematoxylin-eosin-stained smooth muscle tumor was obtained, and pre-treatment and molecular typing recognition model was trained, and the DenSeNet121 model was used for identification.
The molecular typing of LMS is achieved quickly and accurately diagnosed, which reduces human and resource costs, improves popularity, and simplifies the knowledge reserve requirements for the diagnoser.
Smart Images

Figure CN119992546A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical image analysis and machine learning technology, and in particular to a tumor molecular typing recognition system and method based on a neural network. Background Art
[0002] Tumor molecular classification is extremely important in guiding clinical medication and judging patient prognosis. Leiomyosarcoma (LMS) is the most common pathological type of soft tissue sarcoma, accounting for about 20%-24% of all soft tissue sarcomas. It is prone to occur in the abdominal wall, retroperitoneal space, large blood vessels and uterus, and is also seen in the limbs, spine and surrounding soft tissues. Previous studies based on high-throughput sequencing and immunohistochemistry (Immunochemistry, IHC) have found that LMS can be further divided into three molecular subtypes: (1) Subtype I: IHC mainly shows positive ACTG2, SLMAP, LMOD1, CFL2 and MYLK, and patients with this subtype have the best long-term prognosis among LMS (except uterine LMS); (2) Subtype II: IHC mainly shows positive ARL4C, and patients with this subtype have the worst prognosis among all LMS; (3) Subtype III: All LMS except subtypes I and II, and the prognosis of patients with this subtype is in the middle among all LMS.
[0003] Given that tumor molecular typing has important clinical significance in revealing tumor tissue heterogeneity, determining the sensitivity of targeted drugs, and predicting the long-term prognosis of patients, it is necessary to diagnose LMS molecular typing routinely, quickly, and effectively in clinical practice. However, the current LMS molecular typing identification method based on high-throughput sequencing has the disadvantages of complex operation, high manpower and time costs, high cost, and high requirements on the knowledge reserve of pathologists, and is not widely used. Summary of the invention
[0004] Based on the above content, the present invention provides a tumor molecular typing identification system and method based on neural network, which only solves the technical problem of high manpower, financial and time costs in tumor molecular typing identification in the prior art.
[0005] A method for identifying tumor molecular typing based on a neural network, used for identifying smooth muscle tumor molecular typing, comprising:
[0006] Step A1, obtaining a full-field digital pathological section of a smooth muscle tumor to be identified stained with hematoxylin-eosin as a section to be identified;
[0007] Step A2, preprocessing the slice to be identified to obtain image data to be identified formed by the preprocessing;
[0008] Step A3, inputting the image data to be identified into a pre-trained molecular typing recognition model, and the molecular typing recognition model processes the image data to be identified to obtain a molecular typing recognition result of smooth muscle tumor.
[0009] Furthermore, in step A2, the process of preprocessing the slice to be identified includes:
[0010] Step A21, performing oncological annotation on the slice to be identified, marking the region with oncological components in the slice to be identified as a region of interest;
[0011] Step A22, gridding the region of interest into a plurality of image blocks to be identified;
[0012] Step A23, performing Laplace filtering, regularization and signal enhancement processing on the image block to be identified to obtain image data to be identified.
[0013] Furthermore, the molecular typing recognition model is constructed based on the DenSeNet121 model, including four dense blocks, and two adjacent dense blocks are directly connected through a transition layer.
[0014] Furthermore, the training process of the molecular typing recognition model in step A3 includes:
[0015] Step A31, obtaining a number of full-field digital pathological sections of smooth muscle tumors stained with hematoxylin-eosin from a public database as training samples, and obtaining corresponding transcriptome sequencing data from the public database;
[0016] Step A32, performing cluster analysis on the transcriptome test data to determine the molecular typing of the training samples, annotating the molecular typing of the training samples, and preprocessing the training samples to form training data;
[0017] Step A33, using the training data to train the initially constructed neural network model to obtain a molecular typing recognition model.
[0018] Furthermore, in step A33, the molecular typing recognition model is tested using a test set, and the test set formation process includes:
[0019] Step A331, obtaining full-field digital pathological sections of smooth muscle tumors stained with hematoxylin-eosin from several smooth muscle tumor patients in the hospital as test samples, and obtaining corresponding immunohistochemistry data;
[0020] Step A332, based on immunohistochemical data analysis, the tumor molecular typing of each smooth muscle tumor patient is determined, and the molecular typing of the validation samples is annotated, and the annotated test samples are preprocessed to obtain a test set.
[0021] A tumor molecular typing recognition system based on a neural network, used in the aforementioned tumor molecular typing recognition method based on a neural network, comprising:
[0022] A first data acquisition module is used to acquire a full-field digital pathological section of a smooth muscle tumor to be identified stained with hematoxylin-eosin as a section to be identified;
[0023] A preprocessing module, connected to the first data acquisition module, preprocesses the slice to be identified to obtain image data to be identified formed by the preprocessing;
[0024] The molecular typing recognition module is connected to the preprocessing module and is used to input the image data to be recognized into the pre-trained molecular typing recognition model. The molecular typing recognition model processes the image data to be recognized to obtain the molecular typing recognition result of the smooth muscle tumor.
[0025] Furthermore, the preprocessing module includes:
[0026] A region extraction unit is used to annotate the slice to be identified by oncology, and mark the region with oncology components in the slice to be identified as a region of interest;
[0027] The gridding unit is connected to the region extraction unit and is used to grid the region of interest and divide it into a number of image blocks to be identified;
[0028] The block optimization unit is connected to the gridding unit and is used to perform Laplace filtering, regularization and signal enhancement processing on the image block to be identified to obtain the image data to be identified.
[0029] Furthermore, the molecular typing recognition model is constructed based on the DenSeNet121 model, including four dense blocks, and two adjacent dense blocks are directly connected through a transition layer.
[0030] Furthermore, the system also includes a second data acquisition module, a data analysis module, a sample annotation module, and a model training module;
[0031] The second data acquisition module is used to obtain a number of full-field digital pathological sections of smooth muscle tumors stained with hematoxylin-eosin from a public database as training samples, and to obtain corresponding transcriptome sequencing data from the public database;
[0032] A data analysis module, connected to the second data acquisition module, is used to perform cluster analysis on the transcriptome test data to determine the molecular typing of the training samples;
[0033] A sample labeling module, connected to the data analysis module and the first data acquisition module, respectively, for labeling the training samples by molecular typing;
[0034] The preprocessing module is connected to the sample labeling module and is used to preprocess the labeled training samples to form training data;
[0035] The model training module is used to train the initially constructed neural network model using training data to obtain a molecular typing recognition model.
[0036] Further, the system also includes a third data acquisition module and a model testing module;
[0037] The third data acquisition module is used to obtain full-field digital pathological sections of smooth muscle tumors stained with hematoxylin-eosin of several smooth muscle tumor patients in the hospital as test samples;
[0038] The sample annotation module is also connected to the third data acquisition module, and is used to determine the tumor molecular typing of each smooth muscle tumor patient according to the immunohistochemical data analysis, so as to perform molecular typing annotation on the corresponding test samples;
[0039] A preprocessing module is used to preprocess the labeled test samples to form a test set;
[0040] The model testing module is connected to the preprocessing module and is used to test the molecular typing recognition model using the test set.
[0041] The beneficial technical effect of the present invention is that by utilizing a single full-field digital pathology slice image (Whole slide image, WSI) and establishing a deep learning model based on CNN, the molecular typing of LMS can be quickly and accurately diagnosed, which is beneficial to predicting the clinical prognosis of LMS patients and bringing long-term benefits to LMS patients. At the same time, it also provides an accurate, convenient, cost-effective and effective measure for widely carrying out clinical identification of LMS molecular typing and judgment of long-term prognosis in hospitals at all levels. The operation is simple, the knowledge reserve requirements for the diagnostician are reduced, the manpower and resource costs are reduced, and the popularity is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1-Figure 4 This is a flowchart of the steps of a method for identifying tumor molecular typing based on a neural network according to the present invention;
[0043] Figure 5-Figure 6 The figure is a module schematic diagram of a tumor molecular typing recognition system based on a neural network according to the present invention.
[0044] Figure 7-A A comparison result diagram of the accuracy evaluation index after selecting multiple deep learning models for the present invention;
[0045] Figure 7-B A comparison result diagram of the area under the ROC curve evaluation index after selecting multiple deep learning models for training in the present invention;
[0046] Figure 7-C The confusion matrix diagram of LMS patients from the CCY cohort expressed in absolute numbers obtained by the present invention according to the molecular typing recognition model;
[0047] Figure 7-D The confusion matrix diagram of LMS patients from the CCY cohort expressed in percentages obtained by the present invention according to the molecular typing recognition model;
[0048] Figure 7-E It is a display diagram of the performance index Accuracy calculated by the present invention according to the micro-average method and the macro-average method, as well as the performance index Accuracy corresponding to each molecular typing;
[0049] Figure 7-F It is a display diagram of the performance index Recall calculated by the present invention according to the micro-average method and the macro-average method, as well as the performance index Recall corresponding to each molecular typing;
[0050] Figure 7-G It is a display diagram of the performance index Specificity calculated by the present invention according to the micro-average method and the macro-average method and the performance index Specificity corresponding to each molecular typing;
[0051] Figure 7-H It is a display diagram of the performance index Precision calculated by the present invention according to the micro-average method and the macro-average method, as well as the performance index Precision corresponding to each molecular typing;
[0052] Figure 7-I It is a display diagram of the performance indicator F1Score calculated by the present invention according to the micro-average method and the macro-average method, and the performance indicator F1 Score corresponding to each molecular typing;
[0053] Figure 7-J The figure shows the performance index AUROC calculated by the present invention according to the micro-average method and the macro-average method, as well as the performance index AUROC corresponding to each molecular typing;
[0054] Figure 7-K It is a display diagram of the ROC curve calculated by the micro-average method and the macro-average method of the present invention;
[0055] Figure 7-L A display diagram of the ROC curve of molecular subtype I calculated by the present invention at the segmented image block level;
[0056] Figure 7-M A display diagram of the ROC curve of molecular subtype II calculated by the present invention at the segmented image block level;
[0057] Figure 7-N A display diagram of the ROC curve of molecular subtype III calculated by the present invention at the segmented image block level;
[0058] Figure 8-A It shows a visual comprehensive display of the prediction possibility at each segmented image block level for a slice of a patient representative of each molecular subtype selected from the CCY cohort;
[0059] Figure 8-B It indicates that two typical patients of each molecular subtype were selected from the CCY cohort, and the SWI, segmented image block results, CAM heat map and pathological features were displayed;
[0060] Figure 9-A to Figure 9-B It represents the comparison chart of the evaluation index of molecular subtype classification based on the molecular typing recognition model DL_Enabled and pathologist only;
[0061] Figure 9-C A comparison chart showing the evaluation indicators of molecular subtype classification by pathologists alone and pathologists with the assistance of the model;
[0062] Figure 9-D A comparison chart showing the time cost of molecular subtype classification by pathologists alone and pathologists with the assistance of the model, as well as the number of repeated diagnoses of the cut image blocks;
[0063] Figure 9-E A comparison chart showing the confusion matrix of molecular typing recognition by pathologists and the confusion matrix of molecular typing recognition assisted by the model;
[0064] Figure 9-F The figure shows the comparison of the ROC curve obtained by pathologists' molecular typing and the ROC curve obtained by using the model to assist molecular typing. DETAILED DESCRIPTION
[0065] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0066] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.
[0067] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, but they are not intended to limit the present invention.
[0068] See also Figure 1 The present invention provides a method for identifying tumor molecular typing based on a neural network, which is used to identify the molecular typing of smooth muscle tumors, comprising:
[0069] Step A1, obtaining a full-field digital pathological section of a smooth muscle tumor to be identified stained with hematoxylin-eosin as a section to be identified;
[0070] Step A2, preprocessing the slice to be identified to obtain image data to be identified formed by the preprocessing;
[0071] Step A3, inputting the image data to be identified into a pre-trained molecular typing recognition model, and the molecular typing recognition model processes the image data to be identified to obtain a molecular typing recognition result of smooth muscle tumor.
[0072] The inventors analyzed RNA-seq data of 95 LMS patients from the TCGA database and found that the molecular subtypes of LMS are also highly correlated with the degree of differentiation of tumor tissues, showing that the relative risk (RR) of poor differentiation of LMS subtype II compared with other subtypes (I and III) is the highest, while the relative risk of poor differentiation of LMS subtype I is the lowest. Therefore, the present invention uses a single HE-stained LMS pathological section to establish a CNN-based deep learning model for the identification of smooth muscle tumor molecular typing, which has high accuracy, simple operation, short time consumption, low manpower investment, low capital cost, and low requirements for pathologists. It can be widely used in routine pathological diagnosis in hospitals at all levels, thereby benefiting LMS patients in the long term.
[0073] See also Figure 2 , further, in step A2, the process of preprocessing the slice to be identified includes:
[0074] Step A21, performing oncological annotation on the slice to be identified, marking the region with oncological components in the slice to be identified as a region of interest;
[0075] Step A22, gridding the region of interest into a plurality of image blocks to be identified;
[0076] Step A23, performing Laplace filtering, regularization and signal enhancement processing on the image block to be identified to obtain image data to be identified.
[0077] Through oncology annotation, the oncology components in the pathology slides were marked as the region of interest (Region of Interest). Focus on the calculation of the area of interest to reduce the amount of calculation.
[0078] Specifically, in step A22, the region of interest is divided into image blocks according to a predetermined pixel size, such as 512×512, through gridding, and there is partial overlap between the image blocks.
[0079] Laplacian filtering is an edge detection technique that can enhance the edges and details of an image and is used for denoising to make the image clearer. Regularization and signal enhancement can further improve image quality.
[0080] In step A3, the image data to be identified is composed of preprocessed image blocks to be identified. The molecular typing recognition model first performs molecular typing recognition on each preprocessed image block to be identified, and then generates a heat map using the CAM (class activation mapping) technology.
[0081] Furthermore, the molecular typing recognition model is constructed based on the DenSeNet121 model, including four dense blocks, and two adjacent dense blocks are directly connected through a transition layer.
[0082] The molecular typing recognition model includes an input layer, a convolution layer, a pooling layer, a first dense block, a first transition layer, a second dense block, a second transition layer, a third dense block, a third transition layer, a fourth dense block, a fourth transition layer, a global average pooling layer, and a classifier in sequence; the input layer is used to input image data. The convolution layer uses a 7×7 convolution kernel to perform a convolution operation on the input image with a step size of 2, and the pooling layer uses a 3×3 maximum pooling layer for downsampling with a step size of 2 to reduce the size of the feature map. The transition layer is located between two adjacent dense blocks and is responsible for reducing the size of the feature map. The global average pooling layer converts the feature map output by the last dense block into a vector of fixed length. The classifier is usually a fully connected layer plus a softmax activation function for the final classification category output.
[0083] See also Figure 3 Furthermore, the training process of the molecular typing recognition model in step A3 includes:
[0084] Step A31, obtaining a number of full-field digital pathological sections of smooth muscle tumors stained with hematoxylin-eosin from a public database as training samples, and obtaining corresponding transcriptome sequencing data from the public database;
[0085] Step A32, performing cluster analysis on the transcriptome test data to determine the molecular typing of the training samples, annotating the molecular typing of the training samples, and preprocessing the training samples to form training data;
[0086] Step A33, using the training data to train the initially constructed neural network model to obtain a molecular typing recognition model.
[0087] Specifically, the public database is the TCGA database, from which the whole-view digital pathology sections (WSI) and transcriptome sequencing data of leiomyosarcoma are obtained. The preprocessing of the training samples is the same as the preprocessing process of the identification sections, including oncology annotation, gridding, Laplace filtering, regularization, and signal enhancement. The transcriptome test data is analyzed by existing clustering analysis methods to obtain the molecular typing of the training samples to annotate the training samples. Specifically, an unsupervised hierarchical clustering (UHC) method such as an average linkage algorithm is used to analyze the transcriptome sequencing data (RNA sequencing data). The training data is divided into a training set and a validation set, the training set trains the initially constructed model, and the validation set performs parameter tuning on the training model.
[0088] See also Figure 4 Furthermore, in step A33, the molecular typing recognition model is tested using a test set, and the test set formation process includes:
[0089] Step A331, obtaining full-field digital pathological sections of smooth muscle tumors stained with hematoxylin-eosin from several smooth muscle tumor patients in the hospital as test samples, and obtaining corresponding immunohistochemistry data;
[0090] Step A332, based on immunohistochemical data analysis, the tumor molecular typing of each smooth muscle tumor patient is determined, and the molecular typing of the validation samples is annotated, and the annotated test samples are preprocessed to obtain a test set.
[0091] The preprocessing of the test samples was the same as that of the slices to be identified, including oncological annotation, gridding, Laplacian filtering, regularization, and signal enhancement.
[0092] In step A331, full-field digital pathological sections of smooth muscle tumors stained with hematoxylin-eosin from several smooth muscle tumor patients in the hospital are obtained as test samples. These test samples are called CCY cohorts.
[0093] The present invention initially constructs five commonly used neural network architectures, SqueezeNet, DenSeNet121, ResNet50, GoogLeNet, and AlexNet, for training, and then tests them with a test set. The evaluation indicators of the training model using the test set include the area under the ROC curve (AUROC), recall, specificity, precision, F1 score (F1 test value), accuracy, etc. Figure 7-A and Figure 7-BFrom the perspective of the area under the ROC curve and accuracy, after repeated experiments and training of various CNN (convolutional neural network)-based DL models (deep learning models), DenSeNet121 is superior to the other four architecture models. Therefore, the present invention prefers DenSeNet121 as the molecular typing recognition model.
[0094] like Figure 7-C The confusion matrix shown in absolute numbers is Figure 7-D The confusion matrix shown in percentage. The horizontal axis Predicted Subtype represents the three molecular subtypes (I, II, III) predicted by the molecular typing recognition model, and the vertical axis Ground truth represents the true molecular subtype (I, II, III) category. It can be seen from the confusion matrix that the molecular typing categories predicted by the molecular typing recognition model of the present invention and the true molecular typing categories have a high degree of overlap.
[0095] See also Figure 7-E to Figure 7-J , which is a display diagram of the performance indicators Accuracy, Recall, Specificity, Precision, F1 Score, AUROC calculated by the present invention according to the micro-average method and the macro-average method, and the performance indicators Accuracy, Recall, Specificity, Precision, F1 Score, AUROC corresponding to each molecular typing. It can be seen from the figure that the evaluation of each performance indicator of the molecular typing recognition model is relatively good.
[0096] Figure 7-J to Figure 7-N The ROC curves calculated by the present invention according to the micro-average method and the macro-average method and the ROC curves for molecular subtypes I, II, and III are shown in the figure. The horizontal axis False Positive Rate represents the false positive rate, and the vertical axis True Positive Rate represents the positive rate. The ROC curves of each figure are close to the upper left corner, which shows that the molecular typing recognition model of the present invention has better performance.
[0097] Figure 8-AA visualization diagram showing the predicted probability of each segmented image block level according to the slice ASI-HE of a typical patient of each molecular subtype selected from the CCY queue, green represents molecular subtype I, orange represents molecular subtype II, blue represents molecular subtype III, and All Class represents the overall display of the recognition results of each image block to be recognized according to the molecular typing recognition model of the present invention represented by color. A represents the display result of the first selected patient. From the perspective of ALL class, green accounts for the vast majority, and Subtype I probability indicates that the possibility of molecular subtype I is the greatest. B represents the display result of the second selected patient. From the perspective of ALL class, orange accounts for the vast majority, and SubtypeII probability indicates that the possibility of molecular subtype II is the greatest. C represents the display result of the third selected patient. From the perspective of ALL class, blue accounts for the vast majority, and Subtype III probability indicates that the possibility of molecular subtype III is the greatest.
[0098] Figure 8-B It indicates that two typical SWIs of patients of each molecular subtype were selected from the CCY cohort, the recognition results of the segmented image blocks, the heat map generated by the CAM technology, and the display of pathological features.
[0099] Figure 9-A to Figure 9-B It represents the comparison chart of the evaluation indicators of molecular subtype classification based on the molecular typing recognition model DL_Enabled and pathologist only. Figure 9-A It is an accuracy evaluation index, which uses the accuracy evaluation index obtained by the micro-average method and the accuracy evaluation index obtained for each molecular subtype, to compare the molecular typing predicted by the model and the molecular typing classification by the pathologist. Figure 9-B It is a comparison of Recall, Specificity, Precision, and F1 Score between the molecular typing predicted by the model and the molecular typing classification by the pathologist. As can be seen from the figure, from the evaluation of various performance indicators, the recognition of molecular typing using the model is much better than manual recognition.
[0100] Figure 9-C This figure shows the comparison of the evaluation indicators of molecular subtype classification by pathologists only and pathologists with the assistance of the model. As can be seen from the figure, the performance indicators of model-assisted recognition have been greatly improved.
[0101] Figure 9-DThe figure shows a comparison of the time cost for molecular subtype classification by a pathologist alone and by a pathologist with the assistance of a model, as well as the number of repeated diagnoses of the cut image blocks. It can be seen from the figure that after model-assisted recognition, the time cost of recognition is reduced, and the number of image blocks that need to be repeatedly recognized is reduced.
[0102] Figure 9-E A comparison chart showing the confusion matrix of molecular typing recognition by pathologists and the confusion matrix of molecular typing recognition assisted by models. The left side shows the pathologist before using the model assistance, and the right side shows the pathologist after using the model assistance. It can be seen from the figure that the recognition effect is improved after using the model assistance.
[0103] Figure 9-F The figure shows the comparison of the ROC curve obtained by pathologists' molecular typing and the ROC curve obtained by model-assisted molecular typing. #1, #2, #3, #4, and #5 represent five pathologists. On the left are the molecular typing results obtained by the five pathologists before the model was used, and on the right are the molecular typing results obtained by the five pathologists after the model was used. The results before the model was used are biased to the lower side, and after the model was used, they are biased to the upper left side. It can be seen that the ROC performance has improved with the assistance of the model.
[0104] See also Figure 5 The present invention also provides a tumor molecular typing recognition system based on a neural network, which is used to execute the aforementioned tumor molecular typing recognition method based on a neural network, comprising:
[0105] A first data acquisition module (1) is used to acquire a full-field digital pathological section of a smooth muscle tumor to be identified, stained with hematoxylin-eosin, as a section to be identified;
[0106] A preprocessing module (2) is connected to the first data acquisition module (1) to preprocess the slice to be identified to obtain image data to be identified formed by the preprocessing;
[0107] The molecular typing recognition module (3) is connected to the preprocessing module (2) and is used to input the image data to be recognized into a pre-trained molecular typing recognition model. The molecular typing recognition model processes the image data to be recognized to obtain a molecular typing recognition result of a smooth muscle tumor.
[0108] Using a single HE-stained LMS pathological section, a CNN-based deep learning model was established to identify the molecular typing of smooth muscle tumors. This model has high accuracy, simple operation, short time consumption, low manpower investment, and low financial cost. It also has low requirements for pathologists and can be widely used in routine pathological diagnosis in hospitals at all levels, thereby bringing long-term benefits to LMS patients.
[0109] See also Figure 6, further, the preprocessing module (2) comprises:
[0110] A region extraction unit (21) is used to perform oncological annotation on the slice to be identified, and mark the region with oncological components in the slice to be identified as a region of interest;
[0111] A gridding unit (22) is connected to the region extraction unit (21) and is used to grid the region of interest and divide it into a plurality of image blocks to be identified;
[0112] The image block optimization unit (23) is connected to the gridding unit (22) and is used to perform Laplace filtering, regularization and signal enhancement processing on the image block to be identified to obtain image data to be identified.
[0113] Laplacian filtering is an edge detection technique that can enhance the edges and details of an image and is used for denoising to make the image clearer. Regularization and signal enhancement can further improve image quality.
[0114] Furthermore, the molecular typing recognition model is constructed based on the DenSeNet121 model, including four dense blocks, and two adjacent dense blocks are directly connected through a transition layer.
[0115] The molecular typing recognition model includes an input layer, a convolution layer, a pooling layer, a first dense block, a first transition layer, a second dense block, a second transition layer, a third dense block, a third transition layer, a fourth dense block, a fourth transition layer, a global average pooling layer, and a classifier in sequence; the input layer is used to input image data. The convolution layer uses a 7×7 convolution kernel to perform a convolution operation on the input image with a step size of 2, and the pooling layer uses a 3×3 maximum pooling layer for downsampling with a step size of 2 to reduce the size of the feature map. The transition layer is located between two adjacent dense blocks and is responsible for reducing the size of the feature map. The global average pooling layer converts the feature map output by the last dense block into a vector of fixed length. The classifier is usually a fully connected layer plus a softmax activation function for the final classification category output.
[0116] Furthermore, the system also includes a second data acquisition module (4), a data analysis module (5), a sample annotation module (6), and a model training module (7);
[0117] The second data acquisition module (4) is used to obtain a number of hematoxylin-eosin-stained full-view digital pathological sections of smooth muscle tumors from a public database as training samples, and to obtain corresponding transcriptome sequencing data from the public database;
[0118] A data analysis module (5), connected to the second data acquisition module (4), is used to perform cluster analysis on the transcriptome test data to determine the molecular typing of the training samples;
[0119] A sample labeling module (6), connected to the data analysis module (5) and the first data acquisition module (4), respectively, for labeling the training samples by molecular typing;
[0120] A preprocessing module (2), connected to the sample labeling module (6), is used to preprocess the labeled training samples to form training data;
[0121] The model training module (7) is used to train the initially constructed neural network model using training data to obtain a molecular typing recognition model.
[0122] The public database is the TCGA database, from which the whole-view digital pathology sections (WSI) and transcriptome sequencing data of leiomyosarcoma are obtained. The preprocessing of the training samples is the same as the preprocessing process of the identification sections, including oncology annotation, gridding, Laplace filtering, regularization, and signal enhancement. The transcriptome test data is analyzed by existing clustering analysis methods to obtain the molecular typing of the training samples to annotate the training samples. Specifically, unsupervised hierarchical clustering (UHC) methods such as average linkage algorithms are used to analyze transcriptome sequencing data (RNA sequencing data). The training data is divided into a training set and a validation set. The training set trains the initially constructed model, and the validation set performs parameter tuning on the training model.
[0123] Furthermore, the system also includes a third data acquisition module (8) and a model testing module (9);
[0124] The third data acquisition module (8) is used to obtain full-field digital pathological sections of smooth muscle tumors stained with hematoxylin-eosin from several smooth muscle tumor patients in the hospital as test samples, and to obtain corresponding immunohistochemical data;
[0125] The sample labeling module (6) is also connected to the third data acquisition module (8) and is used to determine the tumor molecular typing of each smooth muscle tumor patient based on immunohistochemical data analysis, thereby labeling the test sample with molecular typing;
[0126] A preprocessing module (2) is used to preprocess the labeled test samples to form a test set;
[0127] The model verification module (9) is connected to the preprocessing module (2) and is used to test the molecular typing recognition model using a test set.
[0128] The preprocessing of the test samples was the same as that of the slices to be identified, including oncological annotation, gridding, Laplacian filtering, regularization, and signal enhancement.
[0129] The above are only preferred embodiments of the present invention, and are not intended to limit the implementation methods and protection scope of the present invention. Those skilled in the art should be aware that all solutions obtained by equivalent substitutions and obvious changes made using the description and illustrations of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for identifying tumor molecular typing based on neural network, used for identifying smooth muscle tumor molecular typing, characterized in that: include: Step A1, obtaining a full-field digital pathological section of a smooth muscle tumor to be identified stained with hematoxylin-eosin as a section to be identified; Step A2, preprocessing the slice to be identified to obtain image data to be identified formed by the preprocessing; Step A3, inputting the image data to be identified into a pre-trained molecular typing recognition model, and the molecular typing recognition model processes the image data to be identified to obtain a molecular typing recognition result of smooth muscle tumor.
2. A method for identifying tumor molecular typing based on a neural network as claimed in claim 1, characterized in that: In step A2, the process of preprocessing the slice to be identified includes: Step A21, performing oncological annotation on the slice to be identified, marking the region with oncological components in the slice to be identified as a region of interest; Step A22, gridding the region of interest to divide it into a number of image blocks to be identified; Step A23, performing Laplace filtering, regularization and signal enhancement processing on the image block to be identified to obtain the image data to be identified.
3. A method for identifying tumor molecular typing based on a neural network as claimed in claim 1, characterized in that: The molecular typing recognition model is constructed based on the DenSeNet121 model, and includes four dense blocks, and two adjacent dense blocks are directly connected through a transition layer.
4. A method for identifying tumor molecular typing based on a neural network as claimed in claim 1, characterized in that: The training process of the molecular typing recognition model in step A3 includes: Step A31, obtaining a number of full-field digital pathological sections of smooth muscle tumors stained with hematoxylin-eosin from a public database as training samples, and obtaining corresponding transcriptome sequencing data from the public database; Step A32, performing cluster analysis on the transcriptome test data to determine the molecular typing of the training samples, marking the molecular typing of the training samples, and preprocessing the training samples to form training data; Step A33, using the training data to train the initially constructed neural network model to obtain the molecular typing recognition model.
5. A method for identifying tumor molecular typing based on a neural network as claimed in claim 4, characterized in that: In step A33, the molecular typing recognition model is tested using a test set, and the test set formation process includes: Step A331, obtaining full-field digital pathological sections of smooth muscle tumors stained with hematoxylin-eosin from several smooth muscle tumor patients in the hospital as test samples, and obtaining corresponding immunohistochemistry data; Step A332, based on the immunohistochemical data analysis, determine the tumor molecular typing of each smooth muscle tumor patient, thereby molecular typing the verification samples, and pre-process the labeled test samples to obtain the test set.
6. A tumor molecular typing recognition system based on neural network, characterized in that: A method for performing a tumor molecular typing recognition method based on a neural network as claimed in any one of claims 1 to 5, comprising: A first data acquisition module is used to acquire a full-field digital pathological section of a smooth muscle tumor to be identified stained with hematoxylin-eosin as a section to be identified; A preprocessing module, connected to the first data acquisition module, preprocesses the slice to be identified to obtain image data to be identified formed by the preprocessing; The molecular typing recognition module is connected to the preprocessing module and is used to input the image data to be recognized into a pre-trained molecular typing recognition model. The molecular typing recognition model processes the image data to be recognized to obtain a molecular typing recognition result of smooth muscle tumor.
7. A tumor molecular typing recognition system based on neural network as claimed in claim 6, characterized in that: The preprocessing module comprises: A region extraction unit, used for performing oncological annotation on the slice to be identified, and marking a region with oncological components in the slice to be identified as a region of interest; A gridding unit, connected to the region extraction unit, for gridding the region of interest and dividing it into a plurality of image blocks to be identified; The image block optimization unit is connected to the gridding unit and is used to perform Laplace filtering, regularization and signal enhancement processing on the image block to be identified to obtain the image data to be identified.
8. A tumor molecular typing recognition system based on neural network as claimed in claim 6, characterized in that: The molecular typing recognition model is constructed based on the DenSeNet121 model, and includes four dense blocks, and two adjacent dense blocks are directly connected through a transition layer.
9. A tumor molecular typing recognition system based on neural network as claimed in claim 6, characterized in that: The system also includes a second data acquisition module, a data analysis module, a sample annotation module, and a model training module; The second data acquisition module is used to acquire a number of full-field digital pathological sections of smooth muscle tumors stained with hematoxylin-eosin from a public database as training samples, and to acquire corresponding transcriptome sequencing data from the public database; The data analysis module is connected to the second data acquisition module and is used to perform cluster analysis on the transcriptome test data to determine the molecular typing of the training sample; The sample labeling module is connected to the data analysis module and the first data acquisition module respectively, and is used to label the training samples by molecular typing; The preprocessing module is connected to the sample labeling module and is used to preprocess the labeled training samples to form training data; The model training module is used to train the initially constructed neural network model using the training data to obtain the molecular typing recognition model.
10. A tumor molecular typing recognition system based on neural network as claimed in claim 9, characterized in that: The system also includes a third data acquisition module and a model testing module; The third data acquisition module is used to acquire full-field digital pathological sections of smooth muscle tumors stained with hematoxylin-eosin of several smooth muscle tumor patients in the hospital as test samples; The sample labeling module is also connected to the third data acquisition module, and is used to determine the tumor molecular typing of each smooth muscle tumor patient according to the immunohistochemical data analysis, so as to label the corresponding test sample with molecular typing; The preprocessing module is used to preprocess the labeled test samples to form a test set; The model testing module is connected to the preprocessing module and is used to test the molecular typing recognition model using the test set.