Methods and devices for predicting early hepatocellular carcinoma based on multimodal data, and biomarkers for predicting early hepatocellular carcinoma.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2026-08-14
AI Technical Summary
[0004]1)在复杂任务中,单模态模型的心理利用不足,存在模型准确性和鲁棒性较低的问题
Smart Images

Figure CN122575752A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent medical technology, and in particular to a method and device for predicting early hepatocellular carcinoma based on multimodal data, as well as biomarkers for predicting early hepatocellular carcinoma. Background Technology
[0002] Histopathological diagnosis remains the gold standard for cancer diagnosis. Currently, tumor diagnosis is classified into five levels, with reliability increasing sequentially, and histopathological diagnosis being the highest level. While histopathological diagnosis is a routine clinical diagnostic procedure, only a very small number of highly skilled pathologists are capable of accurately diagnosing early-stage hepatocellular carcinoma (EHCC) due to its ambiguous morphological features. Deep learning-based artificial intelligence models can capture subtle features and have been shown to improve diagnostic performance for difficult pathologies. However, diagnosing early-stage liver cancer solely based on pathological image features remains extremely challenging. There is an urgent need to discover new molecular biomarkers and integrate them with pathological image features to achieve more accurate assisted diagnosis. Therefore, a feature selection and multimodal modeling strategy is required.
[0003] Current methods for early cancer diagnosis mainly rely on deep learning approaches based on single-modal pathological images or molecular profiling. However, they still have the following drawbacks:
[0004] 1) In complex tasks, the psychological utilization of single-modal models is insufficient, resulting in problems with low model accuracy and robustness.
[0005] 2) Single-modal models are susceptible to data noise and outliers, which can lead to a decline in model performance.
[0006] 3) Existing multimodal models integrate all molecular expression profiles, which consumes more computational resources and has a higher cost for clinical application.
[0007] Therefore, there is an urgent need in this field to develop a method and device for predicting early hepatocellular carcinoma based on multimodal data, as well as biomarkers for predicting early hepatocellular carcinoma. This method and device combine the screened biomarkers with pathological sections to obtain a more accurate predictive model for the probability of early hepatocellular carcinoma. Summary of the Invention
[0008] The purpose of this application is to provide a method and device for predicting early hepatocellular carcinoma, as well as biomarkers for predicting early hepatocellular carcinoma. The method, system, and device combine the screened biomarkers with pathological sections to obtain a more accurate predictive model for the probability of early hepatocellular carcinoma.
[0009] The first aspect of this application provides a method for constructing a diagnostic model for early hepatocellular carcinoma features based on multimodal data, comprising the following steps:
[0010] (a) Obtain the patient's pathological sections and RNA molecular expression profiles, and digitally scan and save the pathological sections as digital pathological images;
[0011] (b) Extracting features from the digital pathological image to obtain the pathological feature vector of the digital pathological image;
[0012] (c) Based on the RNA molecule expression profile with tags, the RNA molecule expression profile is divided into three categories by machine learning method, the three categories including cirrhosis, HGDN and eHCC, and the model weight of each gene is extracted from the machine fitting model, thereby obtaining the key genes related to the early hepatocellular carcinoma through the model weight;
[0013] (d) Based on the model weights of the key genes, the expression values of the key genes are added one by one as input to train the deep learning neural network model, thereby obtaining multiple different primary multimodal diagnostic models. The output feature of the primary multimodal diagnostic model is the category label of the pathological slice, and the category label refers to cirrhosis, HGDN and eHCC.
[0014] (e) Evaluate the effectiveness of the multiple different primary multimodal diagnostic models based on their performance metrics, and select the model whose performance metrics no longer show significant improvement as the final multimodal diagnostic model.
[0015] In another preferred embodiment, the plurality of different primary multimodal diagnostic models are multiple different primary multimodal diagnostic models corresponding to different inputs, the different inputs including pathological feature vectors and different numbers of expression values of the key genes.
[0016] In another preferred embodiment, in step (c), the tags in the tagged RNA molecule expression profile are manually labeled.
[0017] In another preferred embodiment, in step (b), the pathological feature vector of the digital pathological image is extracted by a deep learning method.
[0018] In another preferred embodiment, in step (d), according to the model weights of the key genes, the inputs of the multiple different primary multimodal diagnostic models are respectively the pathological feature vector, the pathological feature vector and the expression value of the first key gene, the pathological feature vector and the expression values of the first two key genes, the pathological feature vector and the expression values of the first three key genes, ..., and so on, the pathological feature vector and the expression values of the first n key genes.
[0019] In another preferred example, n = 0-50.
[0020] In another preferred example, n = 30.
[0021] In another preferred example, n = 10.
[0022] In another preferred embodiment, in step (c), the top 10 key genes with the highest weights are selected, namely AARS2, ARHGEF11, RABEPK, ATP6V0A2, CKAP5, ADK, CEP250, DCXR, CYP4F2, and AADAT.
[0023] In another preferred embodiment, the input to the final multimodal diagnostic model is the pathological feature vector of the digital pathological image, the expression value of AARS2, the expression value of ARHGEF11, the expression value of RABEPK, and the expression value of ATP6V0A2.
[0024] In another preferred embodiment, in step (d), the number of the plurality of different primary multimodal diagnostic models is 10, namely, the first primary multimodal diagnostic model, the second primary multimodal diagnostic model, the third primary multimodal diagnostic model, the fourth primary multimodal diagnostic model, ..., and so on, up to the tenth primary multimodal diagnostic model.
[0025] In another preferred embodiment, the input of the first primary multimodal diagnostic model is the feature vector of the digital pathology image; the input of the second primary multimodal diagnostic model is the pathological feature vector and the expression value of AARS2; the input of the third primary multimodal diagnostic model is the pathological feature vector, the expression value of AARS2, the expression value of ARHGEF11, and the expression value of RABEPK; the input of the fourth primary multimodal diagnostic model is the pathological feature vector, the expression value of AARS2, the expression value of ARHGEF11, the expression value of RABEPK, and the expression value of ATP6V0A2, ..., and so on, and the input of the tenth primary multimodal diagnostic model is the pathological feature vector, the expression value of AARS2, the expression value of ARHGEF11, the expression value of RABEPK, the expression value of ATP6V0A2, the expression value of CKAP5, the expression value of ADK, the expression value of CEP250, the expression value of DCXR, the expression value of CYP4F2, and the expression value of AADAT.
[0026] In another preferred embodiment, step (b) includes the following sub-steps:
[0027] (b1) The digital pathological image is segmented to generate multiple patch images;
[0028] (b2) Input the labeled patch images into a deep learning network model and train the deep learning network model. The labels are liver cirrhosis label, HGDN label and eHCC label. Extract the probability of the predicted label of all patch images from the deep learning network model and calculate the statistical features of the predicted label and the corresponding probability. Use the statistical features as the pathological feature vector of the digital pathological image.
[0029] In another preferred embodiment, the statistical features include statistical features of the predicted labels of all patch images, namely, the predicted labels of the three categories: non-candidate region labels, HGDN labels, and eHCC labels, and statistical features of the probabilities of the three categories.
[0030] In another preferred embodiment, the statistical features include, but are not limited to: the frequency of each type of label, the proportion of each type of label, and the predicted probability distribution of each type of label.
[0031] In another preferred embodiment, the deep learning network model is a two-stage multi-scale convolutional neural network model.
[0032] In another preferred embodiment, the two-stage multi-scale convolutional neural network model includes a candidate region network model and an early hepatocellular carcinoma detection network model. When training the candidate region network model, a first type of patch image with non-candidate region labels and a second type of patch image with candidate lesion labels are input into the candidate region network model for training. The non-candidate region labels refer to cirrhosis, and the candidate lesion labels refer to having HGDN or eHCC. The candidate region network model is configured to assign a first predicted label to the input patch image and divide it into non-candidate region patch images with the first predicted label being a non-candidate region label and candidate lesion patch images with the first predicted label being a candidate lesion label. The probabilities of the non-candidate region labels and candidate lesion labels are extracted from the candidate region network model.
[0033] In another preferred embodiment, when training the early hepatocellular carcinoma detection network model, third-class patch images with HGDN labels and fourth-class patch images with eHCC labels are input into the early hepatocellular carcinoma detection network model for training. The early hepatocellular carcinoma detection network model is configured to assign a second predicted label to the input patch images and divide them into HGDN patch images with the second predicted label being HGDN and eHCC patch images with the second predicted label being eHCC. The probabilities of having the HGDN label and having the eHCC label are extracted from the early hepatocellular carcinoma detection network model.
[0034] A second aspect of this application provides a method for predicting early hepatocellular carcinoma based on multimodal data, comprising the following steps:
[0035] (a) Obtain the patient's pathological sections and RNA molecular expression profiles, and digitally scan and save the pathological sections as digital pathological images;
[0036] (b) The digital pathological image is used to extract features to obtain a pathological feature vector;
[0037] (c) Input the pathological feature vector and the expression values of AARS2, ARHGEF11, RABEPK and ATP6V0A2 in the RNA molecular expression profile into the final multimodal diagnostic model trained above, thereby obtaining the category label of the pathological slice.
[0038] A third aspect of this application provides an early hepatocellular carcinoma prediction device based on multimodal data, comprising:
[0039] Memory, used to store computer-executable instructions; and,
[0040] A processor, coupled to the memory, is configured to implement the steps of the method described above when executing the computer-executable instructions.
[0041] A fourth aspect of this application provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the method described above.
[0042] The fifth aspect of this application provides a computer program product including computer-executable instructions, characterized in that the computer-executable instructions, when executed by a processor, implement the steps in the above-described method.
[0043] The sixth aspect of this application provides the use of a biomarker and / or its detection reagent for preparing a reagent or kit for detecting liver cancer, said biomarker being selected from the group consisting of: AARS2, ARHGEF11, RABEPK, ATP6V0A2, CKAP5, ADK, CEP250, DCXR, CYP4F2, AADAT, or combinations thereof.
[0044] In another preferred embodiment, the biomarker is selected from the group consisting of AARS2, ARHGEF11, RABEPK, ATP6V0A2, or combinations thereof.
[0045] In another preferred embodiment, the kit is used for the early detection of liver cancer.
[0046] It should be understood that, within the scope of this invention, the above-described technical features of this invention and the technical features specifically described below (such as in the embodiments) can be combined with each other to form new or preferred technical solutions. Due to space limitations, they will not be described in detail here. Attached Figure Description
[0047] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. It should be understood that the accompanying drawings described below are merely some implementation examples of the present invention, and those skilled in the art can obtain other implementation examples based on these drawings without creative effort.
[0048] Figure 1 This is a flowchart of a method for constructing a diagnostic model for early hepatocellular carcinoma features based on multimodal data according to this application;
[0049] Figure 2 This is the difference in functionality compared to H-DN and eHCC;
[0050] Figure 3 This is the difference in functionality between eHCC and H-DN;
[0051] Figure 4 This refers to the performance of Xgboost using five-fold cross-validation on the discovery queue;
[0052] Figure 5 This refers to Xgboost's performance on independent external test sets;
[0053] Figure 6 This is a heatmap of the expression of 10 genes in all samples (the left side is the discovery cohort, and the right side is the independent external testing cohort).
[0054] Figure 7 The 10 genes with the highest weights in the Xgboost model and their corresponding biological functions;
[0055] Figure 8 This is a schematic diagram of the architecture of the two-stage multi-scale convolutional neural network TMC-net model;
[0056] Figure 9 This is a flowchart of the calculation process for 34-dimensional pathological features;
[0057] Figure 10 This relates to the impact of 0-10 gene features on the accuracy of diagnostic models;
[0058] Figure 11 This relates to the impact of 0-10 gene characteristics on the AUROC of the diagnostic model. Detailed Implementation
[0059] Through extensive and in-depth research, the inventors have developed for the first time a method, system, and device for predicting early-stage hepatocellular carcinoma. The method includes: 1) using a machine learning model to perform gene screening based on feature importance assessment; 2) using deep learning to extract features based on pathological images; and 3) using a forward feature selection method to sequentially add molecular features to the pathological image features to obtain the optimal combination and establish the final machine learning diagnostic model. The method and device of this application aim to learn important molecular features as potential molecular biomarkers through feature selection. Based on this, by integrating pathological image features, a multimodal machine learning model is established to more accurately predict early-stage hepatocellular carcinoma and more efficiently discover novel biomarkers.
[0060] In the following description, many technical details are presented to help the reader better understand this application. However, those skilled in the art will understand that the technical solutions claimed in this application can be implemented even without these technical details and various variations and modifications based on the following embodiments.
[0061] the term
[0062] As used in this article, "eHCC" refers to early-stage liver cancer; all instances of "HCC" in this article refer to "eHCC".
[0063] As used in this article, "HGDN" refers to "high-grade dysplastic nodule of the liver". "DN" and "H-DN" in this article both refer to "HGDN".
[0064] It should be noted that in this patent application, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. In this patent application, if it refers to performing an action according to an element, it means performing the action at least according to that element, including two cases: performing the action only according to that element, and performing the action according to that element and other elements. Expressions such as "multiple," "repeatedly," and "various" include two, two times, two kinds, and more than two, more than two times, and more than two kinds.
[0065] In this invention, all directional indicators (such as up, down, left, right, front, back, etc.) are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indicator will also change accordingly.
[0066] This application possesses at least one of the following advantages.
[0067] (a) This invention provides a multimodal feature selection and integration method, which provides a complete process from input to output and can be applied to experiments for the discovery of new biomarkers as a more efficient computational tool;
[0068] (b) This system adopts a strategy of integrating pathological images and molecular features, providing a more comprehensive feature representation of early liver cancer lesions and establishing a diagnostic model with higher accuracy and more flexible data.
[0069] (c) The forward feature selection method is used to screen and integrate multimodal features, which reduces manual intervention and improves analysis efficiency.
[0070] A method for constructing a multimodal diagnostic model for early hepatocellular carcinoma features
[0071] See Figure 1 This application provides a method for constructing a multimodal diagnostic model for early hepatocellular carcinoma features, comprising the following steps:
[0072] (a) Obtain the patient's pathological sections and RNA molecular expression profiles, and digitally scan and save the pathological sections as digital pathological images;
[0073] (b) Extracting features from the digital pathological image to obtain a pathological feature vector;
[0074] (c) Based on the RNA molecule expression profile with tags, the RNA molecule expression profile is divided into three categories by machine learning method, the three categories include cirrhosis, HGDN and eHCC, the model weight of each gene is extracted from the machine fitting model, and the key genes related to the early hepatocellular carcinoma are obtained through the model weight.
[0075] (d) Based on the model weights of the key genes, the expression values of the key genes are added one by one as input to train the deep learning neural network model, thereby obtaining multiple different primary multimodal diagnostic models. The output feature of the primary multimodal diagnostic model is the category label of the pathological slice, and the category label refers to cirrhosis, HGDN and eHCC.
[0076] (e) The effectiveness of the various primary multimodal diagnostic models is evaluated based on their performance metrics, and the model whose performance metrics no longer show significant improvement is selected as the final multimodal diagnostic model. The final multimodal diagnostic model is the multimodal early hepatocellular carcinoma characteristic diagnostic model.
[0077] Preferably, in step (c), the tags of the tagged RNA molecule expression profile are manually labeled.
[0078] Preferably, the plurality of different primary multimodal diagnostic models are multiple different primary multimodal diagnostic models corresponding to different inputs, wherein the different inputs include pathological feature vectors and expression values of different numbers of the key genes.
[0079] Preferably, in step (d), according to the model weights of the key genes, the inputs of the multiple different primary multimodal diagnostic models are respectively the pathological feature vector, the pathological feature vector and the expression value of the first key gene, the pathological feature vector and the expression values of the first two key genes, the pathological feature vector and the expression values of the first three key genes, ..., and so on, the pathological feature vector and the expression values of the first n key genes.
[0080] Preferably, n = 0-50. More preferably, n = 30. More preferably, n = 10.
[0081] Preferably, in step (c), the top 10 key genes with the highest weights are selected, namely AARS2, ARHGEF11, RABEPK, ATP6V0A2, CKAP5, ADK, CEP250, DCXR, CYP4F2, and AADAT.
[0082] Preferably, the input to the final multimodal diagnostic model is the pathological feature vector of the digital pathological image, the expression value of AARS2, the expression value of ARHGEF11, the expression value of RABEPK, and the expression value of ATP6V0A2.
[0083] In another preferred embodiment, in step (d), the number of the plurality of different primary multimodal diagnostic models is 10, namely, the first primary multimodal diagnostic model, the second primary multimodal diagnostic model, the third primary multimodal diagnostic model, the fourth primary multimodal diagnostic model, ..., and so on, up to the tenth primary multimodal diagnostic model.
[0084] In another preferred embodiment, the input of the first primary multimodal diagnostic model is the feature vector of the digital pathology image; the input of the second primary multimodal diagnostic model is the pathological feature vector and the expression value of AARS2; the input of the third primary multimodal diagnostic model is the pathological feature vector, the expression value of AARS2, the expression value of ARHGEF11, and the expression value of RABEPK; the input of the fourth primary multimodal diagnostic model is the pathological feature vector, the expression value of AARS2, the expression value of ARHGEF11, the expression value of RABEPK, and the expression value of ATP6V0A2, ..., and so on, and the input of the tenth primary multimodal diagnostic model is the pathological feature vector, the expression value of AARS2, the expression value of ARHGEF11, the expression value of RABEPK, the expression value of ATP6V0A2, the expression value of CKAP5, the expression value of ADK, the expression value of CEP250, the expression value of DCXR, the expression value of CYP4F2, and the expression value of AADAT.
[0085] A two-stage multi-scale convolutional neural network model (TMC-net) is used to extract the pathological feature vectors.
[0086] In one embodiment of this application, in step (b), the pathological feature vector of the digital pathological image is extracted using a deep learning method, for example, see [link to relevant documentation]. Figure 8 and Figure 9 .
[0087] Preferably, step (b) includes the following sub-steps:
[0088] (b1) The digital pathological image is segmented to generate multiple patch images;
[0089] (b2) Input the labeled patch images into a deep learning network model and train the deep learning network model. The labels are liver cirrhosis label, HGDN label and eHCC label. Extract the probability of the predicted label of all patch images from the deep learning network model and calculate the statistical features of the predicted label and the corresponding probability. Use the statistical features as the pathological feature vector of the digital pathological image.
[0090] Specifically, in a preferred embodiment, in step (b2), a two-stage multi-scale convolutional neural network model (TMC-net) can be used to extract the pathological feature vector. The training steps of the two-stage multi-scale convolutional neural network model include:
[0091] (b21) Input a first-class patch image with a non-candidate region label and a second-class patch image with a candidate lesion label into a candidate region network model, and train the candidate region network model. The non-candidate region label refers to cirrhosis, and the candidate lesion label refers to having HGDN or eHCC. The candidate region network model is configured to assign a first predicted label to the input patch image and classify it into a non-candidate region patch image with the first predicted label being a non-candidate region label and a candidate lesion patch image with the first predicted label being a candidate lesion label. Extract the probabilities of the two labels, i.e., the categories (non-candidate region label and candidate lesion label), from the candidate region network model. The terms "non-candidate region" and "non-interest region" are used interchangeably.
[0092] (b22) Input the third type patch image with HGDN label and the fourth type patch image with eHCC label into the early hepatocellular carcinoma detection network model, train the early hepatocellular carcinoma detection network model, the early hepatocellular carcinoma detection network model is configured to assign a second predicted label to the input patch image and divide it into DN patch image with the second predicted label having HGDN label and eHCC patch image with the second predicted label having eHCC label, and extract the probabilities of the two labels, i.e. the categories (HGDN label and eHCC label), from the early hepatocellular carcinoma detection network model;
[0093] (b23) Based on the obtained probabilities of cirrhosis areas (prediction of non-candidate areas), HGDN probability, and eHCC probability, calculate statistical features (preferably including 34 indicators), including the number of predictions for each category, the proportion of predictions for each category, and the mean of the probability distribution curves for each category, as the pathological feature vector of the digital pathological image.
[0094] Preferably, the statistical features include, but are not limited to: the frequency of each type of label, the proportion of each type of label, and the predicted probability distribution of each type of label.
[0095] Preferably, the statistical features are selected from the following group: the number of patches predicted as non-candidate regions (also known as non-interest regions), the number of patches predicted as HGDN, the number of patches predicted as eHCC, the total number of patches in the pathological image, the proportion of patches predicted as non-candidate regions to the total number of patches, the proportion of patches predicted as HGDN to the total number of patches, the proportion of patches predicted as eHCC to the total number of patches, the mean of non-candidate region category probability, the standard deviation of non-candidate region category probability, the 25th percentile of non-candidate region category probability, the median (50th percentile) of non-candidate region category probability, the 75th percentile of non-candidate region category probability, the number of patches with a non-candidate region category probability > 0.1, the number of patches with a non-candidate region category probability > 0.2, the number of patches with a non-candidate region category probability > 0.5, the number of patches with a non-candidate region category probability > 0.7, H... Mean probability of GDN class, standard deviation of HGDN class probability, 25th percentile of HGDN class probability, median (50th percentile) of HGDN class probability, 75th percentile of HGDN class probability, number of patches with HGDN class probability > 0.1, number of patches with HGDN class probability > 0.2, number of patches with HGDN class probability > 0.5, number of patches with HGDN class probability > 0.7, mean probability of eHCC class, standard deviation of eHCC class probability, 25th percentile of eHCC class probability, median (50th percentile) of eHCC class probability, 75th percentile of eHCC class probability, number of patches with eHCC class probability > 0.1, number of patches with eHCC class probability > 0.2, number of patches with eEHCC class probability > 0.5, number of patches with eHCC class probability > 0.7, or a combination thereof.
[0096] Preferably, both the candidate region network model and the early hepatocellular carcinoma detection network model are multi-scale convolutional neural networks, including low-resolution convolutional neural networks and high-resolution convolutional neural networks.
[0097] Preferably, in step (b21), each patch image of the first and second types, magnified 20 times, is input into the first low-resolution convolutional neural network of the candidate region network model to extract global features of the input patch images; and
[0098] Each patch image of the first and second classes, magnified 40 times, is input into the first high-resolution convolutional neural network of the candidate region network model to extract detailed features of the input patch image; then the outputs of the first low-resolution convolutional neural network and the first high-resolution convolutional neural network are concatenated together and input into the first multilayer perceptron to obtain the first predicted label for each input patch image.
[0099] Preferably, in step (b22), each of the third and fourth type patch images magnified 20 times is input into the second low-resolution convolutional neural network of the early hepatocellular carcinoma detection network model to extract the global features of the input patch images.
[0100] Each of the third and fourth type patch images, magnified 40 times, is input into the second high-resolution convolutional neural network of the candidate region network model to extract detailed features of the input patch images; then the outputs of the second low-resolution convolutional neural network and the second high-resolution convolutional neural network are concatenated together and input into the second multilayer perceptron to obtain the second predicted label for each input patch image.
[0101] In other embodiments, step (b) further includes calculating statistical features (preferably 34-dimensional features as described above) of the predicted labels and corresponding probabilities of all patch images of the digital pathology image, and then inputting the statistical features into a classifier for training. The classifier is used to predict the final predicted label of the digital pathology image. The final predicted label refers to the DN label and the eHCC label. Preferably, the final predicted label can also be called the slice-level predicted label (slide prediction). The classifier is a support vector machine (SVM) classifier.
[0102] A method for predicting early hepatocellular carcinoma based on multimodal data
[0103] This application also provides a method for predicting early hepatocellular carcinoma based on multimodal data, including the following steps:
[0104] (a) Obtain the patient's pathological sections and RNA molecular expression profiles, and digitally scan and save the pathological sections as digital pathological images;
[0105] (b) The digital pathological image is used to extract features to obtain a pathological feature vector;
[0106] (c) Input the pathological feature vector and the expression values of AARS2, ARHGEF11, RABEPK and ATP6V0A2 in the RNA molecular expression profile into the trained final multimodal diagnostic model to obtain the category label of the pathological slice.
[0107] A multimodal diagnostic system for early hepatocellular carcinoma
[0108] This system comprises the following components: data preprocessing module, feature transformation module, feature selection module, feature integration and model training module, and malignancy risk value output module.
[0109] Data preprocessing module: Collects samples diagnosed as early-stage hepatocellular carcinoma or atypical dysplastic nodules. Pathological images are digitized and segmented into small patch images according to the requirements of characteristic cancer types. RNA expression data undergoes routine quality control and normalization.
[0110] Feature transformation module: Small patch images undergo feature learning through two binary classification deep learning models (the candidate region network model and the early hepatocellular carcinoma detection network model as described above). The sum, maximum, median, or mean of these features are then used as the feature vector for the pathological image. The level of RNA molecule expression is directly used as a molecular feature.
[0111] Feature selection module: For pathological image features, dimensionality reduction is performed on the network output pathological image features by calculating statistical indicators. For RNA expression profiles, after training the machine learning model, the importance score of each feature is extracted from the fitted model as the importance score of each gene. Genes are then sorted in descending order according to their importance scores.
[0112] Feature integration and model training module: Machine learning or deep learning algorithms are used to train the extracted features. During training, genes are added one by one to the pathological features based on their importance scores. This process generates multiple models with different features. All models are evaluated using the same performance metrics (such as AUROC, accuracy, etc.), and the model that no longer shows significant improvement is selected as the final multimodal diagnostic model.
[0113] Malignancy Risk Output Module: After obtaining the fitted model, the model is used for inference. During the inference phase, this module outputs the model's predicted probability of malignancy for the sample being early-stage hepatocellular carcinoma.
[0114] During operation, this system requires pathological images and molecular expression profiles of the samples.
[0115] After preprocessing and segmentation, pathological images are generated into multiple small patch images, which are then trained using supervised or unsupervised deep learning methods. The predicted probabilities of these small patch images are then extracted from the model. Thirty-four indicators, including the number of predicted classes, the proportion of each class, and the mean of the probability distribution curves for each class, are calculated for all small patch images and used as pathological image features based on deep learning.
[0116] Molecular features (such as transcriptomes) are input into a machine learning model for training. Importance scores for each feature are extracted from the fitted model and used as the importance score for each gene. The top ten genes with the highest scores are considered key genes.
[0117] In the multimodal integration process, genes were added one by one according to their importance scores. This process generated 11 models based on different features: pathological model, pathological model and the first gene, pathological model and the first two genes, up to the pathological model and all the first 10 genes. All models followed the same machine learning training and evaluation framework, and the model that no longer showed significant improvement was selected as the final multimodal diagnostic model.
[0118] To make the objectives, technical solutions, and advantages of the present invention clearer, embodiments of the present invention will be described in further detail below with reference to the accompanying drawings. It should be understood that these are merely examples provided to the reader of possible implementations of the present invention and are not intended to limit the scope of the invention.
[0119] Example
[0120] In this embodiment, the process of constructing a diagnostic model for early hepatocellular carcinoma based on multimodal data is given as follows.
[0121] 1. Data Source
[0122] The data for both the model extracting digital image feature vectors and the machine learning model screening genes came from the same patient source. The training and validation sets were from Zhongshan Hospital affiliated with Fudan University. The external independent test set was from Huashan Hospital affiliated with Fudan University.
[0123] 2. H&E Image Preprocessing
[0124] H&E slides were scanned and saved as digital pathological images (hereinafter referred to as HE images) using an Aperiod AT2 digital slide scanner 20X (500nm / pixel resolution) or a NanoZoomer S360 digital slide scanner 40X (220nm / pixel resolution). Each HE image was segmented into non-overlapping patch tiles of 512x512µm.
[0125] 3. Data Processing of RNA Data
[0126] Considering the potential for batches among samples from different sequencing times, we used the `combat_seq` function in the `svaR` package (v.3.38) to adjust the expression data using negative binomial regression. Three sample groups (cirrhosis, H-DN, and eHCC) were set as "variables to be maintained". We used PCA to confirm the effectiveness of batch removal and used the adjusted expression profiles for subsequent modeling.
[0127] 4. Differential expression and functional analysis of RNA data
[0128] Differential expression analysis was performed on the expression profiling data using the DESeq2R package (v.1.30). Using four healthy liver samples as controls, we obtained differentially expressed genes (DEGs) for cirrhosis, H-DN, and eHCC. Significant DEGs were defined as genes with a corrected p-value <0.05 (Wald test, false discovery rate method). In functional enrichment analysis, we set the design matrix to "one-to-other" comparisons to detect specific significant DEGs. The Hallmark gene set and KEGG set were downloaded from the public MsigDB dataset and enriched using the ClsuterProfilerR package.
[0129] 5. Framework for building machine learning models (machine learning based on RNA molecular expression profiling)
[0130] RNA profiling-based machine learning was used to classify cirrhosis, H-DN, and eHCC. The machine learning model construction flowchart consists of three parts: feature selection, model hyperparameter tuning, model fitting, and model evaluation. The machine learning pipeline was built in Python (v.3.9.15) using the following libraries: scikit-learn (v.1.2.0), numpy (v.1.24.1), scipy (v.1.9.3), and pandas (v.1.5.2).
[0131] All samples were divided into a training queue (80%) and a test queue (20%) to complete the classification task.
[0132] 1) In feature selection, we input the transcriptional profiles from the training set above into the model for training, and extract the importance score of each feature from the fitted model as the importance score of each gene. This process is repeated 10 times, and the top 10 genes with the highest average importance scores (G1, G2, ..., G...) are selected. 10 These 10 key features are considered crucial. During downstream training, each model is trained based on these 10 key features.
[0133] 2) Model Hyperparameter Tuning: XGBboost is an advanced machine learning model that can extract more interpretable and important features, and it was used to model RNA expression values. To balance computational burden and model accuracy, we used a random search cross-validation strategy to search for the optimal hyperparameters, setting roc_auc as the scoring function for 5-fold cross-validation.
[0134] 3) Model Fitting and Model Evaluation: To evaluate model performance, we randomly re-split the queue and refitted the model 10 times, evaluating AUROC, balanced accuracy, F1, precision, and recall on a test queue that the model had never seen before.
[0135] 6. Machine learning models integrated with pathological features (multimodal diagnostic models)
[0136] In this embodiment, a two-stage multi-scale convolutional neural network model (TMC-net) is used to extract the pathological feature vector of the digital pathological image. The 34-dimensional feature vector extracted from the two-stage multi-scale convolutional neural network model is used as the pathological feature vector of the digital pathological image. Then, according to the model weights of the key genes, the expression values of the key genes are sequentially added as input to the pathological feature vector to train the deep learning neural network model, thereby obtaining multiple different primary multimodal diagnostic models. The output feature of the primary multimodal diagnostic model is the category label of the pathological slide, which refers to cirrhosis, HGDN, and eHCC. The effectiveness of the multiple different primary multimodal diagnostic models is evaluated according to their performance indicators, and the model whose performance indicators no longer show significant improvement is selected as the final multimodal diagnostic model. Two-stage multi-scale convolutional neural network model (TMC-net)
[0137] Figure 8 A schematic diagram of a two-stage multi-scale convolutional neural network (TMC-net) architecture is presented. During diagnosis, doctors need not only global observation but also detailed analysis of local regions. TMC-net simulates the doctor's diagnostic process.
[0138] Low-magnification observation helps to understand changes in the overall organizational structure;
[0139] High-magnification observation is used to focus on the microscopic features of key lesions.
[0140] Therefore, this application uses two images with different magnifications as input to the network:
[0141] 20x magnified image (overall features): Each patch image magnified 20x is labeled as X0 and used as input to a low-resolution convolutional neural network (Low-Resolution CNN) to extract global features.
[0142] 40x magnified image (local features): Each patch image magnified 40x is divided into four sub-regions—upper left, upper right, lower left, and lower right—labeled as X1, X2, X3, and X4, respectively, and input into a high-resolution convolutional neural network (High-Resolution CNN) to extract detailed features.
[0143] These feature representations at different resolutions are concatenated to provide rich information for further classification tasks. Mathematically, the representation process is as follows:
[0144] Z0 = CNN low-resolution (X0),
[0145] Z l =CNN high-resolution (X l ), l=1,2,3,4
[0146]
[0147] The [·] symbol indicates that multiple sets of features are concatenated into a single vector. The Multilayer Perceptron (MLP) uses these concatenated features to generate the final prediction output and assigns a prediction label to each patch image.
[0148] Two-stage lesion classification
[0149] Although DN is not malignant, H-DN and eHCC share similar histological features, which differ significantly from normal liver histology. In the candidate region extraction stage, the first-stage candidate region network filters out potential lesion areas, i.e., candidate patches that may be DN or eHCC. These patches are then fed into the second-stage eHCC detection network, which is also based on a multi-scale convolutional neural network architecture. This network provides an accurate diagnosis of whether a patch is DN or eHCC. The entire framework is trained end-to-end, allowing for joint optimization of all layers.
[0150] Each patch image, after passing through TMC-net, will be assigned the following predicted labels:
[0151] (1) Non-candidate region
[0152] (2)DN
[0153] (3)eHCC
[0154] In addition, TMC-net will also generate a corresponding probability for each label.
[0155] After patch-level prediction is completed, all patch prediction results for the entire digital pathology slide (WholeSlideImage, WSI) are summarized and analyzed. Statistical characteristics of the patch prediction labels and probabilities are calculated (see details in [link to documentation]). Figure 9 (As per Table 1), each WSI is represented as a vector containing 34-dimensional features, which serves as the pathological feature vector of the digital pathological image described in this application:
[0156]
[0157] These features include:
[0158] a) Frequency of each tag type (e.g., number of patches for DN and eHCC)
[0159] b) The proportion of each type of label (e.g., the patch proportion of DN and eHCC).
[0160] c) Predicted probability distribution of various labels (e.g., mean, standard deviation, quantiles) Table 1 Statistical characteristics of patch images
[0161]
[0162]
[0163] In integrating RNA expression profiles, we sequentially added gene expression values to the HE image features (i.e., the pathological feature vector of the digital pathology image, the aforementioned 34-dimensional feature vector). Specifically, based on gene importance scores, we added one gene at a time in descending order. This process generated 10 models based on different features: the HE feature model and the expression level of the first gene (X = {Z}). HE ,G1}), HE characteristics and expression levels of the first two genes (X={Z HE ,G1,G2}), ..., HE characteristics and expression levels of all 10 genes (X={Z HE ,G1,G2,…,G 10 All models follow the same machine learning training and evaluation framework mentioned above.
[0164] result
[0165] Expression profiling reveals molecular alterations in different liver diseases
[0166] HE images alone are insufficient for accurately identifying complex diseases, especially high-grade dysplastic nodules (H-DN) and early-stage hepatocellular carcinoma (eHCC), two lesions with extremely similar pathological features. This prompts us to further explore the differences in molecular patterns between H-DN and eHCC.
[0167] We performed RNA sequencing on 86 lesions (H-DN, eHCC, and cirrhosis) from 45 patients and 4 normal liver samples. Functional enrichment analysis showed that, compared with other groups, the top 10 significantly upregulated functions in the H-DN samples were amino acid and acid metabolism, as well as oxidative phosphorylation. Figure 2 eHCC samples showed significant changes in amino acid balance, mTORC1 signaling pathway activation, cell cycle regulation, enhanced cell proliferation, and epithelial-mesenchymal transition. Figure 3 ).
[0168] Based on the changes in transcriptional profiles of different liver diseases, we constructed a machine learning framework to classify samples of cirrhosis, H-DN, and eHCC. Samples collected from Zhongshan Hospital affiliated with Fudan University constituted the discovery cohort (sample size = 32), and an XGBoost model was trained on them. We then split the dataset into 10 partitions and used a five-fold cross-validation scheme to optimize the model's hyperparameters and tested its performance. Samples collected from Huashan Hospital affiliated with Fudan University formed the external test cohort (sample size = 18). We confirmed that XGBoost performs well in five-fold cross-validation... Figure 4 ) and external test set ( Figure 5 The F1 scores on the XGBoost model reached 0.9212±0.0225 and 0.8, respectively. Based on the expression values of the top 10 genes with the highest weights in the XGBoost model, significant differences in expression of these 10 genes (AARS2, ARHGEF11, RABEPK, ATP6V0A2, CKAP5, ADK, CEP250, DCXR, CYP4F2, and AADAT) were observed in both the ZHFU cohort and the external HHFU cohort. Figure 6 Consistent with previous enrichment analyses, they are involved in vascular smooth muscle contraction or multiple metabolic pathways. This suggests that machine learning models can establish disease-related features and patterns, leading to more accurate diagnoses. Figure 7 ).
[0169] Machine learning integration with multimodal capabilities
[0170] For complex diseases requiring diagnosis, gene or protein expression is often used to assist pathologists in diagnosing HE slides, thereby improving diagnostic accuracy. Therefore, we further integrated transcriptional features and pathological images to build a predictive model. A total of 48 samples' RNA sequencing and pathological images were analyzed. In the discovery cohort, the training and validation sets were randomly split 10 times in an 8:2 ratio. The 34-dimensional embedding features of each HE image were extracted from the class probabilities output from a two-stage multi-scale convolutional neural network model. Figure 9 Five-fold cross-validation results showed that the model integrating the expression values of 10 genes from pathological image features achieved an accuracy of 0.8231 ± 0.0308. On the external test set, the accuracy of the multimodal integration model reached 0.8462. Figure 10Subsequently, based on the HE image features, we sequentially added the expression values of the top K important genes. The results showed that as the number of genes increased (K from 1 to 10), the model's AUROC and accuracy increased on both the training and test sets, indicating that molecules can effectively aid in the diagnosis of lesions in H&E. Specifically, compared to having no genes, adding the top four genes—AARS2, ARHGEF11, RABEPK, and ATP6A0V2—improved the model's AUROC on the validation set from 0.7188±0.0376 to 0.8875±0.0459, and on the external test set from 0.7000 to 0.9500. Figure 10 Subsequently, as the number of genes increased, the model's accuracy and AUROC on the external test set no longer improved significantly. Figure 11 In summary, we obtained an excellent and robust XG Boost ensemble model (the final multimodal diagnostic model), which in particular achieved accurate classification of H-DN and eHCC.
[0171] Accordingly, embodiments of this application also provide a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the various method embodiments of this application. Computer-readable storage media include permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device. As defined herein, computer-readable storage media does not include transient media, such as modulated data signals and carrier waves.
[0172] Furthermore, embodiments of this application also provide an apparatus for predicting early hepatocellular carcinoma based on multimodal data, including a memory for storing computer-executable instructions and a processor; the processor is used to implement the steps in the above-described method embodiments when executing the computer-executable instructions in the memory. The processor may be a Central Processing Unit (CPU), Graphics Processing Unit (GPU), Digital Signal Processor (DSP), Microcontroller Unit (MCU), Neural Processing Unit (NPU), Application Specific Integrated Circuit (ASIC), Field Programmable Gate Array (FPGA), or other programmable logic devices. The aforementioned memory may be read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or solid-state drive, etc. The steps of the methods disclosed in the embodiments of this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules in the processor.
[0173] Furthermore, embodiments of this application also provide a computer program product, including computer-executable instructions that, when executed by a processor, implement the steps in the above-described method embodiments.
[0174] The specification of this application contains numerous technical features distributed across various technical solutions. Listing all possible combinations of these technical features (i.e., technical solutions) would make the specification excessively lengthy. To avoid this problem, the various technical features disclosed in the above-described invention, the various technical features disclosed in the following embodiments and examples, and the various technical features disclosed in the accompanying drawings can be freely combined to form various new technical solutions (all of which are considered to have been described in this specification), unless such a combination of technical features is technically infeasible. For example, one example discloses feature A+B+C, and another example discloses feature A+B+D+E. Features C and D are equivalent technical means that serve the same function, and technically only one needs to be used; they cannot be used simultaneously. Feature E can technically be combined with feature C. Therefore, the solution A+B+C+D should not be considered as described because it is technically infeasible, while the solution A+B+C+E should be considered as described.
[0175] All documents mentioned in this application are considered to be incorporated in their entirety into the disclosure of this application so that they can serve as a basis for modifications if necessary. Furthermore, it should be understood that after reading the foregoing disclosure of this application, those skilled in the art can make various alterations or modifications to this application, and these equivalent forms also fall within the scope of protection claimed in this application.
Claims
1. A method for constructing a diagnostic model for early hepatocellular carcinoma features based on multimodal data, characterized in that, Includes the following steps: (a) Obtain the patient's pathological sections and RNA molecular expression profiles, and digitally scan and save the pathological sections as digital pathological images; (b) Extracting features from the digital pathological image to obtain the pathological feature vector of the digital pathological image; (c) Based on the RNA molecule expression profile with tags, the RNA molecule expression profile is divided into three categories by machine learning method, the three categories including cirrhosis, HGDN and eHCC, and the model weight of each gene is extracted from the machine fitting model, thereby obtaining the key genes related to the early hepatocellular carcinoma through the model weight; (d) Based on the model weights of the key genes, the expression values of the key genes are added one by one as input to train the deep learning neural network model, thereby obtaining multiple different primary multimodal diagnostic models. The output feature of the primary multimodal diagnostic model is the category label of the pathological slice, and the category label refers to cirrhosis, HGDN and eHCC. (e) Evaluate the effectiveness of the multiple different primary multimodal diagnostic models based on their performance metrics, and select the model whose performance metrics no longer show significant improvement as the final multimodal diagnostic model.
2. The method according to claim 1, characterized in that, In step (d), according to the model weights of the key genes, the inputs of the multiple different primary multimodal diagnostic models are respectively the pathological feature vector, the pathological feature vector and the expression value of the first key gene, the pathological feature vector and the expression values of the first two key genes, the pathological feature vector and the expression values of the first three key genes, ..., and so on, the pathological feature vector and the expression values of the first n key genes.
3. The method according to claim 1, characterized in that, In step (c), the top 10 key genes with the highest weights are selected, namely AARS2, ARHGEF11, RABEPK, ATP6V0A2, CKAP5, ADK, CEP250, DCXR, CYP4F2, and AADAT.
4. The method according to claim 3, characterized in that, The input to the final multimodal diagnostic model is the pathological feature vector of the digital pathological image, the expression value of AARS2, the expression value of ARHGEF11, the expression value of RABEPK, and the expression value of ATP6V0A2.
5. The method according to claim 1, characterized in that, Step (b) includes the following sub-steps: (b1) The digital pathological image is segmented to generate multiple patch images; (b2) Input the labeled patch images into a deep learning network model and train the deep learning network model. The labels are liver cirrhosis label, HGDN label and eHCC label. Extract the probability of the predicted label of all patch images from the deep learning network model and calculate the statistical features of the predicted label and the corresponding probability. Use the statistical features as the pathological feature vector of the digital pathological image.
6. A method for predicting early hepatocellular carcinoma based on multimodal data, characterized in that, Includes the following steps: (a) Obtain the patient's pathological sections and RNA molecular expression profiles, and digitally scan and save the pathological sections as digital pathological images; (b) The digital pathological image is used to extract features to obtain a pathological feature vector; (c) Input the pathological feature vector and the expression values of AARS2, ARHGEF11, RABEPK and ATP6V0A2 in the RNA molecular expression profile into the final multimodal diagnostic model trained by any one of claims 1-5, thereby obtaining the category label of the pathological section.
7. A device for predicting early hepatocellular carcinoma based on multimodal data, characterized in that, include: Memory is used to store executable instructions for a computer; as well as, A processor, coupled to the memory, is configured to implement the steps of the method as described in claim 6 when executing the computer-executable instructions.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions that, when executed by a processor, implement the steps of the method as described in claim 6.
9. A computer program product comprising computer-executable instructions, characterized in that, When the computer-executable instructions are executed by a processor, they implement the steps of the method of claim 6.
10. The use of a biomarker, and / or its detection reagent, characterized in that, For the preparation of reagents or kits for the detection of liver cancer, the biomarkers are selected from the group consisting of: AARS2, ARHGEF11, RABEPK, ATP6V0A2, CKAP5, ADK, CEP250, DCXR, CYP4F2, AADAT, or combinations thereof.