Biomarker model training method and system, electronic equipment and medium
Through targeted spatial multi-omics scanning instruments and algorithm model training, the low-cost, high-throughput and robustness issues of the existing biological analysis framework have been solved, and efficient biomarker model training has been achieved to meet the needs of precision medicine and clinical diagnosis.
Patent Information
- Application Number
- CN202510665333.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-09-16
AI Technical Summary
The existing technology lacks a low-cost, high-throughput and robust bioanalysis framework, resulting in insufficient sample size, high detection costs, high data complexity, and poor model generalization, making it difficult to achieve precision medicine and clinical diagnosis.
Use targeted spatial multi-omics scanning instruments to perform a one-time scan of tissue chips to obtain gene expression information of specific cells. Through algorithm model training and evaluation indicator screening, a high-throughput and high-spatial-resolution biomarker model is obtained.
It meets the sample requirements of high throughput and high spatial resolution, obtains a variety of accurate biomarker models, reduces the cost of single sample testing, and improves detection efficiency and model generalization.
Smart Images

Figure CN120656554A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of biomedical information technology and relates to a biomarker model training method, and in particular to a biomarker model training method, system, electronic equipment and medium. Background Art
[0002] In the field of biomedical information technology, existing data acquisition and analysis technologies still face significant bottlenecks, which restrict the development of precision medicine and clinical diagnosis. First, insufficient sample size is one of the core challenges. Traditional biomarker research mostly relies on small-scale tissue samples or single-omics data. Especially in the study of tumor heterogeneity, mixed sample detection can easily mask key molecular features, making it difficult to identify specific markers. For example, although spatial omics technology can preserve the in situ molecular information of tissues, its high detection cost limits large-scale sample collection. Most studies can only be carried out on dozens of samples, and the data is not representative enough. Secondly, technical costs and data complexity have exacerbated resource barriers. Although the application of technologies such as high-throughput sequencing and multi-color immunofluorescence has increased the data dimension, the equipment investment, reagent consumption and professional manpower costs remain high, which is difficult for small and medium-sized institutions to bear. In addition, the integration of omics data requires complex preprocessing and standardization processes, and the cross-platform data compatibility is poor, which further increases the difficulty and time cost of analysis. Finally, poor model generalization has become a key obstacle to application implementation. Although existing machine learning methods can achieve high accuracy on a single dataset, their performance decreases significantly when verified across populations and institutions due to sample bias, batch effects, or biological heterogeneity.
[0003] Taken together, these limitations highlight the urgency of developing low-cost, high-throughput, and robust bioinformatics analysis frameworks. Summary of the Invention
[0004] The purpose of this application is to provide a biomarker model training method, system, electronic device and medium to solve the problem of the lack of low-cost, high-throughput and robust biological analysis framework in the existing technology.
[0005] In a first aspect, the present application provides a biomarker model training method. The biomarker model training method includes: matching a tissue chip corresponding to the distribution of sample points according to a scanning area; using a targeted spatial multi-omics scanning instrument to perform a one-time scanning experiment on the scanning area to obtain gene expression information of specific cells, including epithelial cells or tumor cells, and immune cells; using the gene expression information of the specific cells to train an algorithm model to obtain a trained model; analyzing the trained model using an evaluation index to obtain performance indicators of each trained model; and screening the trained models according to the performance indicators to obtain a target model.
[0006] In this application, a tissue microarray is customized using sample point distribution to obtain a scanning area. Gene expression information for specific cells is then acquired through a targeted spatial multi-omics scanning experiment. Multiple algorithm models are trained using this gene expression information from specific cells, and these models are then analyzed and screened using evaluation metrics to obtain a target model. This biomarker model training method can simultaneously meet the requirements of high-throughput and high-spatial-resolution samples, obtaining multiple trained models to screen for a more accurate target model.
[0007] In an implementation of the first aspect, matching a tissue chip corresponding to the sample point distribution according to a scanning area includes: analyzing the distribution of paraffin-embedded samples to customize a tissue chip corresponding to the sample point distribution according to the scanning area.
[0008] In one implementation of the first aspect, a one-time scanning experiment is performed on the scanning area using a targeted spatial multi-omics scanning instrument to obtain gene expression information of specific cells, including: using a cell multi-color immunofluorescence staining experiment to identify different types of cells in a specific area of the tissue; using probes to hybridize the expressed genes in the different types of cells and perform directional probe cutting, so as to use a targeted spatial multi-omics scanning instrument to perform a one-time scanning experiment on the scanning area and simultaneously obtain the gene expression information of the specific cells.
[0009] In an implementation of the first aspect, the biomarker model training method further includes: performing data screening, normalization and / or preprocessing on the gene expression information of the specific cells to obtain processed training data.
[0010] In an implementation of the first aspect, the algorithm model is trained using the gene expression information of the specific cells to obtain a trained model, including: grouping and identifying the gene expression information of the specific cells to obtain training data for various algorithm models; and the algorithm model is trained using the corresponding training data to obtain a trained model, wherein the algorithm model includes a generalized linear model, a deep learning model, a gradient boosting machine model, a random forest model, and an extreme gradient boosting model.
[0011] In one implementation of the first aspect, the trained models are analyzed using evaluation indicators to obtain performance indicators of each trained model, including: using area under the curve, accuracy, sensitivity, specificity, F1 score and / or cross-validation to analyze each trained model to obtain performance indicators of each trained model.
[0012] In an implementation of the first aspect, screening the trained model according to the performance indicators to obtain the target model includes: using curve analysis, confusion matrix heat map, and cross-validation result line chart to perform result visualization analysis on the performance indicators to obtain the analysis results of the performance indicators; screening according to the performance indicator analysis results to obtain the target model.
[0013] In the second aspect, the present application provides a biomarker model training system. The biomarker model training system includes: a region acquisition module for matching a tissue chip corresponding to the distribution of sample points according to a scanning area; a region scanning module for performing a one-time scanning experiment on the scanning area using a targeted spatial multi-omics scanning instrument to obtain gene expression information of specific cells, the specific cells including epithelial cells or tumor cells, and immune cells; a model training module for training an algorithm model using the gene expression information of the specific cells to obtain a trained model; a model analysis module for analyzing the trained model using an evaluation index to obtain performance indicators of each trained model; and a model screening module for screening the trained model according to the performance indicators to obtain a target model.
[0014] In a third aspect, the present application provides an electronic device comprising: a memory for storing a computer program; and a processor for executing the computer program stored in the memory, so that the electronic device performs the biomarker model training method described in any one of the first aspects.
[0015] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the biomarker model training method according to any one of the first aspects. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 Shown is a flow chart of the biomarker model training method described in the embodiments of the present application.
[0017] Figure 2 Shown is a schematic diagram of the process of obtaining biomarker model data as described in the examples of this application.
[0018] Figure 3 Shown is a schematic diagram of the algorithm model construction described in the embodiments of the present application.
[0019] Figure 4 Shown is a bar chart of the optimal model performance described in the examples of this application.
[0020] Figure 5Shown are ROC curves of the five types of algorithm models described in the embodiments of this application.
[0021] Figure 6 Shown is a schematic diagram of the structure of the biomarker model training system described in the embodiments of the present application.
[0022] Figure 7 Shown is a structural schematic diagram of an electronic device described in an embodiment of the present application.
[0023] Component number description
[0024] 100 Biomarker Model Training System
[0025] 110 Area Acquisition Module
[0026] 120 Area Scan Module
[0027] 130 Model Training Module
[0028] 140 Model Analysis Module
[0029] 150 Model Screening Module
[0030] 700 Electronic Equipment
[0031] 710 Memory
[0032] 720 processor
[0033] 730 Display
[0034] Steps S11 to S15
[0035] Steps S121 to S122 DETAILED DESCRIPTION
[0036] The following describes the embodiments of the present application through specific examples. Those skilled in the art can easily understand the other advantages and effects of the present application from the content disclosed in this specification. The present application can also be implemented or applied through other different specific embodiments. The details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that the following embodiments and features in the embodiments can be combined with each other unless they conflict.
[0037] It should be noted that in the embodiments of this application, words such as "optionally" or "for example" represent examples, illustrations, or descriptions. Any embodiment or design described in this application as "optionally" or "for example" should not be interpreted as being more preferred or advantageous than other embodiments or designs. Rather, the use of words such as "optionally" or "for example" is intended to present the relevant concepts in a concrete manner.
[0038] In the embodiments of the present application, "at least one" refers to one or more, and "more" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: the existence of A alone, the existence of A and B at the same time, and the existence of B alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can represent: a, b, c, ab, ac, bc or abc, where a, b, c can be single or multiple.
[0039] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present application. Therefore, the illustrations only show components related to the present application and are not drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component can be changed at will, and the component layout type may also be more complicated.
[0040] Malignant tumors have become a major public health challenge threatening human health. Identifying tumor-specific biomarkers facilitates early screening, diagnosis, and personalized treatment. The development of traditional omics technologies has made tumor marker screening possible. However, the use of mixed biomolecules in tissue samples can mask the discovery of specific molecular signatures due to tumor heterogeneity. The emergence of spatial omics technologies not only comprehensively covers transcriptome-encoding genes at the molecular level but also preserves the spatial location of components, facilitating the identification of specific biomarkers based on tissue cell distribution from a histopathological perspective. However, the high cost of spatial omics technologies makes it difficult to identify spatially correlated biomarkers from large numbers of samples. The advent of tissue microarrays has made it possible to monitor the expression of specific molecules in large numbers of samples. However, without omics tools, the discovery of new, more spatially specific biomarkers is difficult. Overall, existing technologies have the following problems: sample size limitation: traditional tissue chip technology is difficult to meet the needs of high throughput and high spatial resolution at the same time; spatial omics technology is expensive: existing spatial omics detection methods are difficult to support large sample size detection to obtain sufficient data; poor model generalization: a single machine learning algorithm is difficult to balance prediction accuracy and feature interpretability.
[0041] To address at least the above-mentioned issues, an embodiment of the present application provides a biomarker model training method. The biomarker model training method includes: matching a tissue chip corresponding to the sample point distribution according to a scanning area; performing a one-time scanning experiment on the scanning area using a targeted spatial multi-omics scanning instrument to obtain gene expression information of specific cells, including epithelial cells or tumor cells, and immune cells; training an algorithm model using the gene expression information of the specific cells to obtain a trained model; analyzing the trained models using evaluation indicators to obtain performance indicators of each trained model; and screening the trained models based on the performance indicators to obtain a target model.
[0042] In the examples of this application, a tissue chip is customized using sample point distribution to obtain a scanning area, and gene expression information of specific cells is obtained through scanning experiments using a targeted spatial multi-omics scanning instrument. Multiple algorithm models are trained using this gene expression information of specific cells, and then the algorithm models are analyzed and screened using evaluation indicators to obtain a target model. This biomarker model training method can simultaneously meet the sample requirements of high throughput and high spatial resolution, obtain multiple trained models, and screen for a more accurate target model.
[0043] The technical solutions in the embodiments of the present application will be described in detail below with reference to the accompanying drawings in the embodiments of the present application.
[0044] The following examples of this application provide a biomarker model training method: Figure 1 Shown is a flow chart of the biomarker model training method described in the embodiments of this application. Figure 1 As shown, the biomarker model training method includes steps S11 to S15.
[0045] Step S11: Matching a tissue chip corresponding to the sample point distribution based on the scan area. Tissue chips are a high-throughput biotechnology tool used to fix dozens to hundreds of tiny tissue samples in a preset array on the same glass slide, enabling simultaneous analysis of large-scale samples.
[0046] In step S12, a targeted spatial multi-omics scanning instrument is used to perform a one-time scanning experiment on the scanning area to obtain gene expression information for specific cells, including epithelial cells, tumor cells, and immune cells. The targeted spatial multi-omics scanning instrument is a GeoMx DSP (GeoMx Digital Spatial Profiler) instrument, which only images the scanning area. Cells outside the scanning area cannot be detected by the instrument.
[0047] Step S13: training the algorithm model using the gene expression information of the specific cells to obtain a trained model.
[0048] Step S14: Analyze the trained models using evaluation indicators to obtain performance indicators of each trained model.
[0049] Step S15: screening the trained model according to the performance index to obtain a target model.
[0050] In some possible implementations, the GeoMx DSP instrument cannot detect tissues that exceed the scanning area of 13.6mm × 33.9mm. Therefore, tissue microarrays of corresponding sizes are matched to the sample point distribution, maximizing coverage within a single GeoMx DSP instrument run to facilitate subsequent data generation and mining. Multicolor immunofluorescence experiments, using antibodies against epithelial cells and immune cells, along with nuclear dyes, combined with HE staining images from clinical slides, allow for simultaneous capture of gene expression information for epithelial cells, tumor cells, and immune cells in the same region. HE staining, also known as hematoxylin and eosin staining, is a staining method widely used in teaching and research in histology, embryology, and pathology. The gene expression data for specific cells is divided into a training dataset and an independent test set, with a training-to-test ratio of 8:2. The algorithm model is trained using the gene expression information for specific cells in the training dataset, and hyperparameters (such as learning rate, tree depth, and regularization coefficient) are optimized using grid search to obtain the trained model. The trained models are analyzed using evaluation indicators to obtain performance indicators of each trained model. The trained models are screened based on the performance indicators to obtain target models as biomarker models.
[0051] In the examples of this application, a tissue chip is customized using sample point distribution to obtain a scanning area, and gene expression information of specific cells is obtained through scanning experiments using a targeted spatial multi-omics scanning instrument. Multiple algorithm models are trained using this gene expression information of specific cells, and then the algorithm models are analyzed and screened using evaluation indicators to obtain a target model. This biomarker model training method can simultaneously meet the sample requirements of high throughput and high spatial resolution, obtain multiple trained models, and screen for a more accurate target model.
[0052] In one embodiment of the present application, matching a tissue chip corresponding to the sample point distribution according to a scanning area includes: analyzing the distribution of paraffin-embedded samples to customize a tissue chip corresponding to the sample point distribution according to the scanning area.
[0053] In some possible implementations, tissue microarrays can be customized using paraffin-embedded samples to capture sufficient sample data for downstream data analysis, based on the size of the tissue microarray. Customized tissue microarrays for DSP (Digital Spatial Profiler) experiments can be used to test more samples in a single experiment, generating more training data for algorithm model training, making the cost of testing a single sample more affordable for testing larger samples.
[0054] In the embodiments of the present application, the method of using customized tissue chips to conduct centralized testing of large samples can reduce the cost of single sample testing and improve sample testing efficiency.
[0055] Figure 2 Shown is a schematic diagram of the process of obtaining biomarker model data described in the embodiments of this application. Figure 2 As shown, step S12 includes steps S121 to S122.
[0056] Step S121 , using a multi-color immunofluorescence staining experiment to identify different types of cells in a specific area of the tissue.
[0057] Step S122 , hybridizing the expressed genes in the different types of cells with probes and performing directional probe cutting, so as to perform a one-time scanning experiment on the scanning area using a targeted spatial multi-omics scanning instrument, and simultaneously obtain the gene expression information of the specific cells.
[0058] In some possible implementations, different cell types in specific areas of tissue are identified using nuclear dyes, PanCK antibodies for marking epithelial cells or tumor cells, and CD45 antibodies for identifying immune cells. Probes hybridized to the sections are then precisely cut to capture gene expression information for different cell types in each sample. The DSP instrument can detect expression information for 18,000 genes across the entire genome, covering all molecular indicators in the sample and facilitating downstream marker screening.
[0059] In other possible implementations, the DSP experiment scanning area is a rectangular area with a range of 13.6mm×33.9mm. In order to meet the needs of different sample array areas, sample points based on 1.5mm diameter and 2mm diameter are designed respectively. According to the arrangement matrix, a single experimental slide can hold 119 samples and 65 samples respectively, realizing the simultaneous scanning of multiple samples. The DSP experiment obtains raw data through high-throughput second-generation sequencing, performs data comparison and quantification, and the data type obtained is a gene expression matrix. Each row and column is the number of sequencing reads obtained by sequencing of each gene expression. The number represents the number of detections. The larger the expression value, the higher the expression level of the gene in the sample. The vertical column is the sample name label and cell type label of each sample.
[0060] In one embodiment of the present application, the biomarker model training method further includes: performing data screening, normalization and / or preprocessing on the gene expression information of the specific cells to obtain processed training data.
[0061] In some possible implementations, the Q3 algorithm is used to preprocess the gene expression information of specific cells. Quantile 3normalization (3rd quartile of all selected targets, referred to as Q3) is a method of data normalization. Q3 means that for all expression distributions, the data of the top 25% expression amount of the AOI distribution are selected as the basis for normalization. This data selection method has been proven to be robust in its representativeness of the overall expression distribution and is recommended as a normalization method for DSP type batch correction. The variation analysis method is used to screen highly variable genes, and the coefficient of variation is calculated by dividing the standard deviation by the mean. The genes are sorted from large to small according to the coefficient of variation, and finally the genes with the highest variation are selected. For example, the top 2000 genes are used as the screened genes by default. The gene expression information of specific cells is screened, normalized and / or preprocessed, and highly variable genes are screened for model construction.
[0062] In one embodiment of the present application, the algorithm model is trained using the gene expression information of the specific cells to obtain a trained model, including: grouping and labeling the gene expression information of the specific cells to obtain training data for various algorithm models; and the algorithm model is trained using the corresponding training data to obtain a trained model, wherein the algorithm model includes a generalized linear model, a deep learning model, a gradient boosting machine model, a random forest model, and an extreme gradient boosting model.
[0063] In some possible implementations, Figure 3 Shown is a schematic diagram of the algorithm model construction described in the embodiment of this application. Figure 3As shown in the figure, the algorithm model construction process includes training the algorithm model with a training set, evaluating the algorithm model performance with a test set, cross-validating the algorithm model's balance, and analyzing the algorithm model's performance using performance indicators. Specifically, five different algorithm models were trained using samples from the training dataset. Except for the extreme gradient boosting model, all models used a grid search method for parameter optimization to obtain the best-performing model. The algorithm models included a generalized linear model (GLM), a deep learning model, a gradient boosting machine model (GBM), a random forest model, and an extreme gradient boosting model (XGBoost).
[0064] For the GLM model, grid search was used to optimize the regularization parameters: alpha (0.1, 0.3, 0.5, 0.7, 0.9) controlled the L1 / L2 mixing ratio, and lambda (0.0001, 0.001, 0.01, 0.1) adjusted the penalty intensity; the model was set to the binomial distribution family, and the category balance and lambda automatic search functions were enabled to adapt to imbalanced data, thereby improving the robustness and interpretability of the model in biomedical data.
[0065] For deep learning models, grid search was performed across hidden layer structures ([256, 128], [512, 256], [512, 512], [768, 384], [1024, 512]), training epochs (30, 50, 70, 100), initial learning rates (0.001, 0.005, 0.01, 0.02), and learning rate decay coefficients (1e-8 to 1e-5). The model was based on a Bernoulli distribution, with a class balancing strategy implemented to ensure a balanced weighting of positive and negative samples.
[0066] For the GBM model, grid search parameters include the number of trees (100, 200, 300, 500), the maximum depth of a single tree (6, 8, 10, 12), the learning rate (0.001-0.02), and the ratio of row sampling to column sampling (0.7, 0.8, 0.9). Class balancing was enabled during training.
[0067] For the random forest model, the parameter search range covers the number of trees (50, 100, 200, 300), maximum depth (6-12), minimum number of node samples (5, 10, 15, 20), and sample-to-feature ratio (0.7, 0.8, 0.9). The model enables class balancing and accelerates training through parallel computing, improving the efficiency of large-scale data processing.
[0068] For the XGBoost model, fixed parameters were set as follows: learning rate (eta = 0.1), maximum tree depth (6), minimum child node weight (1), and sample-to-feature ratio (0.8). The optimization objective was set to binary logistic regression, with AUC as the evaluation metric. Early stopping was performed when the validation set performance did not improve for 10 consecutive rounds.
[0069] With the exception of the XGBoost model, the remaining four models all leveraged the H2O framework for automated grid search. The XGBoost model used fixed parameter configurations. Class balancing was enabled during the training process for each algorithm model to address potential class imbalances, and the same random seed was used to ensure reproducible results.
[0070] In one embodiment of the present application, the trained models are analyzed using evaluation indicators to obtain performance indicators of each trained model, including: using area under the curve, accuracy, sensitivity, specificity, F1 score and / or cross-validation to analyze each trained model to obtain performance indicators of each trained model.
[0071] In some possible implementations, the area under the AUC-ROC curve is used to evaluate the classification ability of the model, the accuracy is used to determine the overall prediction accuracy of the model, the sensitivity is used to determine the true positive rate of the model, the specificity is used to determine the true negative rate of the model, the F1 score is used to determine the harmonic mean of the precision and recall of the model, and five-fold cross-validation stratified sampling is used to ensure that the category distribution in each fold is consistent, so as to obtain the performance indicators of the trained model.
[0072] In one embodiment of the present application, the trained model is screened according to the performance indicators to obtain the target model, including: visually analyzing the performance indicators using curve analysis, confusion matrix heat map, and cross-validation result line chart to obtain the analysis results of the performance indicators; and screening according to the performance indicator analysis results to obtain the target model.
[0073] In some possible implementations, the performance indicators are visualized using curve analysis, confusion matrix heat maps, and cross-validation result line graphs to obtain the analysis results of the performance indicators. The analysis results of the trained model are visualized, and the classification performance and discrimination ability of the model are evaluated using ROC curve analysis. Each curve is labeled with the corresponding AUC (Area Under the Curve) value. The closer the curve is to the upper right corner, the better the model performance. The closer the AUC value is to 1, the stronger the model's ability to distinguish between positive and negative samples. Comparing the ROC curves of different models can intuitively judge the model performance.
[0074] The confusion matrix heatmap details various errors in the model's predictions, including true positives (TP), false positives (FP), true negatives (TN), and false negatives (FN). The heatmap's color depth indicates numerical value. Diagonal elements represent the number of correctly classified samples, while off-diagonal elements represent the number of incorrectly classified samples. This can be used to analyze model performance deviations across different categories. This analysis examines four key metrics: true positives (TP), false positives (FP), true negatives (TN), and false negatives (FN). These metrics include recall (TP / (TP+FN)), precision (TP / (TP+FP)), and overall accuracy ((TP+TN) / total).
[0075] Figure 4 The bar chart of the optimal model performance described in the embodiment of this application is shown. Figure 4 As shown in the figure, the display includes accuracy, precision, sensitivity, specificity, F1 score, and AUC value. The bar height in the display reflects the optimal value of each metric, intuitively showing the performance differences between different models and used to identify the best model for each metric.
[0076] The cross-validation performance box plot is used to demonstrate the stability and distribution characteristics of model performance. The display content includes different models on the horizontal axis and evaluation indicators such as AUC value on the vertical axis. The box plot shows the distribution of 5-fold cross-validation results. The box height in the display content reflects the performance fluctuation range, the median position indicates the typical performance level, and the outliers indicate abnormal performance, which can evaluate the stability and reliability of the model performance. Specifically, the median line shows the trend of data concentration, and the dotted line reflects a wider distribution range. By comparing the median height and the box size, the overall performance and stability of each model on different indicators can be quickly judged. The higher the median and the higher the box position, the better the model is at this indicator. The smaller the box and the shorter the whiskers, the more stable the performance.
[0077] The cross-validation line chart displays the performance trends of the model during the cross-validation process. The horizontal axis shows the number of validation folds, and the vertical axis shows metrics such as the AUC value. Different models are represented by different colored lines. A stable line indicates stable model performance, while a line with large fluctuations indicates that the model is sensitive to the data, helping to identify whether the model is overfitting.
[0078] Screening is performed based on the performance indicator analysis results to obtain the target model. Figure 5 Shown is the ROC curve diagram of the five types of algorithm models described in the embodiment of this application. Figure 5As shown, the horizontal axis is the false positive rate, and the vertical axis is the true positive rate. Each line represents the model's performance in distinguishing between positive and negative classes at different thresholds. As the threshold changes from high to low, the curve extends from the lower left corner to the upper right. A curve closer to the upper left corner indicates a higher true positive rate at a lower false positive rate, indicating better discrimination. The AUC is the area under the curve; a closer approach to 1 indicates a stronger overall model's discriminative ability. An AUC of approximately 0.5 is equivalent to random guessing. Figure 5 The AUC values and corresponding curve colors for the five models are labeled in the lower right corner (red: GLM model, yellow: deep learning model, green: GBM model, blue: random forest model, purple: XGBoost model), allowing for direct comparison of the performance of each model. The AUC values for all five algorithm models are above 0.9, indicating that they are able to identify the majority of positive examples while maintaining a low false positive rate. The deep learning model, with an AUC of 1, is the best performing model. The random forest model and XGBoost model achieve an AUC of 0.96, while the GLM model and GBM model achieve an AUC of 0.92, demonstrating good performance.
[0079] Comprehensive analysis revealed that deep learning models achieved the best performance in terms of AUC, accuracy, F1 score, and sensitivity. All five algorithm models achieved the highest specificity of 1. GLM and deep learning models demonstrated excellent stability. Feature analysis revealed that high-frequency features such as INMT were selected across multiple models, demonstrating their importance. Medium-frequency features such as FCRL1, CLU, and SERPINF1 were present in some models, while low-frequency features were less common. Overall, deep learning models demonstrated the best overall performance, while GLM and DL models demonstrated advantages in stability and reliability.
[0080] Figure 6 Shown is a schematic diagram of the structure of the biomarker model training system described in the embodiment of this application. Figure 6 As shown, the biomarker model training system 100 includes a region acquisition module 110 , a region scanning module 120 , a model training module 130 , a model analysis module 140 and a model screening module 150 .
[0081] The region acquisition module 110 is used to match a tissue chip corresponding to the distribution of sample points according to the scanning region.
[0082] The area scanning module 120 is used to perform a one-time scanning experiment on the scanning area using a targeted spatial multi-omics scanning instrument to obtain gene expression information of specific cells, including epithelial cells or tumor cells, and immune cells.
[0083] The model training module 130 is used to train the algorithm model using the gene expression information of the specific cells to obtain a trained model.
[0084] The model analysis module 140 is used to analyze the trained models using evaluation indicators to obtain performance indicators of each trained model.
[0085] The model screening module 150 is used to screen the trained model according to the performance index to obtain a target model.
[0086] In some possible implementations, in the spatial biomarker detection scenario, any of the five types of trained models (except for models that cannot be output) is used as the target model to predict the sample to be tested, the output feature indicators are ranked in importance, and the gene functions of the top 10 important indicators are interpreted to understand the rationality of the model as a biomarker from the perspective of molecular function.
[0087] It should be noted that the modules 110 to 150 included in the biomarker model training system 100 are Figure 1 Steps S11 to S15 in the biomarker model training method shown correspond to each other and are not described in detail here.
[0088] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices or methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of modules / units is only a logical function division. There may be other division methods in actual implementation. For example, multiple modules or units can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules or units, which can be electrical, mechanical or other forms.
[0089] The modules / units described as separate components may or may not be physically separate, and the components displayed as modules / units may or may not be physical modules, that is, they may be located in one place or distributed across multiple network elements. Some or all of the modules / units may be selected according to actual needs to achieve the purpose of the embodiments of the present application. For example, the functional modules / units in the various embodiments of the present application may be integrated into a processing module, or each module / unit may exist physically separately, or two or more modules / units may be integrated into a single module / unit.
[0090] Those skilled in the art should further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0091] An embodiment of the present application also provides an electronic device. Figure 7 Shown is a schematic diagram of the structure of the electronic device 700 according to an embodiment of the present application. Figure 7 As shown, in this embodiment, the electronic device 700 includes a memory 710 and a processor 720.
[0092] The memory 710 is used to store computer programs; preferably, the memory 710 includes: ROM, RAM, magnetic disk, USB flash drive, memory card or optical disk, etc., various media that can store program codes.
[0093] Specifically, the memory 710 may include a computer system readable medium in the form of a volatile memory, such as a random access memory (RAM) and / or a cache memory. The electronic device 700 may further include other removable / non-removable, volatile / non-volatile computer system storage media. The memory 710 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the various embodiments of the present application. It will be understood that the memory 710 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM) or a programmable read-only memory (PROM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM) and synchronous static random access memory (SSRAM). Memory described in the embodiments of the present invention is intended to include, but not be limited to, these and any other suitable types of memory.
[0094] The processor 720 is connected to the memory 710 and is used to execute the computer program stored in the memory 710 so that the electronic device 700 executes the biomarker model training method described in any embodiment of the present application.
[0095] Optionally, the processor 720 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, and discrete hardware components.
[0096] Optionally, the electronic device 700 in this embodiment may further include a display 730. The display 730 is communicatively connected to the memory 710 and the processor 720, and is used to display a graphical user interface (GUI) interactive interface related to the biomarker model training method described in the embodiment of the present application.
[0097] The present application also provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the biomarker model training method described in any embodiment of the present application is implemented.
[0098] As used in this specification, the terms "component," "module," "system," and the like are used to represent computer-related entities, hardware, firmware, a combination of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. By way of illustration, both an application running on a computing device and a computing device can be a component. One or more components can reside in a process and / or an execution thread, and a component can be located on a computer and / or distributed between two or more computers. In addition, these components can be executed from various computer-readable media having various data structures stored thereon. Components can communicate, for example, via local and / or remote processes based on signals having one or more data packets (e.g., data from two components interacting with another component on a local system, a distributed system, and / or a network, such as the Internet interacting with other systems via signals).
[0099] The descriptions of the processes or structures corresponding to the above figures have different emphases. For parts that are not described in detail in a certain process or structure, please refer to the relevant descriptions of other processes or structures.
[0100] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Anyone skilled in the art may modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or alterations made by one of ordinary skill in the art without departing from the spirit and technical concepts disclosed in this application shall be covered by the claims of this application.
Claims
1. A biomarker model training method, characterized in that: include: Matching the tissue chip corresponding to the sample point distribution according to the scanning area; Performing a one-time scanning experiment on the scanning area using a targeted spatial multi-omics scanning instrument to obtain gene expression information of specific cells, including epithelial cells or tumor cells, and immune cells; Training an algorithm model using the gene expression information of the specific cells to obtain a trained model; Analyzing the trained models using evaluation indicators to obtain performance indicators of each trained model; The trained model is screened according to the performance indicator to obtain a target model.
2. The biomarker model training method according to claim 1, characterized in that Tissue chips that match the distribution of sample points according to the scan area include: The distribution of the paraffin-embedded samples is analyzed to customize a tissue chip corresponding to the distribution of the sample points according to the scanning area.
3. The biomarker model training method according to claim 1, characterized in that A one-time scanning experiment is performed on the scanning area using a targeted spatial multi-omics scanning instrument to obtain gene expression information of specific cells, including: Use multi-color immunofluorescence staining to identify different cell types in specific areas of the tissue; The expressed genes in the different types of cells are hybridized using probes and directional probe cutting is performed, so that a targeted spatial multi-omics scanning instrument is used to perform a one-time scanning experiment on the scanning area, and the gene expression information of the specific cells is obtained at the same time.
4. The biomarker model training method according to claim 1, characterized in that Also includes: Data screening, normalization and / or preprocessing are performed on the gene expression information of the specific cells to obtain processed training data.
5. The biomarker model training method according to claim 1, characterized in that: The algorithm model is trained using the gene expression information of the specific cells to obtain a trained model including: performing grouping and identification processing on the gene expression information of the specific cells to obtain training data for various algorithm models; The algorithm model is trained using corresponding training data to obtain a trained model, wherein the algorithm model includes a generalized linear model, a deep learning model, a gradient boosting machine model, a random forest model, and an extreme gradient boosting model.
6. The biomarker model training method according to claim 1, characterized in that: The trained models are analyzed using evaluation indicators to obtain performance indicators of each trained model, including: The trained models were analyzed using area under the curve, accuracy, sensitivity, specificity, F1 score and / or cross-validation to obtain performance indicators of the trained models.
7. The biomarker model training method according to claim 1, characterized in that: Screening the trained model according to the performance indicators to obtain a target model includes: Performing a visual analysis of the performance indicators using curve analysis, confusion matrix heat map, and cross-validation result line graph to obtain analysis results of the performance indicators; Screening is performed based on the performance indicator analysis results to obtain the target model.
8. A biomarker model training system, characterized in that: include: A region acquisition module, used to match the tissue chip corresponding to the sample point distribution according to the scanning area; A regional scanning module, configured to perform a one-time scanning experiment on the scanning area using a targeted spatial multi-omics scanning instrument to obtain gene expression information of specific cells, including epithelial cells or tumor cells, and immune cells; A model training module, configured to train an algorithm model using the gene expression information of the specific cells to obtain a trained model; A model analysis module is used to analyze the trained models using evaluation indicators to obtain performance indicators of each trained model; The model screening module is used to screen the trained model according to the performance index to obtain a target model.
9. An electronic device, characterized in that: The electronic device comprises: memory for storing computer programs; A processor, wherein the processor is configured to execute the computer program stored in the memory so as to enable the electronic device to perform the biomarker model training method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the biomarker model training method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Biomarker determination method and device, storage medium and electronic equipment
CN114898804A
Method and equipment for constructing tumor curative effect prediction model of PD-1 monoclonal antibody and medium
CN117524294A
Autoimmune disease biomarker screening method based on longitudinal and transverse bidirectional clustering
CN117524315A
Method for training machine learning model to analyze immunohistochemically stained images and computing system performing same
WO2023191472A1