Genomics-based biomarker screening method and system

The genomic method screens cancer-related biomarkers, and uses transcriptome sequencing data and intergenic interaction networks to solve the problem of low specificity in traditional biomarker screening, achieving higher sensitivity of cancer diagnosis and treatment evaluation.

CN119360969BActive Publication Date: 2025-09-02PEKING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411946317.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-09-02
Estimated Expiration
2044-12-27

AI Technical Summary

Technical Problem

During the traditional biomarker screening process, the biomarker specificity is not high and the sensitivity is low, which can easily lead to missed examinations, resulting in missed diagnosis or delayed diagnosis in patients.

Method used

Based on genomics biomarker screening method, the biomarker expression data of significant genes is constructed through transcriptome sequencing data, functional genomics and clinical genomics. The posterior distribution of biomarkers and cancer diseases is estimated using logistic regression models, and combined with the intergenic interaction network, significant interaction genes are screened out, and a list of biomarkers related to cancer diseases is finally obtained.

Benefits of technology

It improves the sensitivity of the biomarker screening process, avoids missed examinations, and ensures early diagnosis and treatment effect evaluation of cancer diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119360969B_ABST
    Figure CN119360969B_ABST
Patent Text Reader

Abstract

The present application discloses a genomics-based biomarker screening method and system, which relates to the field of biological detection technology. The method comprises: determining biomarker expression data associated with significant genes based on transcriptome sequencing data of cancer disease biological samples, and estimating the posterior distribution of the association between biomarkers and cancer diseases through a logistic regression model; determining significant interacting genes based on the posterior distribution and the gene interaction network; and screening to obtain a list of biomarkers associated with cancer diseases based on the significant interacting genes. Since the present application screens the posterior distribution and significant interacting genes through transcriptome sequencing data of cancer disease biological samples, it avoids the situation in which missed detections are caused by low biomarker specificity in the traditional biomarker screening process, thereby obtaining a list of biomarkers with high correlation with cancer diseases, and improving the sensitivity of the biomarker screening process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of biological detection technology, and in particular to a genomics-based biomarker screening method and system. Background Art

[0002] Cancer (malignant tumors) is a major threat to human health and life. Its incidence and mortality rates are increasing year by year worldwide, creating a serious situation. Cancer biomarkers are highly disease-specific, facilitating early diagnosis and evaluating treatment effectiveness. Researchers typically use a variety of methods to identify and discover cancer biomarkers, including immunology, molecular biology, and genomics.

[0003] However, due to various factors such as the interactions between biomarkers, differences in disease samples, differences in biomarker levels between individuals, and the fact that the expression levels of some biomarkers do not change significantly in the early stages of the disease or when the condition is mild, the biomarker screening process has low sensitivity. At the same time, some biomarkers cross-react with other diseases, which can also reduce the specificity of biomarkers, making it easy to miss detections, leading to missed or delayed diagnosis of patients. Summary of the Invention

[0004] The main purpose of this application is to provide a genomics-based biomarker screening method and system, aiming to solve the shortcomings of low biomarker specificity and low screening sensitivity in the traditional biomarker screening process, which easily leads to missed detection and technical problems such as missed diagnosis or delayed diagnosis of patients.

[0005] To achieve the above objectives, the present application proposes a genomics-based biomarker screening method, which comprises:

[0006] Determining biomarker expression data associated with significant genes based on transcriptome sequencing data of cancer disease biological samples, wherein the transcriptome sequencing data is constructed based on functional genomics and clinical genomics;

[0007] estimating the posterior distribution of the association between the biomarker and the cancer disease by a logistic regression model based on the biomarker expression data of the significant genes;

[0008] determining significantly interacting genes in a cancer disease biological sample based on the posterior distribution and the gene-gene interaction network;

[0009] Based on the significantly interacting genes, a list of biomarkers associated with cancer diseases was screened and obtained.

[0010] In one embodiment, the step of determining biomarker expression data associated with significant genes based on transcriptome sequencing data of a cancer disease biological sample comprises:

[0011] Obtain biological samples of various types of cancer diseases;

[0012] Performing transcription amplification on the cancer disease biological sample, and sequencing to obtain transcriptome sequencing data based on functional genomics and clinical genomics;

[0013] Cluster screening is performed on the transcriptome sequencing data to obtain biomarker expression data related to significant genes.

[0014] In one embodiment, the step of performing cluster screening on the transcriptome sequencing data to obtain biomarker expression data associated with significant genes comprises:

[0015] Preliminary screening of the transcriptome sequencing data is performed based on gene expression levels to obtain gene difference data;

[0016] Performing an association test based on the gene difference data to obtain the gene difference data after the association test;

[0017] Performing feature screening on the gene difference data after the association test using a hierarchical clustering algorithm to obtain hierarchical association results between genes and cancer diseases;

[0018] Based on the hierarchical association results, biomarker expression data associated with significant genes are determined.

[0019] In one embodiment, the step of estimating the posterior distribution of the association between the biomarker and the cancer disease using a logistic regression model based on the biomarker expression data of the significant genes comprises:

[0020] Based on the biomarker expression data of the significant genes, a logistic regression model is used to describe the association between the biomarkers and cancer diseases;

[0021] The unnormalized associations between biomarkers and cancer diseases were combined using a Bayesian approach to obtain the posterior distribution of the associations between biomarkers and cancer diseases.

[0022] In one embodiment, the step of determining the significantly interacting genes in the cancer disease biological sample based on the posterior distribution and the gene interaction network comprises:

[0023] Determining significant interactions of genes in a cancer disease state and a normal state based on the posterior distribution and the gene-gene interaction network;

[0024] According to the significant interactions, significantly interacting genes in the cancer disease biological sample are determined.

[0025] In one embodiment, the step of screening and obtaining a list of biomarkers associated with cancer based on the significantly interacting genes comprises:

[0026] Perform feature importance evaluation on the significantly interacting genes to obtain importance scores;

[0027] Sorting the importance scores to obtain a score list;

[0028] The significantly interacting genes are screened based on the score list to obtain a list of biomarkers associated with cancer diseases.

[0029] In one embodiment, the step of performing feature importance assessment on the significantly interacting genes to obtain importance scores comprises:

[0030] Initialize the gradient boosting decision tree model;

[0031] Training the gradient boosting decision tree model using a cancer gene database;

[0032] The feature importance of the significantly interacting genes was evaluated based on the trained gradient boosting decision tree model to obtain importance scores.

[0033] In addition, to achieve the above objectives, the present application also proposes a genomics-based biomarker screening system, the system comprising:

[0034] A sequencing data module is used to determine biomarker expression data associated with significant genes based on transcriptome sequencing data of cancer disease biological samples, wherein the transcriptome sequencing data is obtained based on functional genomics and clinical genomics;

[0035] a posterior distribution module for estimating the posterior distribution of the association between the biomarker and the cancer disease through a logistic regression model based on the biomarker expression data of the significant genes;

[0036] An interaction module, for determining significant interacting genes in a cancer disease biological sample based on the posterior distribution and the gene-gene interaction network;

[0037] The marker screening module is used to screen and obtain a list of biomarkers related to cancer diseases based on the significantly interacting genes.

[0038] In one embodiment, the sequencing data module is further used to obtain biological samples of various types of cancer diseases;

[0039] The sequencing data module is further used to perform transcription amplification on the cancer disease biological sample and sequence to obtain transcriptome sequencing data based on functional genomics and clinical genomics;

[0040] The sequencing data module is further used to perform cluster screening on the transcriptome sequencing data to obtain biomarker expression data related to significant genes.

[0041] In one embodiment, the sequencing data module is further used to perform preliminary screening of the transcriptome sequencing data according to gene expression levels to obtain gene difference data;

[0042] The sequencing data module is further configured to perform an association test based on the gene difference data to obtain the gene difference data after the association test;

[0043] The sequencing data module is further used to perform feature screening on the gene difference data after the association test using a hierarchical clustering algorithm to obtain hierarchical association results between genes and cancer diseases;

[0044] The sequencing data module is further used to determine biomarker expression data associated with significant genes based on the hierarchical association results.

[0045] One or more technical solutions proposed in this application have at least the following technical effects: this application first determines the biomarker expression data associated with significant genes based on the transcriptome sequencing data of cancer disease biological samples, and the transcriptome sequencing data is obtained based on functional genomics and clinical genomics; then, based on the biomarker expression data of the significant genes, the posterior distribution of the association between the biomarker and the cancer disease is estimated through a logistic regression model; then, based on the posterior distribution and the gene interaction network, the significant interacting genes in the cancer disease biological sample are determined; finally, based on the significant interacting genes, a list of biomarkers associated with the cancer disease is screened. Since this application obtains the posterior distribution and significant interacting genes associated with the biomarker and the cancer disease through the transcriptome sequencing data of the cancer disease biological sample, it avoids the situation in which missed detection is caused by the low specificity of the biomarker in the traditional biomarker screening process, thereby obtaining a list of biomarkers with high correlation with the cancer disease, and improving the sensitivity of the biomarker screening process. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0047] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0048] Figure 1 A schematic diagram of the process for the genomics-based biomarker screening method according to Example 1 of this application is provided;

[0049] Figure 2 A schematic diagram of the process for Example 2 of the genomics-based biomarker screening method of this application;

[0050] Figure 3 A schematic diagram of the process for Example 3 of the genomics-based biomarker screening method of this application;

[0051] Figure 4 This is a schematic diagram of the functional modules of the genomics-based biomarker screening system according to an embodiment of the present application.

[0052] The realization of the objectives, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0053] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.

[0054] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0055] It should be noted that the execution entity of this embodiment can be a computing service device with transcript sequencing, posterior distribution calculation, and marker screening functions, such as a personal computer or server, or an electronic device capable of performing the aforementioned functions, or a genomics-based biomarker screening system (screening system) that executes the genomics-based biomarker screening method of this application, and this embodiment is not limited thereto. The following uses the screening system as an example to illustrate this embodiment and the following embodiments.

[0056] Example 1: Based on this, the present application provides a genomics-based biomarker screening method, referring to Figure 1 , Figure 1 A schematic diagram of the process of Example 1 of the genomics-based biomarker screening method of this application is provided.

[0057] In this embodiment, the genomics-based biomarker screening method includes steps S10 to S40:

[0058] Step S10: determining biomarker expression data associated with significant genes based on transcriptome sequencing data of cancer disease biological samples, wherein the transcriptome sequencing data is constructed based on functional genomics and clinical genomics.

[0059] It should be noted that cancer disease biological samples are various biological materials related to cancer.

[0060] For example, biological samples of cancer diseases may include tissue specimens (such as tumor tissue, adjacent cancer tissue), blood (such as whole blood, plasma, serum), body fluids (such as urine, cerebrospinal fluid, pleural effusion, etc.), and cells (such as circulating tumor cells, etc.) obtained from cancer patients.

[0061] It should be noted that transcriptome sequencing data is data obtained by sequencing RNA transcribed from cancer disease biological samples under specific conditions.

[0062] For example, the transcription process may include messenger RNA (mRNA), ribosomal RNA (rRNA), transfer RNA (tRNA), and other non-coding RNAs. Transcriptome sequencing can comprehensively and quantitatively analyze gene expression in cancer disease biological samples.

[0063] It should be noted that biomarker expression data refers to gene data in transcriptome sequencing data that are closely related to cancer diseases and can significantly express biomarkers.

[0064] It should be noted that functional genomics is the study of gene functions, including gene expression regulation, the function of gene products, and the interaction between genes.

[0065] Clinical genomics is the application of genomics technology and research results to clinical practice, involving the use of genomic information to diagnose diseases, predict disease prognosis, etc.

[0066] Constructing transcriptome sequencing data based on the standards of functional genomics and clinical genomics can better understand the molecular mechanisms and individual differences of cancer diseases, provide important support for the screening of biomarkers for cancer diseases and their application in medical treatment, and achieve more accurate and effective screening of biomarkers.

[0067] In one embodiment, transcription amplification can be performed on biological samples of cancer diseases first, and then transcriptome sequencing data can be obtained based on the standards of functional genomics and clinical genomics. The screening system can screen the associated characteristics of cancer diseases based on these data to obtain biomarker expression data of significant genes closely related to cancer diseases.

[0068] Step S20: Based on the biomarker expression data of the significant genes, the posterior distribution of the association between the biomarkers and cancer diseases is estimated by a logistic regression model.

[0069] It should be noted that the logistic regression model is a statistical analysis model used to solve binary (or multi-classification) problems. The logistic regression model can be used to describe the association between biomarkers and cancer diseases and determine the association between biomarkers and cancer diseases.

[0070] It should be noted that biomarkers are biochemical molecules that can identify changes or potential changes in systems, organs, tissues, cells, subcellular structures, and genetic structures. Biomarkers can help more accurately diagnose cancer.

[0071] It should be noted that the posterior distribution is a probability distribution of the association between biomarkers and cancer. The posterior distribution can provide a deeper understanding of the role of biomarkers in cancer development and its uncertainty.

[0072] In one embodiment, biomarker expression data can be used as input variables and combined with relevant information about cancer diseases. The degree of association between biomarkers and cancer diseases can be probabilistically described using a logistic regression model to obtain the posterior distribution of the association between biomarkers and cancer diseases.

[0073] Step S30: Determine the significant interacting genes in the cancer disease biological sample based on the posterior distribution and the gene interaction network.

[0074] It should be noted that the intergenic interaction network is a complex network structure formed by various interaction relationships between numerous genes.

[0075] In organisms, genes do not exist and function in isolation. The expression and function of a gene may be influenced by other genes, with interactions occurring in various ways, including synergistic and antagonistic. These interactions are intertwined, forming a vast network.

[0076] It should be noted that significantly interacting genes are genes with particularly prominent interaction relationships in biological samples of cancer diseases and may play a key role in the occurrence and development of cancer.

[0077] For example, in cancer samples, the interactions of certain genes may change significantly, such as increasing or decreasing their synergistic or antagonistic effects in the cancer state. These genes can be considered to have significant interactions in the cancer. Gene interaction networks can help us better understand the interactions between genes and reveal the associations and synergistic patterns between genes that play a key role in the occurrence and development of cancer.

[0078] In one embodiment, the posterior distribution of the association between biomarkers and cancer diseases, combined with the gene interaction network, can identify genes with particularly prominent interactions that play a key role in the development and progression of cancer. This provides an important basis for accurately screening biomarkers related to cancer diseases and facilitates the discovery of new biomarkers.

[0079] Step S40: Based on the significantly interacting genes, a list of biomarkers related to cancer diseases is screened and obtained.

[0080] It should be noted that the biomarker list is a list of biomarkers related to cancer. Through the biomarker list, we can better understand the relationship between biomarkers and cancer diseases, which helps in disease diagnosis.

[0081] In one embodiment, the significantly interacting genes obtained above can be arranged in order of interaction degree or classified according to different physiological functions, and a list of biomarkers related to cancer diseases can be obtained by screening.

[0082] In this embodiment, transcription amplification can first be performed on a cancer biological sample. Transcriptome sequencing data can then be obtained based on functional genomics and clinical genomics standards. A screening system can then use this data to screen for cancer-associated features, obtaining biomarker expression data for significant genes closely associated with cancer. The biomarker expression data can then be used as input variables and combined with relevant information about the cancer. A logistic regression model can be used to probabilistically describe the degree of association between the biomarker and the cancer, yielding a posterior distribution of the biomarker-cancer association. This posterior distribution of the biomarker-cancer association can then be combined with the gene interaction network to identify genes with particularly prominent interactions that play a key role in the development and progression of cancer. This provides an important basis for accurately screening for biomarkers associated with cancer and facilitates the discovery of new biomarkers. Finally, the significantly interacting genes obtained above can be sorted by interaction level or by physiological function to obtain a list of biomarkers associated with cancer.

[0083] In a feasible embodiment, step S40 of this embodiment may include the steps of: performing feature importance evaluation on the significantly interacting genes to obtain importance scores; sorting the importance scores to obtain a score list; and screening the significantly interacting genes based on the score list to obtain a list of biomarkers related to cancer diseases.

[0084] It should be noted that the importance score is a score that evaluates the importance of those significantly interacting genes based on the degree of association with cancer diseases and is expressed in numerical form.

[0085] In this embodiment, the feature importance of significantly interacting genes can be assessed based on the degree of cancer-disease association to obtain an importance score. These importance scores are then ranked to generate a score list. Finally, the biomolecules corresponding to the top-ranked significantly interacting genes in the score list (e.g., the top ten or dozens) are selected as biomarkers to obtain a list of cancer-related biomarkers. Selecting genes with higher scores as biomarkers helps further focus on key genes.

[0086] In a feasible embodiment, the step of performing feature importance evaluation on the significantly interacting genes and obtaining importance scores described in this embodiment includes: initializing a gradient boosting decision tree model; training the gradient boosting decision tree model using a cancer gene database; and performing feature importance evaluation on the significantly interacting genes based on the trained gradient boosting decision tree model to obtain importance scores.

[0087] It should be noted that the gradient boosting decision tree model uses an ensemble learning approach to combine multiple decision trees to predict and evaluate the feature importance of significantly interacting genes. During the construction process, different decision trees are generated through random feature selection and sample sampling. Together, these decision trees form a "forest" that can be used to evaluate and classify significantly interacting genes.

[0088] It should be noted that the cancer gene database is a database that collects, organizes and stores information on cancer-related genes, such as gene sequences, mutations, expression levels, associations with cancer development, and related clinical information.

[0089] In this embodiment, a gradient boosting decision tree model can be initialized. The gradient boosting decision tree model is then trained using a cancer gene database to obtain a model that has learned information about various cancer-related genes. Finally, the trained gradient boosting decision tree model is used to predict and evaluate the feature importance of significantly interacting genes, obtaining importance scores. This can further provide an important basis for searching for biomarkers associated with cancer diseases using the cancer gene database, thereby improving the accuracy of the biomarker list.

[0090] This embodiment provides a genomics-based biomarker screening method. First, transcriptome amplification can be performed on a cancer biological sample. Transcriptome sequencing data can then be obtained based on functional genomics and clinical genomics standards. The screening system can then use this data to screen for cancer-associated features, obtaining biomarker expression data for significant genes closely associated with cancer. The biomarker expression data can then be used as input variables and combined with relevant information about the cancer. A logistic regression model can be used to probabilistically describe the degree of association between the biomarker and the cancer, obtaining a posterior distribution of the biomarker-cancer association. This posterior distribution of the biomarker-cancer association can then be combined with the gene interaction network to identify genes with particularly prominent interactions that play a key role in the development and progression of cancer. This provides an important basis for accurately screening biomarkers associated with cancer and facilitates the discovery of new biomarkers. Finally, the significantly interacting genes obtained above can be sorted by interaction level or by physiological function to obtain a list of biomarkers associated with cancer. Since this embodiment uses transcriptome sequencing data from cancer disease biological samples to screen and obtain the posterior distribution of biomarkers associated with cancer diseases and significantly interacting genes, it avoids the situation in which missed detections due to low biomarker specificity in the traditional biomarker screening process are avoided. As a result, a list of biomarkers with high correlation with cancer diseases can be obtained, thereby improving the sensitivity of the biomarker screening process.

[0091] Example 2: Based on the first embodiment of this application, in the second embodiment of this application, the same or similar contents as those in the above-mentioned embodiment 1 can be referred to the above introduction and will not be described in detail later. Figure 1 and Figure 2 , Figure 2 A schematic diagram of the process provided for Example 2 of the genomics-based biomarker screening method of this application.

[0092] In this example, step S10 includes steps S11 to S13:

[0093] Step S11: Acquire various types of cancer disease biological samples.

[0094] Step S12: performing transcription amplification on the cancer disease biological sample, and sequencing based on functional genomics and clinical genomics to obtain transcriptome sequencing data.

[0095] It should be noted that transcription amplification is the process of increasing the number of RNA transcripts in a sample through specific technical means. The samples after transcription amplification are then sequenced and analyzed. This sequencing process is to determine the sequence information of these RNA molecules. The sequence information is then screened based on functional genomics and clinical genomics to obtain transcriptome sequencing data.

[0096] Step S13: performing cluster screening on the transcriptome sequencing data to obtain biomarker expression data related to significant genes.

[0097] In this embodiment, biological samples related to cancer (such as cancer cells and cancer tissue) can be subjected to transcriptional amplification. Sequencing analysis is then performed on the amplified samples to determine the sequence information of the RNA molecules in the samples. Subsequently, screening is performed based on functional genomics and clinical genomics to obtain transcriptome sequencing data. By applying the standards of functional genomics and clinical genomics, a better understanding of the molecular mechanisms and individual differences in cancer can be achieved, providing important support for the screening of cancer biomarkers for medical treatment, enabling more accurate and effective biomarker screening.

[0098] In a feasible embodiment, step S13 of this embodiment may include the steps of: performing preliminary screening on the transcriptome sequencing data according to gene expression levels to obtain gene difference data; performing association test on the gene difference data to obtain gene difference data after the association test; performing feature screening on the gene difference data after the association test using a hierarchical clustering algorithm to obtain hierarchical association results between genes and cancer diseases; and determining biomarker expression data associated with significant genes based on the hierarchical association results.

[0099] It should be noted that genetic difference data are obtained by preliminary difference screening based on gene expression levels (for example, comparing the gene expression levels of diseased and healthy individuals).

[0100] Specifically, by comparing the differences in gene expression levels under different conditions (such as normal tissue and cancer tissue), genes with significant changes (increased or decreased) in expression levels are screened out, and these genes form gene difference data.

[0101] It should be noted that the association test is a process of testing based on the standard of significant association with cancer disease after preliminary difference screening.

[0102] It should be noted that the hierarchical clustering algorithm is a cluster analysis method whose basic concept is to form a hierarchical clustering result by gradually merging or splitting samples or data points. The hierarchical clustering algorithm can further aggregate the data with strong correlations with cancer disease status in the genetic difference data after association testing, constructing a hierarchical clustering tree structure, intuitively displaying the hierarchical relationship between samples and genes in the genetic difference data after association testing, and obtaining hierarchical association results.

[0103] In this embodiment, the transcriptome sequencing data can first be subjected to preliminary differential screening based on gene expression levels (e.g., comparing gene expression levels in diseased and healthy samples) to obtain genetic differential data. Following this preliminary differential screening, an association test is then performed using criteria for significant association with cancer to obtain genetic differential data after the association test. Finally, a hierarchical clustering algorithm is used to perform feature screening on the genetic differential data after the association test. Data from the genetic differential data after the association test that are strongly associated with the cancer disease state are further hierarchically aggregated to construct a hierarchical clustering tree structure. This intuitively displays the hierarchical relationships between samples and genes in the genetic differential data after the association test, resulting in hierarchical association results to determine biomarker expression data associated with significant genes. The hierarchical clustering algorithm can thus better understand the distribution and structural characteristics of the hierarchical relationships between cancer samples and genes, thereby improving the accuracy of biomarker expression data.

[0104] In another feasible embodiment, step S30 of this embodiment may include the steps of: determining the significant interactions of genes in the cancer disease state and the normal state based on the posterior distribution and the gene interaction network; and determining the significant interacting genes in the cancer disease biological sample based on the significant interactions.

[0105] It should be noted that significant interaction refers to the obvious correlation between genes in the disease state of cancer and the normal state.

[0106] Specifically, the goal is to determine how the relationships between genes differ in cancer compared to normal conditions. For example, under normal conditions, certain genes may collaborate or constrain each other in a relatively stable manner, but in cancer, their interactions may become abnormally active or inhibited. This significant change in interaction may play a key role in the development of cancer. By identifying such significant interactions, we can gain a deeper understanding of the molecular mechanisms of cancer and provide an important basis for biomarker screening.

[0107] In this embodiment, the posterior distribution of the association between biomarkers and cancer diseases, combined with the intergenic interaction network, can be used to determine the significant interactions between genes in both the cancer disease state and the normal state. This significant interaction can then be used to identify genes with particularly prominent interactions that play a key role in the development and progression of cancer. By identifying significant interactions, we can gain a deeper understanding of the molecular mechanisms of cancer and provide an important basis for biomarker screening.

[0108] This embodiment can perform transcriptional amplification on biological samples related to cancer (such as cancer cells and cancer tissue). Sequencing analysis is then performed on the amplified samples to determine the sequence information of the RNA molecules in the samples. This is then screened using functional genomics and clinical genomics as benchmarks to obtain transcriptome sequencing data. This, combined with the standards of functional genomics and clinical genomics, can provide a better understanding of the molecular mechanisms and individual differences in cancer, providing important support for the screening of cancer biomarkers for medical treatment and enabling more accurate and effective biomarker screening. Furthermore, based on the posterior distribution of the association between biomarkers and cancer, combined with the gene interaction network, significant interactions between genes in both the cancerous state and the normal state can be determined. This significant interaction can then be used to identify genes with particularly prominent interactions that play a key role in the development and progression of cancer. By identifying significant interactions, a deeper understanding of the molecular mechanisms of cancer can be achieved, providing an important basis for biomarker screening.

[0109] Example 3: Based on the first and second embodiments of this application, in the third embodiment of this application, the same or similar contents as those in the first and second embodiments above can be referred to above and will not be described in detail later. Figure 3 , Figure 3 This is a schematic diagram of the process of Example 3 of the genomics-based biomarker screening method of this application.

[0110] In this example, step S20 includes steps S21 and S22:

[0111] Step S21: Based on the biomarker expression data of the significant genes, a logistic regression model is used to describe the association between the biomarkers and cancer diseases.

[0112] Step S22: Using a Bayesian method, unnormalized combinations of the associations between the biomarkers and cancer diseases are combined to obtain a posterior distribution of the associations between the biomarkers and cancer diseases.

[0113] It should be noted that the Bayesian approach is a statistical inference method based on Bayes' theorem. Bayes' theorem describes how to update estimates of event probabilities based on new observations, given certain prior information. Specifically, the posterior distribution of the association between biomarkers and cancer can be calculated using prior information about the association between biomarkers and cancer.

[0114] Specifically, for ease of understanding, the specific processing process is as follows:

[0115] 1. First, build a model for binary outcomes.

[0116] set up Indicates research indicators (i.e., biomarker expression data) that facilitate aggregation. A single , represents a binary disease outcome, represents a biomarker exposure measurement from a reference laboratory (reference measurement), represents biomarker measurements from study-specific local laboratories (local measurements), A vector representing other covariates. It is assumed that for all samples in the study, the local measurement of the biomarker is available, but only a portion of the samples have reference measurements For research A single , If the reference measurement value is available, otherwise .

[0117] Assuming reference measurement value The distribution of The distribution of was heterogeneous across studies. In addition, it can be assumed that the results and local measurements are conditionally independent, given a reference measurement , which means the probability:

[0118] ;

[0119] For convenience, it was assumed that the covariates had no effect on study-specific measurement bias, i.e. In addition, the likelihood function containing all samples can be expressed as follows:

[0120] ;

[0121] in,

[0122] ;

[0123] For convenience, Represents the set of all parameters, which are introduced below. It is research The total number of samples. Further referring to the assumption of conditional independence 1, we can derive the following equation:

[0124] ;

[0125] It can be observed that the likelihood function consists of three components, namely, the biomarker-disease association , reference-local measurement correlation , and the reference prior .

[0126] A logistic regression model with a random intercept term was used to describe the association between biomarkers and disease as follows:

[0127] ;

[0128] in, To study the specific intercept, is the inverse of the logit function. Our main purpose is to estimate , which is the logarithm of the odds ratio (OR), describes the relationship between the biomarker and the disease.

[0129] The model used to describe the reference-local measurement relationship is called the calibration model. and local measurements There is a linear correlation between and All obey the normal distribution, then:

[0130] ;

[0131] and,

[0132] ;

[0133] , , , , , , and Can be seen as Now, we can fully express the likelihood function.

[0134] Since the likelihood function involves the unknown , the use of maximum likelihood estimation introduces challenging integral calculations. In addition, the hierarchical structure of the biological specimen is not incorporated into the estimation of the parameters. Therefore, the Bayesian method is selected to obtain The posterior distribution of . In the Bayesian method, Treated as a latent variable, it can be transformed into an estimable quantity by constructing an appropriate prior distribution. Similarly, the hierarchical structure can also be reflected in The pre-setting of parameters such as , improves the efficiency of estimation.

[0135] 2. Then perform prior distribution.

[0136] Simply put, assuming and The joint prior distribution of is factorized and independent of ,Right now:

[0137] ;

[0138] Since the reference measurements are assumed to be consistent across studies, and should be identically distributed. Therefore, and are not independent. Based on the assumption of normal distribution, the prior distribution of these three variables can be expressed as:

[0139] ;

[0140] and It can be set according to the reference measurement value in the actual scene, or a conjugate prior (normal and inverse gamma distribution) can be used. In our study, a non-informative prior is placed, specifically:

[0141] ;

[0142] for and , it is recommended to use informative priors or weakly informative priors to preserve power and avoid severe instability. For example, using a standard normal distribution, as follows:

[0143] ;

[0144] in, and is the number of covariates.

[0145] For specific study parameters , , and , assuming that any vectors between them are sampled from a common distribution. This corresponds to a "random effects" model. However, in the Bayesian framework, there is no need to assume that the trials were sampled from a superpopulation. Instead, it is replaced by the qualitative assumption of "exchangeability", which means that there is no reason to believe that the studies are systematically different. Here, we continue with the simplest assumption that the parameters of each study are independent samples from a superpopulation distribution controlled by some unknown superparameter. For and , the population distribution is normal distribution, as follows:

[0146]

[0147] And for , which is an inverse gamma distribution, is as follows:

[0148] ;

[0149] in, 、 、 、 、 、 、 and is the corresponding hyperparameter, which is also considered to be part of the unknown parameters. express Follows the inverse gamma distribution, as follows:

[0150] ;

[0151] and and Similar to the prior distribution of , we propose to assign an appropriate weakly informative prior to these hyperparameters. and A standard normal distribution can be used, as follows:

[0152] ;

[0153] for 、 、 、 , you can use the semi-Cauchy distribution version as follows:

[0154] .

[0155] 3. Finally, perform parameter estimation.

[0156] According to the Bayesian formula, the unnormalized joint posterior distribution of the unknown parameters and the reference measurement value can be obtained as follows:

[0157] ;

[0158] Using Markov Chain Monte Carlo (MCMC) methods, such as the Nut (without a U-shaped sampler) 6, it is possible to sample from this distribution. Based on the MCMC sampler, the sample mean is calculated as a point estimate of the parameter, and the 95% highest posterior density interval is calculated as an interval estimate. This results in a posterior distribution for the association between the biomarker and the cancer disease.

[0159] This example aggregates biomarker data from multiple study sources using Bayesian techniques, treating reference measurements from unreanalyzed biospecimens as unobservable latent variables. A two-level study-biospecimen model is developed to describe the relationship between reference measurements, local measurements, and outcomes. This model estimates the posterior distribution of biomarker-disease associations, providing a powerful framework for integrating biomarker data for cancer diseases.

[0160] Reference Figure 4 The genomics-based biomarker screening method of the present application is implemented by a genomics-based biomarker screening system. Figure 4 This is a schematic diagram of the functional modules of the genomics-based biomarker screening system according to an embodiment of the present application. The system includes:

[0161] The sequencing data module 10 is used to determine biomarker expression data related to significant genes based on transcriptome sequencing data of cancer disease biological samples, wherein the transcriptome sequencing data is obtained based on functional genomics and clinical genomics.

[0162] It should be noted that cancer disease biological samples are various biological materials related to cancer.

[0163] For example, biological samples of cancer diseases may include tissue specimens (such as tumor tissue, adjacent cancer tissue), blood (such as whole blood, plasma, serum), body fluids (such as urine, cerebrospinal fluid, pleural effusion, etc.), and cells (such as circulating tumor cells, etc.) obtained from cancer patients.

[0164] It should be noted that transcriptome sequencing data is data obtained by sequencing RNA transcribed from cancer disease biological samples under specific conditions.

[0165] For example, the transcription process may include messenger RNA (mRNA), ribosomal RNA (rRNA), transfer RNA (tRNA), and other non-coding RNAs. Transcriptome sequencing can comprehensively and quantitatively analyze gene expression in cancer disease biological samples.

[0166] It should be noted that biomarker expression data refers to gene data in transcriptome sequencing data that are closely related to cancer diseases and can significantly express biomarkers.

[0167] It should be noted that functional genomics is the study of gene functions, including gene expression regulation, the function of gene products, and the interaction between genes.

[0168] Clinical genomics is the application of genomics technology and research results to clinical practice, involving the use of genomic information to diagnose diseases, predict disease prognosis, etc.

[0169] Constructing transcriptome sequencing data based on the standards of functional genomics and clinical genomics can better understand the molecular mechanisms and individual differences of cancer diseases, provide important support for the screening of biomarkers for cancer diseases and their application in medical treatment, and achieve more accurate and effective screening of biomarkers.

[0170] In one embodiment, transcription amplification can be performed on biological samples of cancer diseases first, and then transcriptome sequencing data can be obtained based on the standards of functional genomics and clinical genomics. The screening system can screen the associated characteristics of cancer diseases based on these data to obtain biomarker expression data of significant genes closely related to cancer diseases.

[0171] The posterior distribution module 20 is configured to estimate the posterior distribution of the association between the biomarker and the cancer disease by using a logistic regression model based on the biomarker expression data of the significant genes.

[0172] It should be noted that the logistic regression model is a statistical analysis model used to solve binary (or multi-classification) problems. The logistic regression model can be used to describe the association between biomarkers and cancer diseases and determine the association between biomarkers and cancer diseases.

[0173] It should be noted that biomarkers are biochemical molecules that can identify changes or potential changes in systems, organs, tissues, cells, subcellular structures, and genetic structures. Biomarkers can help more accurately diagnose cancer.

[0174] It should be noted that the posterior distribution is a probability distribution of the association between biomarkers and cancer. The posterior distribution can provide a deeper understanding of the role of biomarkers in cancer development and its uncertainty.

[0175] In one embodiment, biomarker expression data can be used as input variables and combined with relevant information about cancer diseases. The degree of association between biomarkers and cancer diseases can be probabilistically described using a logistic regression model to obtain the posterior distribution of the association between biomarkers and cancer diseases.

[0176] The interaction module 30 is used to determine the significant interacting genes in the cancer disease biological sample according to the posterior distribution and the gene interaction network.

[0177] It should be noted that the intergenic interaction network is a complex network structure formed by various interaction relationships between numerous genes.

[0178] In organisms, genes do not exist and function in isolation. The expression and function of a gene may be influenced by other genes, with interactions occurring in various ways, including synergistic and antagonistic. These interactions are intertwined, forming a vast network.

[0179] It should be noted that significantly interacting genes are genes with particularly prominent interaction relationships in biological samples of cancer diseases and may play a key role in the occurrence and development of cancer.

[0180] For example, in cancer samples, the interactions of certain genes may change significantly, such as increasing or decreasing their synergistic or antagonistic effects in the cancer state. These genes can be considered to have significant interactions in the cancer. Gene interaction networks can help us better understand the interactions between genes and reveal the associations and synergistic patterns between genes that play a key role in the occurrence and development of cancer.

[0181] In one embodiment, the posterior distribution of the association between biomarkers and cancer diseases, combined with the gene interaction network, can identify genes with particularly prominent interactions that play a key role in the development and progression of cancer. This provides an important basis for accurately screening biomarkers related to cancer diseases and facilitates the discovery of new biomarkers.

[0182] The marker screening module 40 is used to screen and obtain a list of biomarkers related to cancer diseases based on the significantly interacting genes.

[0183] It should be noted that the biomarker list is a list of biomarkers related to cancer. Through the biomarker list, we can better understand the relationship between biomarkers and cancer diseases, which helps in disease diagnosis.

[0184] In one embodiment, the significantly interacting genes obtained above can be arranged in order of interaction degree or classified according to different physiological functions, and a list of biomarkers related to cancer diseases can be obtained by screening.

[0185] This embodiment provides a genomics-based biomarker screening system. First, transcriptome amplification is performed on cancer biological samples. Transcriptome sequencing data is then obtained based on functional genomics and clinical genomics standards. The screening system then uses this data to screen for cancer-associated features, obtaining biomarker expression data for significant genes closely associated with cancer. The biomarker expression data can then be used as input variables and combined with relevant information about the cancer. A logistic regression model is then used to probabilistically describe the degree of association between the biomarker and the cancer, yielding a posterior distribution of the biomarker-cancer association. This posterior distribution of the biomarker-cancer association can then be combined with the gene interaction network to identify genes with particularly prominent interactions that play a key role in the development and progression of cancer. This provides an important basis for accurately screening biomarkers associated with cancer and facilitates the discovery of new biomarkers. Finally, the significantly interacting genes obtained above can be sorted by interaction level or by physiological function to obtain a list of biomarkers associated with cancer. Since this embodiment uses transcriptome sequencing data from cancer disease biological samples to screen and obtain the posterior distribution of biomarkers associated with cancer diseases and significantly interacting genes, it avoids the situation in which missed detections due to low biomarker specificity in the traditional biomarker screening process are avoided. As a result, a list of biomarkers with high correlation with cancer diseases can be obtained, thereby improving the sensitivity of the biomarker screening process.

[0186] Based on the first embodiment of the system of the present application, a second embodiment of the system of the present application is proposed. In the second embodiment of the present application, the same or similar contents as those of the above-mentioned embodiment 1 can be referred to the above introduction and will not be repeated hereafter.

[0187] The sequencing data module 10 in this example is also used to obtain various types of cancer disease biological samples.

[0188] The sequencing data module 10 is further used to perform transcription amplification on the cancer disease biological sample, and to sequence and obtain transcriptome sequencing data based on functional genomics and clinical genomics.

[0189] It should be noted that transcription amplification is the process of increasing the number of RNA transcripts in a sample through specific technical means. The samples after transcription amplification are then sequenced and analyzed. This sequencing process is to determine the sequence information of these RNA molecules. The sequence information is then screened based on functional genomics and clinical genomics to obtain transcriptome sequencing data.

[0190] The sequencing data module 10 is further used to perform cluster screening on the transcriptome sequencing data to obtain biomarker expression data related to significant genes.

[0191] In this embodiment, biological samples related to cancer (such as cancer cells and cancer tissue) can be subjected to transcriptional amplification. Sequencing analysis is then performed on the amplified samples to determine the sequence information of the RNA molecules in the samples. Subsequently, screening is performed based on functional genomics and clinical genomics to obtain transcriptome sequencing data. By applying the standards of functional genomics and clinical genomics, a better understanding of the molecular mechanisms and individual differences in cancer can be achieved, providing important support for the screening of cancer biomarkers for medical treatment, enabling more accurate and effective biomarker screening.

[0192] In another feasible embodiment, the sequencing data module 10 described in this embodiment is also used to perform preliminary screening of the transcriptome sequencing data according to gene expression levels to obtain gene difference data; the sequencing data module 10 is also used to perform association testing based on the gene difference data to obtain gene difference data after the association testing; the sequencing data module 10 is also used to perform feature screening on the gene difference data after the association testing through a hierarchical clustering algorithm to obtain hierarchical association results between genes and cancer diseases; the sequencing data module 10 is also used to determine biomarker expression data associated with significant genes based on the hierarchical association results.

[0193] It should be noted that genetic difference data are obtained by preliminary difference screening based on gene expression levels (for example, comparing the gene expression levels of diseased and healthy individuals).

[0194] Specifically, by comparing the differences in gene expression levels under different conditions (such as normal tissue and cancer tissue), genes with significant changes (increased or decreased) in expression levels are screened out, and these genes form gene difference data.

[0195] It should be noted that the association test is a process of testing based on the standard of significant association with cancer disease after preliminary difference screening.

[0196] It should be noted that the hierarchical clustering algorithm is a cluster analysis method whose basic concept is to form a hierarchical clustering result by gradually merging or splitting samples or data points. The hierarchical clustering algorithm can further aggregate the data with strong correlations with cancer disease status in the genetic difference data after association testing, constructing a hierarchical clustering tree structure, intuitively displaying the hierarchical relationship between samples and genes in the genetic difference data after association testing, and obtaining hierarchical association results.

[0197] In this embodiment, the transcriptome sequencing data can first be subjected to preliminary differential screening based on gene expression levels (e.g., comparing gene expression levels in diseased and healthy samples) to obtain genetic differential data. Following this preliminary differential screening, an association test is then performed using criteria for significant association with cancer to obtain genetic differential data after the association test. Finally, a hierarchical clustering algorithm is used to perform feature screening on the genetic differential data after the association test. Data from the genetic differential data after the association test that are strongly associated with the cancer disease state are further hierarchically aggregated to construct a hierarchical clustering tree structure. This intuitively displays the hierarchical relationships between samples and genes in the genetic differential data after the association test, resulting in hierarchical association results to determine biomarker expression data associated with significant genes. The hierarchical clustering algorithm can thus better understand the distribution and structural characteristics of the hierarchical relationships between cancer samples and genes, thereby improving the accuracy of biomarker expression data.

[0198] The genomics-based biomarker screening system provided by this application adopts the genomics-based biomarker screening method in the above-mentioned embodiment, which can solve the shortcomings of low biomarker specificity and low screening sensitivity in the traditional biomarker screening process, which easily leads to missed detection and causes missed diagnosis or delayed diagnosis of patients. Compared with the prior art, the beneficial effects of the genomics-based biomarker screening system provided by this application are the same as the beneficial effects of the genomics-based biomarker screening method provided by the above-mentioned embodiment, and the other technical features of the genomics-based biomarker screening system are the same as the features disclosed in the above-mentioned embodiment method, which will not be repeated here.

[0199] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the genomics-based biomarker screening method and system of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.

[0200] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A genomics-based biomarker screening method, characterized in that: The method includes: Determining biomarker expression data associated with significant genes based on transcriptome sequencing data of cancer disease biological samples, wherein the transcriptome sequencing data is constructed based on functional genomics and clinical genomics; estimating the posterior distribution of the association between the biomarker and the cancer disease by a logistic regression model based on the biomarker expression data of the significant genes; determining significantly interacting genes in a cancer disease biological sample based on the posterior distribution and the gene-gene interaction network; Based on the significantly interacting genes, a list of biomarkers related to cancer diseases is obtained by screening; The step of determining biomarker expression data associated with significant genes based on transcriptome sequencing data of cancer disease biological samples comprises: obtaining multiple types of cancer disease biological samples; performing transcriptional amplification on the cancer disease biological samples; performing sequencing analysis on the cancer disease biological samples after transcriptional amplification to determine sequence information of RNA molecules in the cancer disease biological samples, and screening based on functional genomics and clinical genomics to obtain transcriptome sequencing data; performing a preliminary screening of the transcriptome sequencing data based on gene expression levels to obtain gene difference data; performing an association test on the gene difference data based on a criterion of significant association with cancer to obtain gene difference data after the association test; performing feature screening on the gene difference data after the association test using a hierarchical clustering algorithm, hierarchically aggregating data that is strongly associated with the cancer disease state in the gene difference data after the association test, constructing a hierarchical clustering tree structure to display the hierarchical relationship between samples and genes in the gene difference data after the association test, and obtaining hierarchical association results between genes and cancer diseases; and determining biomarker expression data associated with the significant genes based on the hierarchical association results; The step of estimating the posterior distribution of the association between the biomarker and the cancer disease using a logistic regression model based on the biomarker expression data of the significant genes comprises: establishing a binary outcome model based on the biomarker expression data of the significant genes: ; in, Indicates research indicators that promote aggregation, namely biomarker expression data; for research A single , represents a binary disease outcome; represents the biomarker exposure measurement from a reference laboratory, i.e., the reference measurement; represents biomarker measurements from study-specific local laboratories, i.e., local measurements, A vector representing other covariates; Logistic regression models were used to describe the association between biomarkers and cancer disease: ; in, is the study-specific intercept; is the inverse of the logit function; is the logarithm of the odds ratio, which is used to describe the relationship between the biomarker and the disease; yes Elements of Represents the set of all parameters; By using a Bayesian approach, we combined the unnormalized associations of biomarkers with cancer diseases, treated the reference measurements of biospecimens that were not reanalyzed as unobservable latent variables, and established a two-level study-biospecimen model to describe the relationship between reference measurements, local measurements, and outcomes to estimate the posterior distribution of the biomarker-cancer disease association: ; in, ; in, It is research The total number of samples for the study A single , If the reference measurement value is available, otherwise .

2. The method according to claim 1, wherein The step of determining the significant interacting genes in the cancer disease biological sample according to the posterior distribution and the gene interaction network comprises: Determining significant interactions of genes in a cancer disease state and a normal state based on the posterior distribution and the gene-gene interaction network; According to the significant interactions, significantly interacting genes in the cancer disease biological sample are determined.

3. The method according to claim 1, wherein The step of screening and obtaining a list of biomarkers associated with cancer diseases based on the significantly interacting genes comprises: Perform feature importance evaluation on the significantly interacting genes to obtain importance scores; Sorting the importance scores to obtain a score list; The significantly interacting genes are screened based on the score list to obtain a list of biomarkers associated with cancer diseases.

4. The method according to claim 3, wherein The step of performing feature importance assessment on the significantly interacting genes to obtain importance scores comprises: Initialize the gradient boosting decision tree model; Training the gradient boosting decision tree model using a cancer gene database; The feature importance of the significantly interacting genes was evaluated based on the trained gradient boosting decision tree model to obtain importance scores.

5. A genomics-based biomarker screening system, characterized in that: The genomics-based biomarker screening system performs the genomics-based biomarker screening method according to claim 1, and the system comprises: A sequencing data module is used to determine biomarker expression data associated with significant genes based on transcriptome sequencing data of cancer disease biological samples, wherein the transcriptome sequencing data is obtained based on functional genomics and clinical genomics; a posterior distribution module for estimating the posterior distribution of the association between the biomarker and the cancer disease through a logistic regression model based on the biomarker expression data of the significant genes; An interaction module, for determining significant interacting genes in a cancer disease biological sample based on the posterior distribution and the gene-gene interaction network; The marker screening module is used to screen and obtain a list of biomarkers related to cancer diseases based on the significantly interacting genes.

Citation Information

Patent Citations

  • Biomarker combination and screening method thereof

    CN116106401A

  • Screening method and application of ovarian cancer biomarker

    CN117625793A