Molecular marker screening method and device based on sequencing, equipment and storage medium
By using sequencing-based molecular biomarker screening methods, utilizing gene variant sites and expression profiles, weighted gene co-expression networks, and miRNA-gene interaction networks, combined with machine learning algorithms, highly specific and sensitive biomarkers for lung adenocarcinoma were screened out. This solved the problem of insufficient sensitivity of traditional biomarkers and enabled more accurate early diagnosis and treatment.
Patent Information
- Application Number
- CN202511060126.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-07-31
- Filing Date
- 2025-07-30
- Publication Date
- 2026-02-03
AI Technical Summary
In the early diagnosis of lung cancer, the sensitivity and specificity of traditional biomarkers are insufficient, resulting in poor diagnosis and treatment outcomes for lung adenocarcinoma. Furthermore, current sequencing technologies cannot fully assess the potential efficacy of biomarkers.
After obtaining sample sequencing data and performing quality control processing, gene mutation sites and target gene expression profiles were detected. Combined with weighted gene co-expression network analysis and miRNA-gene molecular interaction network, candidate biomarkers were screened, and machine learning algorithms were used to calculate importance scores to screen out target biomarkers with higher specificity and sensitivity.
It improves the specificity and sensitivity of molecular markers for lung adenocarcinoma, enabling more accurate early diagnosis, providing multidimensional assessment, and enhancing the effectiveness of treatment.
Smart Images

Figure CN121459918A_ABST
Abstract
Description
[0001] This invention claims priority to Chinese Patent Application No. 2024110419532, filed with the Chinese Patent Office on July 31, 2024, entitled "Method, Apparatus, Device and Storage Medium for Drug Recommendation Based on Sequencing Data", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This invention belongs to the field of bioinformatics technology, specifically relating to a sequencing-based molecular marker screening method, apparatus, device, and storage medium. Background Technology
[0003] Lung adenocarcinoma is one of the most common histological types of lung cancer, accounting for approximately 40% of all lung cancers, and its incidence and mortality rates are high worldwide. Significant progress has been made in the treatment of lung adenocarcinoma, with treatment methods including surgery, radiotherapy, chemotherapy, targeted therapy, and immunotherapy.
[0004] Currently, early detection and diagnosis of lung cancer rely on imaging techniques such as X-rays and CT scans. Low-dose spiral CT (LDCT), in particular, remains the primary screening method for high-risk groups. Compared to X-rays, LDCT can significantly improve the detection rate of lung cancer, but its high false-positive rate and radiation exposure are insurmountable drawbacks. Pathological diagnosis is the gold standard for diagnosing non-small cell lung cancer (NSCLC), but tissue biopsy procedures have significant limitations and risks, and in many cases, obtaining tissue samples is extremely difficult.
[0005] The effectiveness of traditional biomarkers is largely limited by the spatial and temporal heterogeneity of tumors. The sensitivity and / or specificity of hematologic tumor markers such as carcinoembryonic antigen (CEA), neuron-specific enolase (NSE), and the soluble fragment of cytokeratin 19 (Cyfra21-1) are not satisfactory. This has resulted in a persistently low 5-year survival rate for lung adenocarcinoma despite the discovery and application of many molecular markers in its diagnosis and treatment. Therefore, the accurate discovery and screening of molecular markers for early-stage lung cancer is crucial for cancer diagnosis and treatment.
[0006] Traditional biomarker screening relies on comparing the differential expression of biomarkers. However, this method suffers from the problem of too many differentially expressed biomarkers, resulting in a large screening workload and low specificity and sensitivity of the selected biomarkers. Furthermore, some biomarkers may not accurately indicate the disease in actual diagnosis. In addition, although some studies have used sequencing technology to explore biomarkers for early lung cancer diagnosis, these studies are mostly limited to single-molecule level explorations, thus failing to comprehensively evaluate the potential efficacy of these biomarkers, leading to low specificity and sensitivity of the selected biomarkers. Summary of the Invention
[0007] The purpose of this invention is to provide a sequencing-based molecular biomarker screening method, apparatus, device, and storage medium, which can comprehensively evaluate the potential efficacy of biomarkers from multiple dimensions, thereby improving the specificity and sensitivity of the screened biomarkers.
[0008] The first aspect of this invention discloses a sequencing-based molecular biomarker screening method, comprising:
[0009] Obtain sample sequencing data and perform quality control processing on the sample sequencing data to obtain clean sequencing data;
[0010] Based on the clean sequencing data, gene mutation sites were detected and obtained;
[0011] Based on the clean sequencing data, the expression profile of the target gene was obtained by analysis;
[0012] Weighted gene co-expression network analysis was performed on the clean sequencing data to obtain the analysis results;
[0013] Co-expression analysis was performed on the clean sequencing data to construct a miRNA-gene molecular interaction network. Topological analysis was then performed on the miRNA-gene molecular interaction network to identify key nodes.
[0014] Candidate biomarkers were screened based on the gene mutation sites, the target gene expression map, the analysis results, and the key nodes of the miRNA-gene molecular interaction network.
[0015] Machine learning algorithms are used to calculate the importance scores of all candidate biomarkers;
[0016] Sort all candidate markers by importance score from highest to lowest, and select the top-ranked candidate markers and their combinations as target markers.
[0017] In some embodiments, obtaining sample sequencing data includes:
[0018] Cancerous tissue from patients clinically diagnosed with cancer and their matching adjacent non-cancer tissues were obtained as paired samples; the RNA of the paired samples was subjected to miRNA sequencing and RNA-seq sequencing to obtain initial miRNA sequencing data and initial gene sequencing data, which were then combined to obtain sample sequencing data.
[0019] In some embodiments, detecting gene mutation sites based on the clean sequencing data includes:
[0020] The clean sequencing data is aligned to a human reference genome to obtain alignment results, and gene mutation sites are detected based on the alignment results.
[0021] In some embodiments, the clean sequencing data includes clean gene sequencing data and miRNA sequencing data; based on the clean sequencing data, an expression profile of the target gene is obtained through analysis, including:
[0022] Clean gene sequencing data is aligned to a human reference genome to obtain the first gene data. After quantifying the first gene data to obtain gene expression values, a comparison group is set up according to the normal group and lung adenocarcinoma, and differential analysis is performed on the comparison group to obtain the first gene expression map.
[0023] Clean miRNA sequencing data was aligned to the human reference genome to obtain second gene data. After quantifying the second gene data to obtain miRNA expression values, comparison groups were set up according to the normal group and lung adenocarcinoma, and differential analysis was performed on the comparison groups to obtain the second gene expression profile.
[0024] Based on the first gene expression map and the second gene expression map, the target gene expression map is obtained.
[0025] The importance scores of all molecular markers were calculated using machine learning algorithms, including:
[0026] A biomarker evaluation model is constructed based on machine learning algorithms such as random forest and stepwise regression. The biomarker evaluation model is then trained to obtain a trained biomarker evaluation model.
[0027] The gene mutation sites and target gene expression maps are input into the trained biomarker evaluation model to obtain the importance score of each candidate biomarker.
[0028] In some embodiments, training the biomarker evaluation model to obtain a trained biomarker evaluation model includes:
[0029] If N training sets are obtained in advance, then n samples with replacement are randomly selected from the N training sets for a single decision tree as training samples for that decision tree.
[0030] Let the number of input features of the training samples be M, where m is less than M. Then, when splitting at each node of each decision tree, we randomly select m input features from the M input features, and then select the best one from the m input features to split, until all training samples of that node belong to the same class.
[0031] In some embodiments, after obtaining the target marker, the method further includes:
[0032] Obtain the expression information of the target biomarker in pan-cancer; the process of obtaining the expression information of the target biomarker in pan-cancer includes: obtaining the expression information of the target biomarker in the GEPIA and dbDEMC databases to verify the target biomarker;
[0033] The target biomarker was validated using its expression information in pan-cancer.
[0034] A second aspect of this invention discloses a sequencing-based molecular marker screening device, comprising:
[0035] Acquisition unit, used to acquire sample sequencing data;
[0036] The quality control unit is used to perform quality control processing on sample sequencing data to obtain clean sequencing data;
[0037] A mutation detection unit is used to detect and obtain gene mutation sites based on the clean sequencing data;
[0038] The first analysis unit is used to analyze and obtain the target gene expression map based on the clean sequencing data;
[0039] The second analysis unit is used to perform weighted gene co-expression network analysis on the clean sequencing data to obtain analysis results.
[0040] The third analysis unit is used to perform co-expression analysis on the clean sequencing data, construct a miRNA-gene molecular interaction network, perform topological analysis on the miRNA-gene molecular interaction network, and identify key nodes.
[0041] The screening unit is used to screen candidate biomarkers based on the gene mutation sites, the target gene expression map, the analysis results, and the key nodes of the miRNA-gene molecular interaction network.
[0042] The computational unit is used to calculate the importance scores of all candidate markers using machine learning algorithms;
[0043] The determining unit is used to sort all candidate markers from largest to smallest according to their importance scores, and to obtain a specified number of candidate markers and their combinations that rank highest as target markers.
[0044] A third aspect of the present invention discloses an electronic device, including a memory storing executable program code and a processor coupled to the memory; the processor calls the executable program code stored in the memory to execute the sequencing-based molecular marker screening method disclosed in the first aspect.
[0045] A fourth aspect of the present invention discloses a computer-readable storage medium storing a computer program, wherein the computer program causes a computer to perform the sequencing-based molecular marker screening method disclosed in the first aspect.
[0046] The beneficial effects of this invention are as follows: Clean sequencing data is obtained by acquiring sample sequencing data and performing quality control processing; gene mutation sites are detected based on the clean sequencing data; target gene expression maps are analyzed based on the clean sequencing data; weighted gene co-expression network analysis is performed on the clean sequencing data to obtain analysis results; co-expression analysis is performed on the clean sequencing data to construct a miRNA-gene molecular interaction network; topological analysis is performed on the miRNA-gene molecular interaction network to identify key nodes; candidate biomarkers are screened based on gene mutation sites, target gene expression maps, analysis results, and key nodes of the miRNA-gene molecular interaction network; and the importance score of each candidate biomarker is calculated using machine learning algorithms; a specified number of candidate biomarkers and their combinations, ranked from largest to smallest, are obtained as target biomarkers. This allows for a comprehensive multi-dimensional evaluation of the potential efficacy of biomarkers, thereby improving the specificity and sensitivity of the screened biomarkers. Attached Figure Description
[0047] The accompanying drawings illustrate specific examples of the technical solutions described in this invention and, together with the detailed embodiments, form part of the specification, serving to explain the technical solutions, principles, and effects of this invention.
[0048] Unless otherwise specified or defined, the same reference numerals in different figures represent the same or similar technical features, and different reference numerals may be used to represent the same or similar technical features.
[0049] Figure 1 This is a flowchart of a sequencing-based molecular biomarker screening method disclosed in an embodiment of the present invention;
[0050] Figure 2 This is an overall flowchart of a sequencing-based molecular marker screening method disclosed in an embodiment of the present invention;
[0051] Figure 3 These are some of the gene mutation sites disclosed in the embodiments of this invention;
[0052] Figure 4 This is the first gene expression map disclosed in the embodiments of the present invention;
[0053] Figure 5 This is the second gene expression map disclosed in the embodiments of the present invention;
[0054] Figure 6 This is a diagram of the gene clustering module disclosed in an embodiment of the present invention;
[0055] Figure 7 This is a diagram of the miRNA clustering module disclosed in an embodiment of the present invention;
[0056] Figure 8 This is a correlation diagram between the orange gene module disclosed in the embodiments of the present invention and clinical applications;
[0057] Figure 9 This is a clinical correlation diagram of the blue gene module disclosed in the embodiments of the present invention;
[0058] Figure 10 This is a molecular interaction network diagram disclosed in the embodiments of the present invention;
[0059] Figure 11 This is the ROC curve of hsa-miR-133a-3p of the miRNA disclosed in the embodiments of the present invention;
[0060] Figure 12 This is the ROC curve of hsa-miR-133b of the miRNA disclosed in the embodiments of the present invention;
[0061] Figure 13 This is the ROC curve of hsa-miR-139-3p of the miRNA disclosed in the embodiments of the present invention;
[0062] Figure 14 This is the ROC curve of hsa-miR-139-5p of the miRNA disclosed in the embodiments of the present invention;
[0063] Figure 15 This is the ROC curve of hsa-miR-21-3p of the miRNA disclosed in the embodiments of the present invention;
[0064] Figure 16 This is the ROC curve of hsa-miR-30c-2-3p of the miRNA disclosed in the embodiments of the present invention;
[0065] Figure 17 This is the ROC curve of CA4 in mRNA disclosed in the embodiments of the present invention;
[0066] Figure 18 This is the ROC curve of CCT3 of mRNA disclosed in the embodiments of the present invention;
[0067] Figure 19 This is the ROC curve of EPCAM of mRNA disclosed in the embodiments of the present invention;
[0068] Figure 20 This is the ROC curve of LINC00092 of lncRNA disclosed in the embodiments of the present invention;
[0069] Figure 21 This is the ROC curve of MYO16-AS1 of lncRNA disclosed in the embodiments of the present invention;
[0070] Figure 22 This is the ROC curve of TRHDE-AS1 of lncRNA disclosed in the embodiments of the present invention;
[0071] Figure 23 This is the ROC curve of the first target marker combination disclosed in the embodiments of the present invention;
[0072] Figure 24 This is the ROC curve diagram of the second target marker combination disclosed in the embodiments of the present invention;
[0073] Figure 25 This is the ROC curve of the third target marker combination disclosed in the embodiments of the present invention;
[0074] Figure 26 This is the ROC curve of the fourth target marker combination disclosed in the embodiments of the present invention;
[0075] Figure 27 This is the ROC curve of the fifth target marker combination disclosed in the embodiments of the present invention;
[0076] Figure 28 This is the first thermal image disclosed in the embodiments of the present invention;
[0077] Figure 29 This is the second heat map disclosed in the embodiments of the present invention;
[0078] Figure 30 This is the ROC curve of CA4 in the TCGA database gene disclosed in this embodiment of the invention;
[0079] Figure 31 This is the ROC curve of CCT3 in the TCGA database gene disclosed in the embodiments of the present invention;
[0080] Figure 32 This is the ROC curve of EPCAM in the TCGA database gene disclosed in the embodiments of the present invention;
[0081] Figure 33 This is the ROC curve of the first target biomarker combination in the TCGA database genes disclosed in this embodiment of the invention;
[0082] Figure 34 This is the ROC curve of the second target biomarker combination in the TCGA database genes disclosed in this embodiment of the invention;
[0083] Figure 35 This is the ROC curve of the third target biomarker combination in the TCGA database genes disclosed in the embodiments of the present invention;
[0084] Figure 36 This is the ROC curve of the fourth target biomarker combination in the TCGA database genes disclosed in this embodiment of the invention;
[0085] Figure 37 This is the ROC curve of the fifth target biomarker combination in the TCGA database genes disclosed in this embodiment of the invention;
[0086] Figure 38 This is a distribution map of the expression value of the CA4 gene in pan-cancer in the GEPIA database disclosed in the embodiments of the present invention;
[0087] Figure 39 This is a distribution map of the expression value of the CCT3 gene in pan-cancer in the GEPIA database disclosed in the embodiments of the present invention;
[0088] Figure 40 This is a distribution map of the expression value of the EPCAM gene in pan-cancer in the GEPIA database disclosed in the embodiments of the present invention;
[0089] Figure 41 This is a distribution map of the expression value of the LINC00092 gene in pan-cancer in the GEPIA database disclosed in the embodiments of the present invention;
[0090] Figure 42 This is a distribution map of the expression value of the MYO16-AS1 gene in pan-cancer in the GEPIA database disclosed in the embodiments of the present invention;
[0091] Figure 43 This is a distribution map of the expression value of the TRHDE-AS1 gene in pan-cancer in the GEPIA database disclosed in the embodiments of the present invention;
[0092] Figure 44 This is a distribution map of the expression value of the HMGB3 gene in pan-cancer in the GEPIA database disclosed in the embodiments of the present invention;
[0093] Figure 45 This is a distribution map of the expression value of the SPP1 gene in pan-cancer in the GEPIA database disclosed in the embodiments of the present invention;
[0094] Figure 46This is a heatmap of differential expression of miRNA genes in different cancer types in the dbDEMC database disclosed in this embodiment of the invention;
[0095] Figure 47 This is a schematic diagram of the structure of a sequencing-based molecular marker screening device disclosed in an embodiment of the present invention;
[0096] Figure 48 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of the present invention.
[0097] Explanation of reference numerals in the attached figures:
[0098] 401. Acquisition Unit; 402. Quality Control Unit; 403. Variation Detection Unit; 404. First Analysis Unit; 405. Second Analysis Unit; 406. Third Analysis Unit; 407. Synthesis Unit; 408. Determination Unit; 501. Memory; 502. Processor. Detailed Implementation
[0099] Unless otherwise specified or defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. When combined with the technical solutions of the invention in a real-world scenario, all technical and scientific terms used herein may also have meanings corresponding to the purpose of achieving the technical solutions of the invention. The terms "first," "second," etc., used herein are merely for distinguishing names and do not represent a specific number or order. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0100] It should be noted that when a component is considered "fixed" to another component, it can be directly fixed to the other component or there can be an intervening component; when a component is considered "connected" to another component, it can be directly connected to the other component or there can be an intervening component; when a component is considered "mounted" on another component, it can be directly mounted on the other component or there can be an intervening component; when a component is considered "placed" on another component, it can be directly placed on the other component or there can be an intervening component.
[0101] Unless otherwise specified or defined, the terms "described" or "the" as used herein refer to the technical features or technical content mentioned or described prior to the relevant section, which may be the same as or similar to the technical features or technical content mentioned herein. Furthermore, the terms "comprising" and "having," and any variations thereof, as used herein, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the steps or units listed, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such processes, methods, products, or apparatus.
[0102] This invention discloses a sequencing-based molecular biomarker screening method, which can be implemented through computer programming. The method can be executed by electronic devices such as computers, laptops, and tablets, or by sequencing-based molecular biomarker screening devices embedded in electronic devices; this invention does not limit this to any particular type. To facilitate understanding of this invention, specific embodiments will be described in more detail below with reference to the accompanying drawings.
[0103] like Figure 1 As shown, the method includes the following steps 110 to 170:
[0104] 110. Obtain sample sequencing data and perform quality control processing on the sample sequencing data to obtain clean sequencing data.
[0105] This invention can be applied to the screening of molecular markers for auxiliary diagnosis of various cancers, such as lung adenocarcinoma. The sample sequencing data includes initial gene sequencing data and initial miRNA sequencing data. The specific implementation method for obtaining the sample sequencing data includes: obtaining cancerous tissue from a clinically diagnosed cancer patient and its matched adjacent non-cancer tissue as paired samples; performing miRNA sequencing and RNA-seq sequencing on the RNA of the paired samples to obtain initial miRNA sequencing data and initial gene sequencing data respectively, and merging them to obtain the sample sequencing data.
[0106] In the miRNA sequencing process, the raw data files obtained from sequencing are converted into raw sequencing reads through base calling analysis. These reads are called Raw Data or Raw Reads and contain the sequence information of the sequencing reads and their corresponding sequencing quality information. This data is considered as the obtained sample sequencing data and is finally stored in FASTQ file format.
[0107] The raw data obtained from sequencing (i.e., sample sequencing data) contains low-quality reads with adapters. To ensure the quality of subsequent information analysis, the raw reads must undergo quality control filtering to obtain clean reads (i.e., clean sequencing data). The raw data quality control filtering steps are as follows: remove low-quality reads; remove reads with 5' adapter contamination; remove reads without 3' adapter sequences; remove reads without insert fragments; remove reads shorter than 15 nt. The final clean sequencing data includes clean gene sequencing data and miRNA sequencing data.
[0108] 120. Based on clean sequencing data, detect and obtain gene mutation sites.
[0109] Specifically, the gene variation analysis process includes: aligning clean sequencing data to a human reference genome to obtain alignment results, processing the alignment results, and identifying gene loci with mutation frequencies greater than or equal to a preset threshold as gene variation sites.
[0110] 130. Based on clean sequencing data, analyze and obtain the expression map of the target gene.
[0111] Specifically, step 130 may include the following steps 1301 to 1303 (not shown):
[0112] 1301. The clean gene sequencing data is aligned to the human reference genome to obtain the first gene data. After quantifying the first gene data to obtain the gene expression value, a comparison group is set up according to the normal group and the lung adenocarcinoma group. The difference analysis is performed on the comparison group to obtain the first gene expression map.
[0113] 1302. The clean miRNA sequencing data was aligned to the human reference genome to obtain the second gene data. After quantifying the second gene data to obtain the miRNA expression value, a comparison group was set up according to the normal group and the lung adenocarcinoma group. The difference analysis was performed on the comparison group to obtain the second gene expression map.
[0114] 1303. Based on the first gene expression map and the second gene expression map, obtain the target gene expression map.
[0115] 140. Perform weighted gene co-expression network analysis on the clean sequencing data to obtain the analysis results.
[0116] Specifically, based on gene expression values, miRNA expression values, combined with clinical data and sample characteristics, weighted gene co-expression network (WGCNA) analysis was performed on clean sequencing data to obtain WGCNA analysis results.
[0117] 150. Perform co-expression analysis on clean sequencing data, construct a miRNA-gene molecular interaction network, perform topological analysis on the miRNA-gene molecular interaction network, and identify key nodes.
[0118] Specifically, based on gene expression values and miRNA expression values, co-expression analysis was performed on clean sequencing data to construct a miRNA-gene molecular interaction network. Topological analysis was then conducted on the miRNA-gene molecular interaction network to identify key nodes.
[0119] 160. Based on gene mutation sites, target gene expression maps, analysis results, and key nodes of the miRNA-gene molecular interaction network, candidate biomarkers were screened.
[0120] Preferably, the screening criteria are that the following conditions are met simultaneously:
[0121] 1) Expression difference: The expression level of the biomarker in early cancer samples should be statistically different from that in adjacent normal tissues;
[0122] 2) Association with WGCNA core modules: Biomarkers should be distributed in the core color modules of WGCNA analysis that are closely related to tumor occurrence and development;
[0123] 3) Key nodes in the network: The biomarker should occupy a core position in the miRNA-gene molecular interaction network and have multiple direct and indirect interaction relationships;
[0124] 4) Gene mutation frequency: Genes with high mutation frequencies are more likely to be directly related to the occurrence and development of tumors.
[0125] Specifically, a molecular marker is considered as a candidate marker if it meets the following criteria: differential expression, association with the WGCNA core module, being a key node in the miRNA-gene molecular interaction network, and being a gene mutation site.
[0126] 170. Using machine learning algorithms, calculate the importance scores of all candidate markers.
[0127] Specifically, a biomarker evaluation model is constructed based on machine learning algorithms such as random forest and stepwise regression, and the biomarker evaluation model is trained to obtain a trained biomarker evaluation model.
[0128] The gene mutation sites and target gene expression maps are input into the trained biomarker evaluation model to obtain the importance score of each candidate biomarker.
[0129] The step of training the biomarker evaluation model to obtain a trained biomarker evaluation model includes:
[0130] If N training sets are obtained in advance, then n samples with replacement are randomly selected from the N training sets for a single decision tree as training samples for that decision tree.
[0131] Let the number of input features of the training samples be M, where m is less than M. Then, when splitting at each node of each decision tree, we randomly select m input features from the M input features, and then select the best one from the m input features to split, until all training samples of that node belong to the same class.
[0132] 180. Sort all candidate markers from highest to lowest importance score, and select the top-ranked candidate markers and their combinations as target markers.
[0133] Based on gene expression profiling, WGCNA analysis results, and integrated analysis of miRNA-gene molecular interaction networks, the most promising candidate biomarkers and their combinations were ultimately obtained. In this embodiment of the invention, the candidate biomarkers among the target biomarkers include at least one of the following: CA4, CCT3, EPCAM, hsa-miR-133a-3p, hsa-miR-133b, hsa-miR-139-3p, hsa-miR-139-5p, hsa-miR-21-3p, hsa-miR-30c-2-3p, LINC00092, MYO16-AS1, and TRHDE-AS.
[0134] After obtaining the target marker, the process also includes:
[0135] Obtain the expression information of the target biomarker in pan-cancer; the process of obtaining the expression information of the target biomarker in pan-cancer includes: obtaining the expression information of the target biomarker in the GEPIA and dbDEMC databases to verify the target biomarker;
[0136] The target biomarker was validated using its expression information in pan-cancer.
[0137] Specifically, to validate the target biomarkers, in this embodiment of the invention, different population cohorts, different sequencing methods, and external independent datasets from different sample sources were used to validate the discovered target biomarkers, ensuring the reliability and reproducibility of the results. By using the expression information of the target biomarkers in different databases and datasets, the expression information of the target biomarkers in different cancers can be obtained. By comparing cancer samples with normal samples, it can be seen that the target biomarkers are not only significantly differentially expressed in lung cancer, but also exhibit similar expression patterns in other cancers.
[0138] Example 1
[0139] This embodiment collected high-quality samples from 94 patients with lung adenocarcinoma. Each patient provided matched cancerous tissue and adjacent adjacent normal tissue. Thirty stage I patients were randomly selected as the training set, and the other 64 samples were used as an expanded validation set. The samples were then processed in step 110 to obtain clean sequencing data.
[0140] Step 120: Comprehensive whole transcriptome sequencing (RNA-seq) and microRNA (miRNA) sequencing were performed on the sample, followed by detailed mutation analysis based on the RNA-seq sequencing data. For example... Figure 3 As shown, genes with a mutation frequency greater than or equal to a preset threshold are considered high-frequency mutated genes, i.e., gene mutation sites. The preset threshold can be set to 10%.
[0141] Table 1 Importance scores of molecular markers calculated by machine learning algorithms
[0142]
[0143]
[0144] Step 130:
[0145] By comparing the expression profiles of cancerous tissue and adjacent normal tissue, the expression profiles of genes and miRNAs specific to lung adenocarcinoma were determined. Furthermore, the expression information of normal tissue and tumor samples was compared, such as... Figure 4 and Figure 5As shown, specific examples include 10,453 differentially expressed genes and 410 differentially expressed miRNAs between normal tissue and patients with early-stage lung adenocarcinoma. Further molecular functional analysis identified key molecules closely associated with different tumor stages and specific pathological types: collagen family members such as COL10A1 and COL11A1 were significantly overexpressed in stage I. This characteristic change is mainly associated with the enhancement of basic physiological functions of lung tissue and the acceleration of substance synthesis and metabolism, particularly the remodeling of cell structure, enhanced biosynthesis of steroid hormones, and optimization of protein digestion and absorption. These findings reveal the profound impact of early-stage lung adenocarcinoma on the lung tissue microenvironment. As the disease progresses to stage II, the gene expression profile exhibits new trends, with enhanced cell cycle regulation, senescence processes, and DNA replication becoming dominant, suggesting a significant increase in the proliferative capacity of cancer cells. Simultaneously, the immune system response begins to manifest, with negative regulation of immune pathways such as complement and coagulation cascades, possibly reflecting subtle changes in the immune balance within the tumor microenvironment. In stage III, although most upregulation pathways are similar to those in stage II, the scope of immunosuppression expands further, involving negative regulation of multiple important immune pathways such as the tumor necrosis factor signaling pathway and IL-7-mediated cell differentiation. As the disease progresses to advanced stages, the pathways and functions involved in the gene expression profile become even more complex. In summary, changes in gene expression at different stages of lung adenocarcinoma not only reveal the molecular mechanisms of disease progression but also provide potential therapeutic targets and prognostic indicators for clinical practice.
[0146] Table 2. Expression data of some genes
[0147]
[0148] Step 140:
[0149] To further explore the association between gene expression modules and clinical characteristics, a combined analysis was performed using the WGCNA method. Figures 6-9 As shown, this method clusters highly correlated genes into different modules by calculating co-expression relationships between genes, and assesses the correlation between these modules and the clinical phenotypes of lung adenocarcinoma (such as tumor stage and survival). Through WGCNA analysis, several gene modules highly associated with the occurrence of lung adenocarcinoma have been successfully identified.
[0150] Step 150:
[0151] Based on gene and miRNA expression data, complex molecular interaction networks were constructed, which visually demonstrate the interactions between genes and between miRNAs and target genes. This analysis revealed the regulatory mechanisms of key signaling pathways in lung adenocarcinoma and identified several potential regulatory hubs. Figure 10As shown, the interaction network of some core genes and miRNAs is displayed. Orange nodes represent mRNA, green nodes represent miRNA, and blue nodes represent lncRNA molecules. The larger the network node, the more important the molecule at that node.
[0152] Step 160: Analyze gene mutation sites and gene expression values as model factors, and give an importance score for each factor.
[0153] Table 3 Importance scores of miRNA biomarkers obtained based on machine learning algorithms
[0154]
[0155] Table 4 Importance scores of genomic biomarkers obtained based on machine learning algorithms
[0156]
[0157]
[0158] Step 170:
[0159] A molecular biomarker is selected as a candidate biomarker when it simultaneously meets the following criteria: (1) significantly differentially expressed genes, where the significantly differentially expressed genes are those with log2FC>=1 and qvalue<0.05, and miRNAs are obtained through target gene expression mapping; (2) gene mutation sites; (3) key nodes in the miRNA-gene molecular interaction network; and (4) correlation with clinical data in WGCNA analysis. Simultaneously, a machine learning algorithm is used to obtain the importance scores and rankings of all candidate biomarkers. Finally, based on the candidate markers and the results of the machine learning algorithm, one or more of the following were selected as target markers: CA4, CCT3, EPCAM, hsa-miR-133a-3p, hsa-miR-133b, hsa-miR-139-3p, hsa-miR-139-5p, hsa-miR-21-3p, hsa-miR-30c-2-3p, LINC00092, MYO16-AS1, TRHDE-AS1, SPP1, and HMGB3.
[0160] To further evaluate the accuracy of the target biomarker in distinguishing cancer patients from healthy individuals, the subject characteristic curve (ROC) was plotted using the R package "pROC". Figures 11-27 As shown, AUC values, sensitivity, and specificity are analyzed to determine the diagnostic efficacy of indicators alone or in combination. A higher AUC (closer to 1) indicates better model performance. An AUC ≥ 0.9 indicates an excellent model. Figure 23The first target marker combination is CA4_hsa-miR-133a-3p+hsa-miR-139-5p+hsa-miR-30c-2-3p, the second target marker combination is hsa-miR-133b+hsa-miR-139-3p+LINC00092+MYO16-AS1+TRHDE-AS1+LINC00092, the third target marker combination is CCT3+EPCAM+hsa-miR-21-3p, the fourth target marker combination is hsa-miR-21-3p+CCT3+EPCAM+SPP1, and the fifth target marker combination is hsa-miR-21-3p+EPCAM+HMGB3+SPP1.
[0161] Validate the results using both internal and external data:
[0162] Internally: To verify the accuracy and generality of the selected biomarkers, we used the remaining 64 samples for validation. The expression values of the selected biomarkers in cancer and adjacent tissues showed statistical differences, and the expression trends were consistent with those of the selected 30 samples.
[0163] Externally: Further, we systematically integrated and analyzed data from multiple independent cohorts in the TCGA (The Cancer Genome Atlas), a large-scale cancer genomics database, and the GEO (Gene Expression Omnibus) database, ensuring that at least two datasets were derived from blood samples. We used differential expression algorithms such as limma and DESeq to compare the expression levels of selected biomarkers in early-stage cancer samples and normal samples to determine their significance and specificity. To visually represent the results, we constructed a heatmap to depict the expression patterns of the selected biomarkers in the TCGA and GEO data cohorts. In the heatmap, blue highlights indicate significant expression of the biomarker in the corresponding data, consistent with the expression trend observed in this experiment. Figures 28-29 As shown.
[0164] To further evaluate the accuracy of the screened molecular markers in distinguishing cancer patients from healthy individuals in independent samples, receiver operating characteristic (ROC) curves were also plotted using TCGA data. Figures 31-37 As shown.
[0165] ROC curve analysis shows that the AUC values of the target biomarkers in the TCGA dataset are all close to 1, which successfully verifies the accuracy and reliability of the target molecular biomarkers in distinguishing between cancer patients and healthy individuals.
[0166] Measurement of the expression value of target biomarkers in pan-cancer
[0167] To comprehensively evaluate and validate the stability and consistent expression of the target biomarkers across various cancer backgrounds, we utilized widely recognized database resources: GEPIA and dbDEMC. These two databases not only contain abundant gene expression data but also cover a wide range of cancer types and samples, providing a foundation for our analysis.
[0168] Combining the pan-cancer analysis results of GEPIA and dbDEMC, such as Figures 38-46 As shown, the gene distribution in the database is clearly consistent with the aforementioned expression trends. This high degree of consistency not only verifies the widespread presence and stable expression of the selected molecular markers in different cancer backgrounds, but also reveals their great potential as diagnostic markers for multiple cancer types.
[0169] Finally, by combining molecular expression data from the GEPIA and dbDEMC databases, expression information of the selected molecular markers in different cancer types was obtained. A systematic evaluation of the expression patterns and potential value of these molecular markers across pan-cancer populations provides a new perspective for the diagnosis and treatment of lung cancer and other cancers.
[0170] In summary, by implementing the embodiments of the present invention, gene mutation results, gene expression data, rich clinical information, and complex intermolecular interaction network information are organically integrated in multiple dimensions. Through in-depth analysis in multiple dimensions, the potential efficacy of these molecules is comprehensively evaluated.
[0171] Therefore, this invention reveals the remarkable efficacy of molecular biomarkers or combinations thereof. This discovery is not only strongly supported by internal sequencing data, but also demonstrates high confirmability and consistency in a wide range of datasets from different ethnic groups, as well as in serum, plasma, and exosome samples obtained through non-invasive methods. In other words, the obtained tissue biomarkers can also serve as blood biomarkers.
[0172] like Figure 4 As shown, this embodiment of the invention discloses a sequencing-based molecular marker screening device, including an acquisition unit 401, a quality control unit 402, a variant detection unit 403, a first analysis unit 404, a second analysis unit 405, a third analysis unit 406, a screening unit 407, a calculation unit 408, and a determination unit 409, wherein...
[0173] Acquisition unit 401 is used to acquire sample sequencing data;
[0174] Quality control unit 402 is used to perform quality control processing on sample sequencing data to obtain clean sequencing data;
[0175] The variant detection unit 403 is used to detect and obtain gene variant sites based on clean sequencing data;
[0176] The first analysis unit 404 is used to analyze and obtain the expression map of the target gene based on clean sequencing data;
[0177] The second analysis unit 405 is used to perform weighted gene co-expression network analysis on clean sequencing data to obtain analysis results.
[0178] The third analysis unit 406 is used to perform co-expression analysis on clean sequencing data, construct miRNA-gene molecular interaction network, perform topological analysis on miRNA-gene molecular interaction network, and identify key nodes;
[0179] The screening unit 407 is used to screen candidate biomarkers based on the gene mutation sites, the target gene expression map, the analysis results, and the key nodes of the miRNA-gene molecular interaction network.
[0180] The computing unit 408 is used to calculate the importance scores of all candidate markers using a machine learning algorithm.
[0181] The determination unit 409 is used to sort all candidate markers from largest to smallest according to their importance scores, and obtain a specified number of candidate markers and their combinations that rank highly as target markers.
[0182] Optionally, the acquisition unit 401 is specifically used to acquire cancer tissue from patients clinically diagnosed with cancer and their matching adjacent non-cancer tissue as paired samples; to perform miRNA sequencing and RNA-seq sequencing on the RNA of the paired samples to obtain initial miRNA sequencing data and initial gene sequencing data, respectively, and to merge them to obtain sample sequencing data.
[0183] Optionally, the above-mentioned variant detection unit 403 is specifically used to align the clean sequencing data to a human reference genome to obtain alignment results, and to detect gene variant sites based on the alignment results.
[0184] Optionally, clean sequencing data includes clean gene sequencing data and miRNA sequencing data; the first analysis unit 404 mentioned above includes the following subunits not shown:
[0185] The first differential analysis subunit is used to align clean gene sequencing data to the human reference genome to obtain first gene data. After quantifying the first gene data to obtain gene expression values, a comparison group is set up according to the normal group and lung adenocarcinoma, and differential analysis is performed on the comparison group to obtain the first gene expression map.
[0186] The second differential analysis subunit is used to align clean miRNA sequencing data to the human reference genome to obtain second gene data. After quantifying the second gene data to obtain miRNA expression values, comparison groups are set up according to the normal group and lung adenocarcinoma, and differential analysis is performed on the comparison groups to obtain the second gene expression profile.
[0187] The acquisition subunit is used to obtain the target gene expression map based on the first gene expression map and the second gene expression map.
[0188] like Figure 5 As shown, an embodiment of the present invention discloses an electronic device, including a memory 501 storing executable program code and a processor 502 coupled to the memory 501;
[0189] The processor 502 calls the executable program code stored in the memory 501 to execute the sequencing-based molecular marker screening method described in the above embodiments.
[0190] This invention also discloses a computer-readable storage medium storing a computer program that causes a computer to execute the sequencing-based molecular marker screening methods described in the above embodiments.
[0191] The purpose of the above embodiments is to reproduce and derive the technical solution of the present invention by way of example, and to fully describe the technical solution, purpose and effect of the present invention. The purpose is to enable the public to have a more thorough and comprehensive understanding of the disclosure of the present invention, and not to limit the scope of protection of the present invention.
[0192] The above embodiments are not an exhaustive list based on the present invention, and there may be many other embodiments not listed. Any substitutions and improvements made without departing from the concept of the present invention are within the protection scope of the present invention.
Claims
1. A sequencing-based method for screening molecular biomarkers, characterized in that, include: Obtain sample sequencing data and perform quality control processing on the sample sequencing data to obtain clean sequencing data; Based on the clean sequencing data, gene mutation sites were detected and obtained; Based on the clean sequencing data, the expression profile of the target gene was obtained by analysis; Weighted gene co-expression network analysis was performed on the clean sequencing data to obtain the analysis results; Co-expression analysis was performed on the clean sequencing data to construct a miRNA-gene molecular interaction network. Topological analysis was then performed on the miRNA-gene molecular interaction network to identify key nodes. Candidate biomarkers were screened based on the gene mutation sites, the target gene expression map, the analysis results, and the key nodes of the miRNA-gene molecular interaction network. Machine learning algorithms are used to calculate the importance scores of all candidate biomarkers; Sort all candidate markers by importance score from highest to lowest, and select the top-ranked candidate markers and their combinations as target markers.
2. The sequencing-based molecular biomarker screening method as described in claim 1, characterized in that, Obtain sample sequencing data, including: Cancerous tissue from patients clinically diagnosed with cancer and their matching adjacent non-cancer tissues were obtained as paired samples; the RNA of the paired samples was subjected to miRNA sequencing and RNA-seq sequencing to obtain initial miRNA sequencing data and initial gene sequencing data, which were then combined to obtain sample sequencing data.
3. The sequencing-based molecular biomarker screening method as described in claim 1, characterized in that, Based on the clean sequencing data, gene mutation sites were detected, including: The clean sequencing data is aligned to a human reference genome to obtain alignment results, and gene mutation sites are detected based on the alignment results.
4. The sequencing-based molecular biomarker screening method according to any one of claims 1 to 3, characterized in that, The clean sequencing data includes clean gene sequencing data and miRNA sequencing data; Based on the clean sequencing data, the expression profiles of the target genes were analyzed and obtained, including: Clean gene sequencing data is aligned to a human reference genome to obtain the first gene data. After quantifying the first gene data to obtain gene expression values, a comparison group is set up according to the normal group and lung adenocarcinoma, and differential analysis is performed on the comparison group to obtain the first gene expression map. Clean miRNA sequencing data was aligned to the human reference genome to obtain second gene data. After quantifying the second gene data to obtain miRNA expression values, comparison groups were set up according to the normal group and lung adenocarcinoma, and differential analysis was performed on the comparison groups to obtain the second gene expression profile. Based on the first gene expression map and the second gene expression map, the target gene expression map is obtained.
5. The sequencing-based molecular marker screening method according to any one of claims 1 to 3, characterized in that, The importance scores of all molecular markers were calculated using machine learning algorithms, including: A biomarker evaluation model is constructed based on machine learning algorithms such as random forest and stepwise regression. The biomarker evaluation model is then trained to obtain a trained biomarker evaluation model. The gene mutation sites and target gene expression maps are input into the trained biomarker evaluation model to obtain the importance score of each candidate biomarker.
6. The sequencing-based molecular biomarker screening method as described in claim 5, characterized in that, The step of training the biomarker evaluation model to obtain a trained biomarker evaluation model includes: If N training sets are obtained in advance, then n samples with replacement are randomly selected from the N training sets for a single decision tree as training samples for that decision tree. Let the number of input features of the training samples be M, where m is less than M. Then, when splitting at each node of each decision tree, we randomly select m input features from the M input features, and then select the best one from the m input features to split, until all training samples of that node belong to the same class.
7. The sequencing-based molecular biomarker screening method according to any one of claims 1-5, characterized in that, After obtaining the target marker, the process also includes: Obtain the expression information of the target biomarker in pan-cancer; the process of obtaining the expression information of the target biomarker in pan-cancer includes: obtaining the expression information of the target biomarker in the GEPIA and dbDEMC databases to verify the target biomarker; The target biomarker was validated using its expression information in pan-cancer.
8. A sequencing-based molecular marker screening device, characterized in that, include: Acquisition unit, used to acquire sample sequencing data; The quality control unit is used to perform quality control processing on sample sequencing data to obtain clean sequencing data; A mutation detection unit is used to detect and obtain gene mutation sites based on the clean sequencing data; The first analysis unit is used to analyze and obtain the target gene expression map based on the clean sequencing data; The second analysis unit is used to perform weighted gene co-expression network analysis on the clean sequencing data to obtain analysis results. The third analysis unit is used to perform co-expression analysis on the clean sequencing data, construct a miRNA-gene molecular interaction network, perform topological analysis on the miRNA-gene molecular interaction network, and identify key nodes. The screening unit is used to screen candidate biomarkers based on the gene mutation sites, the target gene expression map, the analysis results, and the key nodes of the miRNA-gene molecular interaction network. The computational unit is used to calculate the importance scores of all candidate markers using machine learning algorithms; The determining unit is used to sort all candidate markers from largest to smallest according to their importance scores, and to obtain a specified number of candidate markers and their combinations that rank highest as target markers.
9. An electronic device, characterized in that, It includes a memory storing executable program code and a processor coupled to the memory; the processor calls the executable program code stored in the memory to execute the sequencing-based molecular marker screening method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program causes a computer to perform the sequencing-based molecular marker screening method according to any one of claims 1 to 7.
Citation Information
Patent Citations
NMIBC prognosis prediction molecular marker, screening method and modeling method
CN114582425A
Method and system for screening panthenic cancer early screening molecular marker based on whole genome bisulfite sequencing data
CN115424666A
Method for acquiring skeletal muscle early injury time inference gene set system screened based on next-generation sequencing technology and model construction method of skeletal muscle early injury time inference gene set system
CN116364310A
Screening method of biomarkers related to preeclampsia lysosomes and application of screening method
CN117012276A