Gene combination lung cancer screening diagnosis and targeted therapy method and product
Through screening and targeted therapy methods of NLRC4, PLEKHN1, RASIP1 and SPP1 gene combinations, the high false positive rate and limited selection problems of existing lung cancer screening and targeted therapy are solved, and the effects of high accuracy and personalized treatment are achieved.
Patent Information
- Application Number
- CN202510270851.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-05-30
AI Technical Summary
Existing lung cancer screening and targeted therapy methods have high false positive rates, limited targeted therapy options and drug resistance caused by tumor heterogeneity, which is difficult to meet the needs of efficient screening in clinical practice.
The genotype or expression level of these genes was detected by specific probes or detection reagents, combined with the mathematical model of gene combination risk assessment, and the lung cancer screening diagnosis was performed, and targeted drugs were developed for these genes for personalized treatment.
High accuracy and sensitivity of lung cancer screening is achieved, false positive rates are reduced, and more personalized targeted treatment plans are provided, overcoming the limitations of the existing technology.
Smart Images

Figure CN120060474A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of screening diagnosis, subtype identification and treatment of lung cancer, and specifically relates to a method and product for screening diagnosis and targeted treatment of lung cancer with a gene combination. Background Art
[0002] As a fatal disease, the incidence of lung cancer has been continuously rising globally, becoming one of the major challenges in the field of public health. Early lung cancer often lacks specific symptoms, resulting in most patients being diagnosed at an advanced stage, which greatly limits the treatment effect and the survival rate of patients.
[0003] To address the above challenges, the medical community has introduced various lung cancer screening tools and technologies, such as low-dose spiral CT scanning, biomarker detection, etc., in order to improve the detection rate of early lung cancer. At the same time, targeted therapy, as part of precision medicine, realizes more personalized treatment of lung cancer by identifying specific gene mutations or protein changes in cancer cells. These methods have improved the prognosis of lung cancer patients to a certain extent and provided new hope for those patients with clear driver gene mutations.
[0004] However, the existing methods for lung cancer screening and targeted treatment still have many deficiencies. On the one hand, although the existing screening means are helpful for early detection of lesions, they may be accompanied by a high false positive rate, bringing unnecessary anxiety to patients and the burden of further examinations. On the other hand, for most non-small cell lung cancer patients, the available targeted treatment options are limited, and due to the high heterogeneity within and between tumors, the treatment with a single gene target is prone to drug resistance problems. In addition, the number of lung cancer-related genes discovered in previous studies is huge, but the prediction performance is poor, the accuracy and sensitivity are not high, and it is difficult to meet the high-efficiency screening requirements in clinical practice. Therefore, it is particularly important to develop a lung cancer screening diagnosis and targeted treatment plan based on an optimized gene combination, aiming to overcome the limitations of the existing technology and provide a more accurate and effective medical solution. Summary of the Invention
[0005] Object of the Invention: To solve the above problems, the present invention provides a method and product for screening diagnosis and targeted treatment of lung cancer with a gene combination.
[0006] Technical Solution: A gene combination includes the following genes: NLRC4, PLEKHN1, RASIP1, and SPP1 .
[0007] A product for screening diagnosis of lung cancer includes: at least one specific probe or detection reagent for detecting NLRC4, PLEKHN1, RASIP1, and SPP1 genotype or expression level.
[0008] A product for targeted treatment of lung cancer includes: forNLRC4, PLEKHN1, RASIP1, and SPP1 Targeted drugs for one or more of the genes and their subtypes, and a drug delivery device for delivering the targeted drugs into a patient's body.
[0009] A method for screening and diagnosing lung cancer using the gene combination as described above, comprising the following steps: Obtain a biological sample from a subject; Use a specific probe or assay to actually detect the NLRC4, PLEKHN1, RASIP1, and SPP1 Genotype or expression level of the gene in the biological sample to obtain detection data; Analyze the detection data using a gene combination risk assessment mathematical model to identify lung cancer risk or diagnose lung cancer subtypes to obtain an analysis result; Based on the analysis result, determine whether the subject has lung cancer and the specific subtype of lung cancer.
[0010] In a further embodiment, it comprises the following steps: Determine the specific subtype of lung cancer; Select a suitable targeted drug according to the lung cancer subtype; wherein, the targeted drug is: for NLRC4, PLEKHN1, RASIP1, and SPP1 Targeted drugs for one or more of the genes and their subtypes; Administer the targeted drug through a drug delivery device to achieve personalized treatment.
[0011] Use of the gene combination as described above in screening and diagnosing lung cancer.
[0012] Use of the gene combination as described above in targeted therapy for lung cancer.
[0013] For the method of screening and diagnosing lung cancer as described above, after obtaining the expression values of these genes, the expression formula of the gene combination risk assessment mathematical model is as follows: ; In the formula, Represents the risk assessment value corresponding to the formula , Takes values of 1, 2, 3; Is an arbitrary real number, Is the intercept corresponding to within the formula , , , And Are the coefficients corresponding to within the formula , , , And Respectively represent the gene expression values of NLRC4, PLEKHN1, RASIP1, and SPP1; The maximum risk value is calculated using the following formula: : ; Based on the maximum risk value Calculate the probability of cancer : ; like , then it is judged as cancer, among which This is a pre-set judgment threshold based on the type of cancer.
[0014] In a further embodiment, the following steps are also included to construct a four-dimensional visualization table, and the specific steps are as follows: Selecting three specific genes as basic elements for constructing a three-dimensional space coordinate axis, and constructing a three-dimensional space coordinate system based on the basic elements; In the three-dimensional space coordinate system, different individuals will determine their specific positions in space based on the expression values of the three genes corresponding to them, so as to show the distribution of patients and healthy people in the three-dimensional space; At the same time, different colors are used to indicate the approximate risk of lung cancer. The intuitive difference in colors enables clinicians to quickly understand the risk level of lung cancer for different individuals. The three-dimensional space displays individual distribution and the risk probability of lung cancer indicated by color to form a four-dimensional visualization chart.
[0015] Beneficial effects: (1) Mathematically formulated gene combination risk assessment model: The present invention models the gene combination through a set of clear mathematical formulas, each of which contains the gene name and the corresponding explanatory coefficient. The positive or negative value of the coefficient in the formula directly reveals the positive or negative impact of high or low expression of each gene on the risk of lung cancer, greatly enhancing the interpretability of the model. This quantitative relationship between gene expression and lung cancer risk has not appeared in previous studies, filling the gap in the interpretability of existing gene screening methods.
[0016] (2) Original four-dimensional visualization chart: Combining the four-dimensional visualization of gene expression distribution and probability bar chart, the present invention clearly shows the contribution of gene combinations to lung cancer risk. This intuitive graphical method helps researchers and clinicians understand the synergistic and independent effects between genes, thereby optimizing diagnostic decisions.
[0017] (3)Precise and reproducible gene screening: In response to the challenge reported in the literature that "almost every gene is associated with cancer", the present invention screens out only 4 core genes from tens of thousands of genes through tissue sample analysis, achieving a diagnostic accuracy of nearly 99.99%. This breakthrough result overcomes the bottleneck in traditional gene screening that makes it difficult to achieve targeted therapy due to the large number of genes.
[0018] (4)Synergistic effect of core genes: The 4 screened core genes are distributed on different chromosomal strands, and their individual expression changes are not significant, but the combined (synergistic) effect in lung cancer recognition reaches the optimal level. This synergistic relationship is not a traditional biological gene interaction, but a synergistic mode with both mathematical and biological equivalence, providing a new perspective for genomic analysis.
[0019] (5)Meeting the goals of international frontier research: The present invention theoretically solves the 15th of the 23 biomathematics challenges in the 21st century proposed by the Defense Advanced Research Projects Agency ( DARPA ), namely the genomic space geometry problem. By establishing a mathematical space model of gene expression, the present invention not only has biological significance but also meets the scientific requirements of geometric modeling and multi-dimensional data analysis. Brief Description of the Drawings
[0020] Figure 1 is the first four-dimensional visualization chart of lung adenocarcinoma.
[0021] Figure 2 is the second four-dimensional visualization chart of lung adenocarcinoma.
[0022] Figure 3 is the first four-dimensional visualization chart of lung squamous cell carcinoma.
[0023] Figure 4 is the second four-dimensional visualization chart of lung squamous cell carcinoma. Detailed Description of the Invention
[0024] Example 1 Explain how the present invention uses NLRC4, PLEKHN1, RASIP1, and SPP1 gene combinations for lung cancer screening diagnosis and treatment through Example 1, which specifically includes the following steps: Step S 1. Obtain tissue samples from the tumor and adjacent areas of the subject through surgical operations (such as lobectomy), bronchoscopic biopsy, or fine needle aspiration biopsy.
[0025] Step S 2. Sample preparation: 1) Tissue fixation: It is planned to fix the tissue with 10% neutral buffered formalin to maintain cell morphology and DNA / RNA stability; 2) Tissue embedding: After dehydrating the fixed tissue, embed it in paraffin ( FFPE , formalin-fixed paraffin embedding); 3) Sectioning and staining: After sectioning the paraffin tissue, determine the regions of cancerous tissue and adjacent non-cancerous tissue by HE staining (hematoxylin-eosin staining).
[0026] Step S 3. Nucleic acid extraction.
[0027] Step S 4. Design and synthesize specific probes or detection reagents for NLRC4, PLEKHN1, RASIP1, and SPP1 genes, and sequence to detect whether the combination of these four genes constitutes a high-risk type for lung cancer.
[0028] Thus, early-stage and disease-stage screening for lung cancer can be carried out: Individuals with the abnormal genotype of simultaneous mutations in these 4 genes will have a higher probability of developing lung cancer than the general population.
[0029] Simultaneously performing multiple simultaneous detections of the 4 genes, or separately detecting and then aggregating the sequencing results of the 4 genes, are all possible implementation methods.
[0030] Step S 5. According to the genotype differences and expression differences of the subjects, use a gene combination risk assessment mathematical model to classify lung cancer, and place the new sample in a four-dimensional space chart to see which region the lung cancer sample falls into. NLRC4, PLEKHN1, RASIP1, and SPP1
[0031] Step S 6. Design and synthesize drugs targeting NLRC 4. PLEKHN 1. SPP 1 or RASIP 1 to regulate the expression level of one or more genes, or intervene at any link in their expression pathway (including but not limited to DNA expression, RNA expression, protein expression, biological activity of proteins, etc.), thereby performing targeted therapy for lung cancer.
[0032] Step S 7. During the above-mentioned targeted therapy process, detect NLRC4, PLEKHN1, RASIP1, and SPP1 the expression level and expression pathway of genes, monitor and evaluate the treatment effect, and adjust the treatment plan according to the results.
[0033] Step S 8. Design or synthesize new editing tools for NLRC 4. PLEKHN 1. SPP 1 or RASIP 1 genes, or use existing general editing tools (such as CRISPRtechnology) for preventing or treating lung cancer.
[0034] For example, if NLRC the expression value of 4 is increased, and SPP the expression value of 1 is decreased, and the other two genes are determined to be increased or decreased according to their positions in the 4D graph. Editing one or more of these four genes can reduce the incidence probability of non-diseased individuals or treat diseased individuals.
[0035] Furthermore, in step S 5, the expression formula of the gene combination risk assessment mathematical model is as follows: ; In the formula, represents the risk assessment value corresponding to formula , takes values of 1, 2, 3; is an arbitrary real number, is the intercept corresponding to formula inside, , , and are the coefficients corresponding to formula inside, , , and respectively represent the gene expression values of NLRC4, PLEKHN1, RASIP1, and SPP1; The maximum risk value is calculated using the following formula: ; Based on the maximum risk value , the cancer probability is calculated: ; If , it is determined to have cancer, where is the judgment threshold preset according to the cancer type.
[0036] For further illustration by example, using the gene expression values calculated by the Illumina platform and RNAseq technology according to log2(fpkm + 1) in the GDC TCGA Lung Adenocarcinoma (LUAD) updated on May 10, 2024, the present application obtains the following specific formula corresponding to the above formula: ; ; ; Further, taking lung adenocarcinoma as an example, set , then after calculation, if , it is determined that the patient has lung adenocarcinoma, otherwise the patient does not have lung adenocarcinoma. Currently, the determination level of this formula is: 99.66% precision, 99.62% sensitivity, and 99.99% specificity. Table 1 gives other partial data for reference.
[0037] Table 1 Data set of the mathematical model for gene combination risk assessment of lung adenocarcinoma patients In another embodiment, regarding lung squamous cell carcinoma, using the gene expression values calculated according to log2(fpkm + 1) based on the Illumina platform and RNAseq technology updated on May 9, 2024 for GDC TCGA Lung Squamous Cell Carcinoma (LUSC lung squamous cell carcinoma, squamous cell lung cancer), the present application obtains the following specific formula corresponding to the above formula: ; ; ; . The probability risk of identifying lung squamous cell carcinoma is the same as the identification steps for lung adenocarcinoma. Currently, the determination level of the obtained formula is: 99.99% precision, 99.99% sensitivity, and 99.99% specificity. Table 2 gives other partial data for reference.
[0038] Table 2 Data set of the mathematical model for gene combination risk assessment of lung squamous cell carcinoma patients Example 2 Based on what is disclosed in Example 1, this example further describes the construction process of the thinking visualization table in step 5. Refer to Figures 1 to 4 , and the specific steps are as follows: Select three specific genes as the basic elements for constructing the three-dimensional space coordinate axes, and construct a three-dimensional space coordinate system based on these basic elements; In the three-dimensional space coordinate system, different individuals will determine their specific positions in the space according to the expression values of these three genes corresponding to them, so as to show the distribution of patients and healthy people in this three-dimensional space; At the same time, use different colors to identify the approximate risk probability of having lung cancer, so that clinicians can quickly understand the risk levels of different individuals with lung cancer through the intuitive difference in colors; Combining the display of individual distribution in three-dimensional space and the probability of lung cancer risk indicated by colors constitutes a four-dimensional visualization chart.
[0039] In other words, once the expression values of 4 genes of a new potential patient are collected, clinicians can select the best treatment plan based on the relative position of the patient's indicators in the four-dimensional visualization chart and the treatment conditions of other patients (which requires referring to the literature and accumulation).
[0040] Example 3 This example uses four core datasets to NLRC4, PLEKHN1, RASIP1, and SPP1 conduct actual verification on the lung cancer screening and diagnosis of gene combinations. The four core datasets are respectively: Dataset I is Nature the lung adenocarcinoma ( LUAD ) dataset published in a journal, containing 576 samples (517 tumor samples and 59 normal samples). The samples used in this application are the latest samples updated on May 10, 2024 (530 tumor samples and 59 normal samples).
[0041] Dataset II is Nature the comprehensive genomic characteristics of squamous cell lung cancer ( LUSC ) published in a publication, containing 552 (501 tumor samples and 51 normal samples).
[0042] Dataset III is the European sample for the classification and survival prediction of non-small cell lung cancer ( NSCLC ) based on gene expression, containing 156 samples (91 tumor samples and 65 normal samples).
[0043] Dataset IV is the Japanese lung adenocarcinoma data, containing 226 samples (158 variant samples and 68 normal samples).
[0044] Meanwhile, this example also uses eleven datasets from the National Center for Biotechnology Information of the United States ( National Center for Biotechnology Information , abbreviated as NCBI ) GSE 10245, GSE 18842, GSE 19188, GSE 19804, GSE 27262, GSE 43458, GSE 60052, GSE 75037, GSE 116946, GSE 151103, GSEThe experiment verification of 171415 has a consistently good effect.
[0045] At the same time, comparisons were made between lung adenocarcinoma and squamous cell lung cancer, between genomic expressions of smoking and never-smoking lung cancer patients, and between genomic expressions of tumor cells and macrophages in lung cancer, and it was found that these four genes have a high recognition efficacy.
[0046] In this embodiment, upstream and downstream analyses of the core genes were also carried out, and it was found that those being concerned in the literature and currently in the industry EGFR 、 KRAS 、 TP 53 is downstream of the genes listed in this invention application, indicating that the functions of these three genes are limited.
[0047] The method of this embodiment achieved 99.99% sensitivity and 99.99% specificity in predicting all 576 samples in their respective groups in the data set I ; In the data set II in classifying all 552 samples in their respective groups, it achieved 99.99% sensitivity and 99.99% specificity.
[0048] For the data set III in classifying all 156 samples in their respective groups, the sensitivity was 97.8% and the specificity was 99.99%.
[0049] In the data set IV 224 samples were divided into their respective groups, and its sensitivity and specificity were 99.99% and 95% respectively. The above data are all superior to any previously disclosed method.
[0050] Therefore, the originality and application potential of the NLRC4, PLEKHN1, RASIP1, and SPP1 gene combination for lung cancer screening and diagnosis proposed in this application are significantly superior to traditional gene-based lung cancer screening methods, providing solid theoretical and technical support for the precise diagnosis, precise treatment, and personalized medicine of lung cancer.
[0051] Moreover, this application does not require multiple statistical tests, overcoming the uncertainty of statistical inference. By detecting the genotypes and expression levels of these four genes and substituting them into a mathematical formula, it is possible to relatively accurately determine whether an individual has lung cancer, which subtype of lung cancer it belongs to, and the malignancy degree of lung cancer. This screening method has high sensitivity and specificity, can achieve precise identification of the gene components of lung cancer, and improve the cure rate of lung cancer.
[0052] This application specifically points out that under the conditions of the mathematical formula with 99.99% accuracy, the four genes screened out are confirmed as the internal dominant genes of lung cancer and play a decisive role: 1) Distinction between intrinsic master genes and extrinsic variables: These core genes directly determine whether an individual has lung cancer, while other clinical variables (such as age, gender, lifestyle habits, etc.) only serve as extrinsic factors. Although they may affect the prognosis and survival time of patients, they do not play a dominant role in the risk of developing cancer. This discovery subverts the traditional view and further emphasizes the unique position of core genes in cancer identification.
[0053] 2) Universality across races and sequencing technologies: Although slight differences in coefficients in the mathematical formula may be caused by different ethnic groups or gene sequencing methods, the diagnostic accuracy rate of this application always remains at 99.99%. This characteristic verifies the universality and stability of these four genes as intrinsic genes for lung cancer, perfectly reflecting the equivalence relationship between mathematical modeling and biological phenomena.
[0054] 3) Prediction ability for the carcinogenic stage and malignancy degree: This application can not only accurately diagnose lung cancer patients, but also predict the development stage and malignancy degree of carcinogenesis by judging the samples of adjacent tissues. Traditionally, the malignancy degree of cancer is mostly measured by the size or volume of the lesion, while the core gene model of this application provides a brand-new means for malignancy degree assessment based on the molecular level.
[0055] 4) Breakthrough recurrence prediction ability: Cancer recurrence usually stems from the lesions of adjacent tissues, and this application has an important breakthrough in the analysis of adjacent tissues, which can warn of potential recurrence risks, thus providing key guidance for clinical intervention and personalized treatment. This application realizes the perfect integration from mathematical formulas to biological mechanisms. The screening and prediction system constructed based on four core genes not only provides theoretical support for lung cancer diagnosis, but also opens up new paths in the fields of cancer staging, malignancy degree assessment and recurrence prediction. This achievement will lay a solid foundation for the early detection and precise treatment of lung cancer.
[0056] The effects of the four genes in predicting lung cancer are very significant. These four genes were found to have a sensitivity of nearly 99.99% and a specificity of 99.99% in the studies of adenocarcinoma of the lung ( LUAD ), and squamous cell carcinoma of the lung ( LUSC ). Previous studies have failed to find shared genes between adenocarcinoma of the lung and squamous cell carcinoma of the lung. This application is the first breakthrough to discover core shared genes in lung cancer research. At the same time, after introducing GPT 2, FAM 220 A 、 SGPL 1, PCOLCE 2, GABPB 1- IT 1 these five genes in non-small cell lung cancer ( NSCLC ), and ALK positive and EGFR / KRAS / ALK In the study of negative lung adenocarcinoma ( LUAD ), an accuracy of 97.8% to 99.99% was also shown.
[0057] In addition, in terms of treatment, the NLRC4, PLEKHN1, RASIP1, and SPP1 gene combination and its synergistic effect are used for targeted therapy of lung cancer. Through drug design targeting these four genes, precise strikes on lung cancer cells can be achieved, effectively inhibiting the growth and spread of lung cancer cells and improving the treatment effect of lung cancer. This targeted therapy method has advantages such as less side effects and remarkable curative effects, and can significantly improve the quality of life of lung cancer patients.
Claims
1. A gene combination, characterized in that: These genes include: NLRC4, PLEKHN1, RASIP1, and SPP1 。 2. A product for lung cancer screening and diagnosis, characterized in that: include: At least one specific probe or detection reagent for detecting NLRC4, PLEKHN1, RASIP1, and SPP1 genotype or expression level.
3. A product for targeted treatment of lung cancer, characterized in that: include: against NLRC4, PLEKHN1, RASIP1, and SPP1 Targeted drugs for one or more genes and their subtypes, and drug delivery devices for delivering the targeted drugs into a patient's body.
4. A method for lung cancer screening and diagnosis using the gene combination as claimed in claim 1, characterized in that: The following steps are involved: obtaining biological samples from subjects; Using specific probes or detection to actually detect the NLRC4, PLEKHN1, RASIP1, and SPP1 The genotype or expression level of the gene is used to obtain the test data; Analyzing the test data using a gene combination risk assessment mathematical model to identify lung cancer risk or diagnose lung cancer subtypes to obtain analysis results; Based on the analysis results, it is determined whether the subject has lung cancer and the specific subtype of lung cancer.
5. A method for targeted treatment of lung cancer using the gene combination as claimed in claim 1, characterized in that: The following steps are involved: Identify specific subtypes of lung cancer; Select appropriate targeted drugs according to the subtype of lung cancer; wherein the targeted drugs are: NLRC4, PLEKHN1, RASIP1 and SPP1 Drugs targeting one or more of the genes and their subtypes; The targeted drug is administered through a drug delivery device to achieve personalized treatment.
6. Use of the gene combination as claimed in claim 1 in lung cancer screening and diagnosis.
7. Use of the gene combination as claimed in claim 1 in targeted therapy for lung cancer.
8. The method for lung cancer screening and diagnosis according to claim 4, characterized in that: After obtaining the expression values of these genes, the expression of the gene combination risk assessment mathematical model is as follows: ; In the formula, Representation formula The corresponding risk assessment value is The values of are 1, 2, 3; is any real number, For the formula The corresponding intercept is , , and For the formula The corresponding coefficients are , , and represent the gene expression values of NLRC4, PLEKHN1, RASIP1, and SPP1, respectively; The maximum risk value is calculated using the following formula: : ; Based on the maximum risk value Calculate the probability of cancer : ; like , then it is judged as cancer, among which This is a pre-set judgment threshold based on the type of cancer.
9. The method for lung cancer screening and diagnosis according to claim 4, characterized in that: It also includes the following steps to build a four-dimensional visualization table: Selecting three specific genes as basic elements for constructing a three-dimensional space coordinate axis, and constructing a three-dimensional space coordinate system based on the basic elements; In the three-dimensional space coordinate system, different individuals will determine their specific positions in space based on the expression values of the three genes corresponding to them, so as to show the distribution of patients and healthy people in the three-dimensional space; At the same time, different colors are used to indicate the approximate risk of lung cancer. The intuitive difference in colors enables clinicians to quickly understand the risk level of lung cancer for different individuals. The three-dimensional space displays individual distribution and the risk probability of lung cancer indicated by color to form a four-dimensional visualization chart.
Citation Information
Patent Citations
SE10245C1