MRNAs marker and kit for early diagnosis of pulmonary signet-ring cell carcinoma

12 mRNA markers were discovered through RNA sequencing and machine learning methods, and the ultimate random tree risk score model was constructed, which solved the problem of insufficient sensitivity and specificity of early diagnosis of lung mark ring cell carcinoma in the prior art, and achieved high sensitivity and specificity early diagnosis.

CN120330331AInactive Publication Date: 2025-07-18HANGZHOU FIRST PEOPLES HOSPITAL +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510463064.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art is difficult to provide high sensitivity and specific biomarkers for the early diagnosis of lung marking ring cell carcinoma, there are problems with high false positives in imaging and pathological examinations, and insufficient sensitivity and specificity of immunohistochemistry examinations, which cannot meet the needs of early diagnosis.

Method used

The mRNA expression profiles of cancer tissues and control tissues of lung marker cell carcinoma patients were detected by RNA sequencing technology. Combined with statistical and machine learning methods, 12 differentially expressed mRNA markers FASN, RGPD2, ACSL1, ABCA3, MT2A, ITPKC, FKBP5, CEACAM6, HP, CAPN8, FCGBP and MUC16 were found, and the ultimate random tree risk score model was constructed for diagnosis.

Benefits of technology

The constructed risk scoring model is sensitive to early lung mark ring cell carcinoma diagnosis >80% and specificity >60%, improving the accuracy of early diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120330331A_ABST
    Figure CN120330331A_ABST
Patent Text Reader

Abstract

The invention discloses an mRNAs marker and a kit for early diagnosis of pulmonary signet-ring cell carcinoma. The marker is a combination of FASN, RGPD2, ACSL1, ABCA3, MT2A, ITPKC, FKBP5, CEACAM6, HP, CAPN8, FCGBP and MUC16, and is characterized in that the marker is a combination of FASN, RGPD2, ACSL1, ABCA3, MT2A, ITPKC, FKBP5, CEACAM6, HP, CAPN8 and MUC16. According to the invention, FFPE cancer tissue of a patient with primary pulmonary signet-ring cell carcinoma diagnosed by an existing detection means in clinic is used as a sample, para-carcinoma tissue is used as a control sample, mRNA expression profiles of the cancer tissue and the control tissue are detected by adopting an RNA sealing technology, and 12 differentially expressed mRNA markers are found by applying statistics and machine learning methods. Based on the found 12 markers, a risk scoring model with high sensitivity and specificity is constructed, and the sensitivity gt of the model to early-stage lung signet-ring cell carcinoma diagnosis; the specificity is gt; 60%.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of biological detection, and particularly relates to mRNAs markers and kits for early diagnosis of pulmonary signet-ring cell carcinoma. Background Art

[0002] Signet-ring cell carcinoma (SRCC) is a special type of mucus-secreting adenocarcinoma. It is named because under the microscope, the cytoplasm of the cells is rich and filled with mucus, and the nucleus is squeezed to one side of the cytoplasm, showing a "signet ring" shape. It often occurs in the gastrointestinal tract, breast, bladder, prostate and other parts, while primary SRCC in the lungs is very rare. When the World Health Organization (WHO) released the classification of pulmonary tumors, pathology and genetics in 2004, primary pulmonary signet-ring cell carcinoma was classified as a subtype of solid adenocarcinoma that can secrete mucus in the lungs. In the latest international multidisciplinary classification standard for lung adenocarcinoma, primary pulmonary signet-ring cell carcinoma is no longer regarded as an independent subtype of lung adenocarcinoma, but only a cytological feature. The pathogenesis of pulmonary signet-ring cell carcinoma is not yet clear, but its malignancy is extremely high. Due to the strong concealment of clinical symptoms and the rapid progression of the disease, most cases are found in the advanced stage, missing the best opportunity for radical surgical treatment, and the 5-year survival rate is almost zero. However, the 5-year survival rate of early-stage patients can reach 55% after complete surgical resection. Therefore, early screening and early diagnosis of pulmonary signet-ring cell carcinoma are crucial for improving the prognosis of patients.

[0003] Currently, the methods used for the diagnosis of pulmonary signet-ring cell carcinoma in clinical practice mainly include imaging examinations (chest CT scan, PET-CT, etc.), pathological examinations, immunohistochemical examinations, etc. However, these diagnostic methods all have certain limitations. The CT manifestations of pulmonary signet-ring cell carcinoma are mostly peripheral lung cancers, which can be distributed in each lung lobe. The lesions are round or oval, mostly with a large diameter, uneven edges, clear boundaries, and signs of lobulation and short spicules. Some have pleural indentation signs. Imaging examinations have the problem of high false positive rate and need to be differentiated from other types of lung cancers, and finally need to be confirmed by histopathology. Pathological examinations such as bronchoscopy also have the problem of high false positive rate. Cells with a "signet ring" appearance not only appear in signet-ring cell carcinoma, but also in malignant lymphoma, and even in some benign diseases such as pseudomembranous colitis. Immunohistochemical examinations such as markers NapsinA (novel aspartic protease A), TTF-1 (thyroid transcription factor-1), and CK7 / CK20, etc., have poor sensitivity or specificity for the diagnosis of pulmonary signet-ring cell carcinoma and cannot meet the needs of early diagnosis. It is difficult to distinguish primary pulmonary signet-ring cell carcinoma from metastatic carcinomas of other organs, especially when the cell morphology is atypical or the cell proportion is small and easy to be ignored. Developing highly specific biomarkers to improve the diagnostic ability of pulmonary signet-ring cell carcinoma is an urgent clinical need.

[0004] At present, there are few reports on the research of molecular-level biomarkers in the early diagnosis of pulmonary signet ring cell carcinoma. There are two relatively relevant research literatures: Literature 1 (PMID: 39910169): Researchers found the activation of ATR signal and the upregulation of DNA repair proteins such as FANCD2 and RECQL5 in LSRCC through deep vision proteomics, suggesting that replication stress can be used as a potential diagnostic biomarker. However, this study was only based on one patient, and the universality of its results needs to be further verified. Literature 2 (PMID: 20022810) focused on gastric signet ring cell carcinoma. Researchers analyzed 353 gastric samples from two independent patient subgroups in Japan through microRNA microarray, compared the microRNA expression patterns between non-tumor mucosa and cancer samples. Among 160 non-tumor mucosa and cancer paired samples, 22 microRNAs were upregulated and 13 microRNAs were downregulated in gastric cancer, and 292 (83%) samples were correctly distinguished by this feature. However, this study mainly evaluated the relationship between microRNA expression and the progression and prognosis of gastric cancer, and did not further explore the potential of differentially expressed genes as diagnostic biomarkers.

[0005] Currently, there is no reported biomarker for the early detection of pulmonary signet ring cell carcinoma that has both high sensitivity and high specificity. How to develop a biomarker for the early diagnosis of pulmonary signet ring cell carcinoma has become one of the urgent problems to be solved in this field. Summary of the Invention

[0006] The purpose of the present invention is to provide an mRNAs marker and a kit for the early diagnosis of pulmonary signet ring cell carcinoma.

[0007] The technical solutions adopted by the present invention to achieve the above purpose are as follows:

[0008] The present invention provides an mRNAs marker for the early diagnosis of pulmonary signet ring cell carcinoma, and the mRNAs marker is a combination of FASN, RGPD2, ACSL1, ABCA3, MT2A, ITPKC, FKBP5, CEACAM6, HP, CAPN8, FCGBP and MUC16.

[0009] The present invention also provides the application of a detection reagent for the expression level of the mRNAs marker in the preparation of a composition or a kit for evaluating, diagnosing or monitoring pulmonary signet ring cell carcinoma.

[0010] On the other hand, the present invention provides a kit for the early diagnosis of pulmonary signet ring cell carcinoma, and the kit includes a reagent for detecting the expression level of the mRNAs marker.

[0011] As a specific implementation, the expression values of the mRNAs markers in cancer tissues and control tissues are obtained by RNA sequencing technology for the reagent.

[0012] As a specific implementation, the kit further includes a risk scoring model, and the formula of the model is as follows.

[0013]

[0014] X i represents the marker expression value of sample i, and P(X i ) is the probability of the corresponding classification of sample i predicted by the diagnostic model, where 0 represents a healthy individual and 1 represents a patient with pulmonary signet-ring cell carcinoma; the sum of the probabilities of all classifications is equal to 1, and the classification with the maximum probability is taken as the final prediction result of sample i.

[0015] As a specific implementation, the model algorithm adopted by the risk scoring model is the extreme random tree.

[0016] The present invention also provides a method for diagnosing early pulmonary signet-ring cell carcinoma using the kit described above, including the following steps:

[0017] Step 1: Collect cancer tissues of patients suspected of having pulmonary signet-ring cell carcinoma by clinical diagnosis, and use RNA sequencing to obtain the expression levels of mRNAs markers in the cancer tissues;

[0018] Step 2: Substitute the obtained expression values of the mRNAs markers into the risk scoring model formula of claim 5 to calculate the probabilities of the patient being predicted as healthy and having pulmonary signet-ring cell carcinoma;

[0019] Step 3: Select the output result with the highest probability and give the prediction result of the risk of the patient suffering from pulmonary signet-ring cell carcinoma.

[0020] The English abbreviation annotations involved in the present invention are as follows:

[0021] FASN: Fatty acid synthase

[0022] RGPD2: Containing RANBP2-like and GRIP domains 2

[0023] ACSL1: Long-chain acyl-CoA synthetase 1

[0024] ABCA3: ATP-binding cassette subfamily A member 3

[0025] MT2A: Metallothionein 2A

[0026] ITPKC: Inositol 1,4,5-trisphosphate 3-kinase C

[0027] FKBP5: FKBP prolyl isomerase 5

[0028] CEACAM6: Carcinoembryonic antigen-related cell adhesion molecule 6

[0029] HP: Haptoglobin

[0030] CAPN8: Calpain 8

[0031] FCGBP: Fcγ-binding protein

[0032] MUC16: Mucin 16

[0033] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0034] 1. The present invention uses RNA sequencing technology to detect the mRNA expression profiles of cancer tissues and control tissues of patients with pulmonary signet-ring cell carcinoma, and applies statistical and machine learning methods to discover 12 differentially expressed mRNA markers, namely FASN, RGPD2, ACSL1, ABCA3, MT2A, ITPKC, FKBP5, CEACAM6, HP, CAPN8, FCGBP, and MUC16.

[0035] 2. Based on the 12 markers discovered, the present invention constructs a risk scoring model with high sensitivity and specificity. The sensitivity of the model for the diagnosis of early pulmonary signet-ring cell carcinoma is >80%, and the specificity is >60%. Brief Description of the Drawings

[0036] Figure 1 It is the immunohistochemical staining result diagram in Example 1.

[0037] Figure 2 It is the ROC curve of the training cohort of the extremely randomized tree model algorithm in Example 3.

[0038] Figure 3 It is the ROC curve of the validation cohort of the extremely randomized tree model algorithm in Example 3. Detailed Embodiments

[0039] The technical solutions of the present invention will be described in detail below in conjunction with the embodiments. The reagents and biological materials used below are all commercial products unless otherwise specified.

[0040] Example 1: Screening of mRNA Markers

[0041] (1) Research Cohort and Clinical Information

[0042] A total of 23 cases of cancer tissues from patients with pulmonary signet-ring cell carcinoma were included, and another 17 cases of adjacent tissues were collected as healthy controls. Finally, a total of 40 tissue samples were included in the study. Among them, 28 samples were randomly selected as the training cohort, including 16 cases of pulmonary signet-ring cell carcinoma tissues and 12 cases of adjacent tissues. The validation cohort consisted of 12 cases, including 7 cases of cancer tissues and 5 cases of adjacent tissues. The specific grouping information is shown in Table 1.

[0043] Table 1

[0044]

[0045] (2) Preparation and sectioning of FFPE samples

[0046] The preparation of FFPE samples is completed in two parts: formaldehyde fixation and paraffin embedding. Tissue samples of appropriate size are placed in 10% neutral buffered formalin (also known as formaldehyde) fixative and fixed at room temperature for 18 - 24 h. After fixation, the samples are taken out and rinsed with running water for several minutes to remove residual fixative and possible crystals. After the tissue samples are fixed, the water in them needs to be removed to facilitate the full penetration of the embedding medium. Gradient dehydration with alcohol: soak in 75% ethanol, 85% ethanol, 95% ethanol, absolute ethanol (I), and absolute ethanol (II) for 1 min each. After dehydration, soak in xylene twice, 1 min each time, and then place in melted paraffin. The wax-impregnated tissue samples are placed in a mold, and melted paraffin is poured in, and then the paraffin is allowed to cool and solidify. In this way, the tissue samples are embedded in paraffin.

[0047] The fully cooled embedded tissue blocks are fixed on the sample holder of the microtome, and the section thickness of the microtome is adjusted, usually 4 - 5 μm thick. The cut tissue sections are gently removed from the knife with a brush and spread in warm water (40 °C). After the spread sections are firmly attached to the glass slides, they are baked at 60 °C for 2 h.

[0048] (3) Immunohistochemical detection

[0049] The sections were dewaxed with xylene and then successively placed in absolute ethanol (I), absolute ethanol (II), 95% ethanol, 85% ethanol, 70% ethanol and pure water, each immersion for 5 min. Antigen retrieval was performed with 10 mM sodium citrate (pH 6.0), and then the sections were treated with 3% hydrogen peroxide aqueous solution at room temperature for 10 min and blocked with 5% BSA at room temperature for 30 min. Primary antibody information: Napsin A (ab73021, abcam), CK7 (ab181598, abcam) and TTF-1 (ab76013, abcam). The antibodies were incubated at room temperature for 1 h, washed thoroughly with TBST, and incubated with rabbit secondary antibody (A0208, Beyotime) or mouse secondary antibody (A0216, Beyotime) at room temperature for 30 min, washed four times with TBST for 10 min each time, sealed after color development, and observed under a microscope for the location and expression level of the antigen in cells or tissues. See Figure 1 , which is the immunohistochemical staining result diagram. The staining results showed that Napsin A and CK7 presented obvious brown or yellowish-brown staining. Specifically, diffuse brown granular precipitates were visible in the cytoplasm, while there was no obvious staining in the intercellular matrix, indicating that Napsin A and CK7 were expressed in the cytoplasm of the tissue cells. TTF1 presented obvious yellowish-brown staining and was mainly expressed in the nucleus. Figure 1 The results showed that the alveolar epithelial markers Napsin A, CK7 and TTF1 in the patient's tumor cells were all positive, strongly suggesting that the patient had lung adenocarcinoma.

[0050] (4) RNA extraction from FFPE samples

[0051] Refer to the product manual, and use ReadiMag FFPE RNA Kit (3DMed, Shanghai) to extract total RNA from FFPE samples, and finally elute the RNA with 60 μL RNase-free Water. Use NanoDrop to detect the purity of RNA, Qubit to quantify the RNA concentration, and use Agilent 2100 bioanalyzer and the supporting RNA analysis kit (5067-1514, Agilent) to detect the fragment distribution of mRNA (evaluate the size of the DV200 value).

[0052] (5) Detection of mRNA expression level

[0053] Refer to the product manual, use the KAPA RiboErase HMR kit (KK8481, Kapa) to remove ribosomal RNA, and use the KAPA RNA Hyper kit (KK8542, Kapa) for mRNA library construction. Load 400 ng of total RNA for each sample, and then perform rRNA removal, DNase digestion, first and second strand synthesis, adapter ligation, and library amplification. Use Agencourt Ampure XP-PCR purification beads (A63881, Beckman) to purify the PCR enrichment product, and finally elute the library DNA with 20 μL of nuclease-free water. Use the Invitrogen Qubit 4.0 fluorometer and the accompanying reagent Qubit dsDNA HS Assay Kit (Q32854, Thermofisher) to quantify the DNA concentration. Use the Agilent 4150 Bioanalyzer and the accompanying chip and reagent HighSensitivity D1000 ScreenTape & Reagents (5067-5584 & 5067-5585, Agilent) to detect the distribution of library DNA fragments. Sequence using the Illumina NovaSeq platform with a sequencing strategy of PE150 and a sequencing data volume of 12 G for each library.

[0054] (6) Sequencing data analysis process

[0055] Based on the RNA sequencing detection technology, obtain the expression levels of mRNA in the cancer tissues and adjacent tissues of patients with pulmonary signet ring cell carcinoma. The analysis process of the sequencing data is as follows:

[0056] Sequencing data alignment. After removing the sequencing adapters from the RNA sequencing data using the fastp software (version: 0.23.4), use the STAR software (version: 2.7.11b) to align the sequencing data to the human reference genome hg38 (genome download link: http: / / hgdownload.soe.ucsc.edu / goldenPath / hg38 / bigZips / ).

[0057] mRNA expression quantification. Use the RSEM software (version: 1.3.3) to quantitatively count the number of reads aligned to different mRNAs.

[0058] mRNA annotation. Use the Gencode v43 database to annotate the mRNA, and retain the mRNAs annotated as known for subsequent analysis.

[0059] mRNA filtering. For the training cohort, mRNAs that aligned to at least 10 reads in all training cohort samples were retained for subsequent analysis. For the validation cohort, the mRNAs selected from the training cohort were retained for subsequent analysis.

[0060] mRNA expression normalization. The original mRNA expression levels of the training cohort samples were normalized using the trimmed mean of M-values (TMM) method, and the validation cohort samples were processed with the same parameters.

[0061] (7) Discovery of biomarkers

[0062] Samples were grouped according to the pathological examination results. Based on the mRNA expression levels in the training cohort, statistical and machine learning methods were used to discover mRNAs that could distinguish patients with pulmonary signet-ring cell carcinoma from healthy controls as biomarkers. The process is as follows:

[0063] 1) Grouping of the training cohort. According to the pathological examination results of the samples, the samples in the training cohort were divided into two groups: healthy controls and pulmonary signet-ring cell carcinoma.

[0064] 2) Screening of biomarkers using statistical methods. The U-test and T-test statistical methods were used to calculate the significance of the differences in the expression of all mRNAs between the two groups. Statistical significance was evaluated using the P-value, and a P-value ≤ 0.05 was considered statistically significant. First, mRNAs with both median and mean expression levels greater than 5 in the two groups were selected. Subsequently, mRNAs that met at least one of the following two screening criteria were retained for subsequent analysis: (a) According to the statistical test results of the U-test, the absolute value of the logarithm of the fold change in median expression (logFC median) > 1, and the U-test P-value ≤ 0.05; (b) According to the statistical test results of the T-test, the absolute value of the logarithm of the fold change in mean expression (logFC mean) > 1, and the T-test P-value ≤ 0.05.

[0065] 3) Screening biomarkers using machine learning methods. To determine the final markers for constructing a risk scoring model, multiple machine learning algorithms were used for biomarker screening, including 4 machine learning algorithms: Linear Model, random forest, extremely randomized trees, and Ensemble methods. The main process is as follows: (a) Use the mRNA screened in 2) above as the initial biomarker and evaluate the baseline performance of the biomarker; (b) Evaluate the feature importance of each biomarker in the model, retain the biomarkers with importance > 0, and retrain the model; (c) Evaluate the feature importance of the biomarkers in the new model again, and repeat the process of (b) until the feature importance of the biomarkers is > 0 in each model. These biomarkers were used for subsequent analysis. The classification effect and feature importance were trained and evaluated on the training set data, and 12 biomarkers were screened, specifically FASN, RGPD2, ACSL1, ABCA3, MT2A, ITPKC, FKBP5, CEACAM6, HP, CAPN8, FCGBP, and MUC16.

[0066] Example 2: Construction of risk scoring model

[0067] Taking healthy individuals and lung signet ring cell carcinoma as the classification prediction targets, using the biomarkers discovered in Example 1, 4 machine learning algorithms, namely Extremely Randomized Trees, Linear Model, Random Forest, and Ensemble methods, were used. Different hyper-parameters were preset for each algorithm, and 6 diagnostic models of different machine learning algorithms were trained on the entire training cohort sample data. The formula for the final risk scoring model is as follows:

[0068]

[0069] X i represents the biomarker expression value of sample i, and P(X i ) is the probability of the classification corresponding to sample i predicted by the diagnostic model, where 0 represents a healthy individual and 1 represents a patient with lung signet ring cell carcinoma. The sum of the probabilities of all classifications is equal to 1, and the classification with the maximum probability is taken as the final prediction result of sample i. The trained model is saved on the hard disk in the form of a file, and the model prediction result can be obtained by inputting the biomarker expression value of the sample when calling the model.

[0070] Example 3: Evaluation and validation of the performance of the risk scoring model

[0071] In the training cohort, the 5-fold cross-validation method was used to evaluate the classification performance and feature weights of the model, and the generalization performance of the model was evaluated in the validation cohort. The model evaluation and validation metrics included the area under the receiver operating characteristic curve (AUC, ranging from 0 to 1), accuracy (ranging from 0 to 1), positive predictive value (ranging from 0 to 1), negative predictive value (ranging from 0 to 1). The specificity (ranging from 0 to 1) and sensitivity (ranging from 0 to 1) of the model were evaluated, and higher values indicated better classification performance of the model.

[0072] For the 12 markers FASN, RGPD2, ACSL1, ABCA3, MT2A, ITPKC, FKBP5, CEACAM6, HP, CAPN8, FCGBP, and MUC16, the evaluation results of individual markers in the training cohort and the validation cohort are shown in Table 2.

[0073] Table 2

[0074]

[0075] The data results in Table 2 showed that the evaluation effects of individual markers were poor and the generalization performance was relatively poor. A high value in either the training cohort alone or the validation cohort alone indicated poor generalization performance. It was better that the AUC values were close in the training cohort and the validation cohort.

[0076] The 12 markers were combined, and the 5-fold cross-validation method was used to evaluate the classification performance and feature weights of the model in the training cohort. Combining the pathological examination results, the prediction results of the risk of pulmonary signet-ring cell carcinoma for each patient were given according to the model formula and program. The model evaluation metrics included the area under the receiver operating characteristic curve (AUC, ranging from 0 to 1), accuracy (ranging from 0 to 1), positive predictive value (ranging from 0 to 1), negative predictive value (ranging from 0 to 1), sensitivity (ranging from 0 to 1), and specificity (ranging from 0 to 1). The evaluation results of the combination of 12 markers in the training cohort are shown in Table 3. The results showed that in the training cohort, this risk score model had high AUC, accuracy, positive predictive value, negative predictive value, specificity, and sensitivity, and the model had excellent prediction performance. Among them, the model algorithm using extremely randomized trees had the best prediction performance. See Figure 2 , which is the ROC curve of the training cohort for the extremely randomized tree model algorithm.

[0077] Table 3

[0078]

[0079] To verify the performance of the risk score model in predicting pulmonary signet-ring cell carcinoma, another independent cohort was selected as the validation cohort to evaluate the classification effect and feature weights of the model. Using the pathological test results as the ground truth, the model evaluation metrics included AUC (value range 0-1), accuracy (value range 0-1), positive predictive value (value range 0-1), negative predictive value (value range 0-1), sensitivity (value range 0-1), and specificity (value range 0-1). The evaluation results of the combined use of 12 markers in the validation cohort are shown in Table 4. The results showed that in the validation cohort, this risk score model had high AUC, accuracy, positive predictive value, negative predictive value, specificity, and sensitivity, and the model had excellent predictive performance. See Figure 3 , which is the ROC curve of the validation cohort for the extreme random tree model algorithm.

[0080] Table 4

[0081]

[0082]

[0083] Example 4: Application of the risk score model

[0084] 1) Collect the cancer tissues of patients with clinically diagnosed suspected pulmonary signet-ring cell carcinoma, and use RNA sequencing to obtain the expression levels of biomarkers in the cancer tissues;

[0085] 2) Substitute the obtained expression values of FASN, RGPD2, ACSL1, ABCA3, MT2A, ITPKC, FKBP5, CEACAM6, HP, CAPN8, FCGBP, and MUC16 into the trained model formula to calculate the probabilities of each patient being predicted as healthy and having pulmonary signet-ring cell carcinoma;

[0086] 3) Select the output result with the highest probability and give the prediction result of the risk of each patient having pulmonary signet-ring cell carcinoma.

[0087] In summary, the present invention uses RNA sequencing technology to detect the mRNA expression profiles of cancer tissues and healthy control tissues of patients with pulmonary signet-ring cell carcinoma, and applies statistical and machine learning methods to discover 12 differentially expressed mRNA markers, namely FASN, RGPD2, ACSL1, ABCA3, MT2A, ITPKC, FKBP5, CEACAM6, HP, CAPN8, FCGBP, and MUC16. Based on the 12 discovered markers, a risk score model with high sensitivity and specificity was constructed. The sensitivity of the model for the diagnosis of early pulmonary signet-ring cell carcinoma > 80%, and the specificity > 60%.

[0088] The above are only some preferred embodiments of the present invention, and the present invention is not limited to the content of the embodiments. For those skilled in the art, various changes and modifications can be made within the scope of the concept of the technical solution of the present invention, and any changes and modifications made are within the protection scope of the present invention.

Claims

1. mRNAs markers for early diagnosis of pulmonary signet ring cell carcinoma, characterized in that: The mRNAs markers are a combination of FASN, RGPD2, ACSL1, ABCA3, MT2A, ITPKC, FKBP5, CEACAM6, HP, CAPN8, FCGBP, and MUC16.

2. Use of a detection reagent for obtaining the expression level of the mRNAs markers described in claim 1 in the preparation of a composition or kit for evaluating, diagnosing, or monitoring pulmonary signet-ring cell carcinoma.

3. A kit for the early diagnosis of pulmonary signet-ring cell carcinoma, characterized in that: The kit includes a reagent for detecting the expression level of the mRNAs markers described in claim 1.

4. The kit for early diagnosis of pulmonary signet ring cell carcinoma according to claim 3, wherein: The reagent respectively obtains the expression values of the mRNAs markers in cancer tissue and control tissue through RNA sequencing technology.

5. The kit for early diagnosis of pulmonary signet ring cell carcinoma according to claim 3, characterized in that: The kit further includes a risk scoring model, and the formula of the model is as follows. Denote the biomarker expression value of sample i, which is the probability of the classification corresponding to sample i predicted by the diagnostic model, where 0 represents healthy individuals and 1 represents patients with pulmonary signet ring cell carcinoma; the sum of the probabilities of all classifications is equal to 1, and the classification with the largest probability is taken as the final prediction result of sample i.

6. The kit for early diagnosis of pulmonary signet ring cell carcinoma according to claim 5, characterized in that: The model algorithm used in the risk scoring model is the extreme random tree.

7. A method for diagnosing early pulmonary signet-ring cell carcinoma using the kit described in claim 3, comprising the following steps: Step 1, collect cancer tissue from a patient with a clinical diagnosis suspected of pulmonary signet-ring cell carcinoma, and use RNA sequencing to obtain the expression level of the mRNAs markers in the cancer tissue; Step 2, substitute the obtained expression values of the mRNAs markers into the risk scoring model formula described in claim 5 to calculate the probabilities of the patient being predicted as healthy and having pulmonary signet-ring cell carcinoma; Step 3, select the output result with the highest probability and give a prediction result of the risk of the patient having pulmonary signet-ring cell carcinoma.