Markers for detecting lung cancer and their uses and systems
By using specific DNA methylation markers and second-generation sequencing technology, cfDNA methylation levels in blood samples are analyzed, and the sensitivity and specificity of existing lung cancer detection methods are insufficient, achieving early and non-invasive and efficient lung cancer screening and diagnosis.
Patent Information
- Application Number
- CN202111188052.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-12
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2041-10-12
AI Technical Summary
The existing lung cancer detection methods are insufficient in terms of sensitivity and specificity, making it difficult to achieve early and non-invasive efficient screening and diagnosis.
Lung cancer is detected by analyzing cfDNA methylation levels in blood samples using one or more specific DNA methylation markers, including CG site sequences located on different chromosomes, combined with bisulfite treatment and second-generation sequencing techniques.
Early and non-invasive detection of lung cancer is achieved, the sensitivity and accuracy of the detection is improved, and the development of cancer can be monitored in real time.
Smart Images

Figure SMS_2 
Figure SMS_3 
Figure SMS_4
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biotechnology, and in particular to a marker for detecting lung cancer, and a use and system thereof. Background Art
[0002] Lung cancer is one of the most common cancers with the highest morbidity and mortality rates worldwide. Conventional lung cancer screening methods include low-dose spiral CT (LDCT) and protein markers such as carcinoembryonic antigen (CEA), squamous cell carcinoma antigen (SCC), and neuron-specific enolase (NSE). However, these methods vary in sensitivity and specificity. Currently, DNA methylation has been shown to be tissue-specific and can be used for early cancer detection. The methylation signature of circulating tumor DNA (ctDNA) can also be used to track the primary tumor site.
[0003] cfDNA (cell-free DNA) is a small fragment of DNA that is free in peripheral blood. It originates from normal cells or tumor cells and metabolism, and contains genetic information such as somatic mutations and DNA methylation. Among them, ctDNA methylation can be detected early in tumor development and has good stability, making ctDNA methylation a highly valuable marker in tumor liquid biopsies. BRIEF DESCRIPTION OF THE DRAWINGS
[0004] Figure 1 This is a flow chart of the biomarker research of the present invention (N represents healthy samples, and C represents lung cancer samples).
[0005] Figure 2 Figure 3 is the ROC curve of the model constructed based on 5 markers in the training set.
[0006] Figure 3 The ROC curve of the model constructed based on 5 markers in the test set.
[0007] Figure 4 ROC curve of the test set for the model constructed based on two of the five markers.
[0008] Figure 5 ROC curve of the model constructed based on 3 out of 5 markers in the test set.
[0009] Figure 6 ROC curve of the model constructed based on 4 out of 5 markers in the test set. Summary of the Invention
[0010] The object of the present invention is to provide a marker for detecting lung cancer, its use and system.
[0011] The specific technical solutions of the present invention are as follows:
[0012] 1. A marker for detecting lung cancer, characterized in that the marker comprises one or more of the following methylation markers:
[0013] The first marker comprises more than 80%, preferably more than 90%, and preferably all CG sites in the sequence identical to or reverse complementary to the nucleotide sequence at position 151445000-151450000 of chromosome 1.
[0014] The second marker comprises more than 80%, preferably more than 90%, and preferably all CG sites in the sequence identical to or reverse complementary to the nucleotide sequence at positions 191183500-191188500 of chromosome 2.
[0015] The third marker comprises more than 80%, preferably more than 90%, and preferably all CG sites in the sequence identical or reverse complementary to the nucleotide sequence located at position 191184000-191189000 of chromosome 2,
[0016] A fourth marker comprising more than 80%, preferably more than 90%, and preferably all of the CG sites in the sequence identical or reverse complementary to the nucleotide sequence located at position 68566500-68571500 of chromosome 4, and
[0017] The fifth marker includes more than 80%, preferably more than 90%, of all CG sites in the sequence that is identical to or reverse complementary to the nucleotide sequence located at positions 30601500-30606500 of chromosome 11, and preferably includes all CG sites.
[0018] 2. The marker according to item 1, characterized in that
[0019] The first marker includes more than 80%, preferably more than 90%, and preferably all CG sites in the nucleotide sequence shown in SEQ ID NO: 1 or the reverse complementary sequence of the sequence shown in SEQ ID NO: 1.
[0020] The second marker includes more than 80%, preferably more than 90%, and preferably all CG sites in the nucleotide sequence shown in SEQ ID NO: 2 or the reverse complementary sequence of the sequence shown in SEQ ID NO: 2.
[0021] The third marker includes more than 80%, preferably more than 90%, and preferably all CG sites in the nucleotide sequence shown in SEQ ID NO: 3 or the reverse complementary sequence of the sequence shown in SEQ ID NO: 3.
[0022] The fourth marker includes more than 80%, preferably more than 90%, and preferably all CG sites in the nucleotide sequence shown in SEQ ID NO: 4 or the reverse complementary sequence of the sequence shown in SEQ ID NO: 4, and
[0023] The fifth marker includes more than 80%, preferably more than 90%, of all CG sites in the nucleotide sequence shown in SEQ ID NO: 5 or the reverse complementary sequence of the sequence shown in SEQ ID NO: 5, and preferably includes all CG sites.
[0024] 3. A first marker comprising more than 80%, preferably more than 90%, and preferably all, of the CG sites in the sequence identical or reverse complementary to the nucleotide sequence at positions 151445000-151450000 of chromosome 1; or
[0025] A second marker comprising more than 80%, preferably more than 90%, and preferably all, of the CG sites in the sequence identical to or reverse complementary to the nucleotide sequence at positions 191183500-191188500 of chromosome 2; or
[0026] A third marker comprising more than 80%, preferably more than 90%, and preferably all, of the CG sites in the sequence identical or reverse complementary to the nucleotide sequence located at position 191184000-191189000 of chromosome 2; or
[0027] A fourth marker comprising more than 80%, preferably more than 90%, and preferably all, of the CG sites in the sequence identical or reverse complementary to the nucleotide sequence located at position 68566500-68571500 of chromosome 4; or
[0028] A fifth marker comprising more than 80%, preferably more than 90%, and preferably all of the CG sites in the sequence identical or reverse complementary to the nucleotide sequence located at positions 30601500-30606500 of chromosome 11,
[0029] Use in detecting lung cancer or in preparing a kit for detecting lung cancer.
[0030] 4. The use according to item 3, characterized in that
[0031] The first marker includes more than 80%, preferably more than 90%, and preferably all CG sites in the nucleotide sequence shown in SEQ ID NO: 1 or the reverse complementary sequence of the sequence shown in SEQ ID NO: 1.
[0032] The second marker includes more than 80%, preferably more than 90%, and preferably all CG sites in the nucleotide sequence shown in SEQ ID NO: 2 or the reverse complementary sequence of the sequence shown in SEQ ID NO: 2.
[0033] The third marker includes more than 80%, preferably more than 90%, and preferably all CG sites in the nucleotide sequence shown in SEQ ID NO: 3 or the reverse complementary sequence of the sequence shown in SEQ ID NO: 3.
[0034] The fourth marker includes more than 80%, preferably more than 90%, and preferably all CG sites in the nucleotide sequence shown in SEQ ID NO: 4 or the reverse complementary sequence of the sequence shown in SEQ ID NO: 4, and
[0035] The fifth marker includes more than 80%, preferably more than 90%, of all CG sites in the nucleotide sequence shown in SEQ ID NO: 5 or the reverse complementary sequence of the sequence shown in SEQ ID NO: 5, and preferably includes all CG sites.
[0036] 5. A system for detecting lung cancer, characterized in that the system comprises:
[0037] A sample collection module, which is used to collect samples from subjects;
[0038] A sample processing module, which is used to extract cfDNA from the sample and perform methylation treatment for sequencing; preferably, the methylation treatment is performed using bisulfite;
[0039] A sequencing module, which sequences the samples processed by the sample processing module;
[0040] an analysis module for analyzing the sequencing results to determine the comprehensive methylation levels of the methylation markers in the sample;
[0041] The determination module is used to determine whether the subject has lung cancer based on the comprehensive methylation level obtained by analysis.
[0042] 6. The system according to item 5, characterized in that
[0043] The markers include one or more of the following methylation markers:
[0044] The first marker comprises more than 80%, preferably more than 90%, and preferably all CG sites in the sequence identical to or reverse complementary to the nucleotide sequence at position 151445000-151450000 of chromosome 1.
[0045] The second marker comprises more than 80%, preferably more than 90%, and preferably all CG sites in the sequence identical to or reverse complementary to the nucleotide sequence at positions 191183500-191188500 of chromosome 2.
[0046] The third marker comprises more than 80%, preferably more than 90%, and preferably all CG sites in the sequence identical or reverse complementary to the nucleotide sequence located at position 191184000-191189000 of chromosome 2,
[0047] A fourth marker comprising more than 80%, preferably more than 90%, and preferably all of the CG sites in the sequence identical or reverse complementary to the nucleotide sequence located at position 68566500-68571500 of chromosome 4, and
[0048] The fifth marker includes more than 80%, preferably more than 90%, of all CG sites in the sequence that is identical to or reverse complementary to the nucleotide sequence located at positions 30601500-30606500 of chromosome 11, and preferably includes all CG sites.
[0049] 7. The system according to item 6, characterized in that
[0050] The first marker includes more than 80%, preferably more than 90%, and preferably all CG sites in the nucleotide sequence shown in SEQ ID NO: 1 or the reverse complementary sequence of the sequence shown in SEQ ID NO: 1.
[0051] The second marker includes more than 80%, preferably more than 90%, and preferably all CG sites in the nucleotide sequence shown in SEQ ID NO: 2 or the reverse complementary sequence of the sequence shown in SEQ ID NO: 2.
[0052] The third marker includes more than 80%, preferably more than 90%, and preferably all CG sites in the nucleotide sequence shown in SEQ ID NO: 3 or the reverse complementary sequence of the sequence shown in SEQ ID NO: 3.
[0053] The fourth marker includes more than 80%, preferably more than 90%, and preferably all CG sites in the nucleotide sequence shown in SEQ ID NO: 4 or the reverse complementary sequence of the sequence shown in SEQ ID NO: 4, and
[0054] The fifth marker includes more than 80%, preferably more than 90%, of all CG sites in the nucleotide sequence shown in SEQ ID NO: 5 or the reverse complementary sequence of the sequence shown in SEQ ID NO: 5, and preferably includes all CG sites.
[0055] 8. The system according to item 6 or 7, characterized in that, in the analysis module, the comprehensive methylation level of the methylation marker refers to the one calculated based on the methylation level of one or more selected methylation markers.
[0056] 9. The system according to item 8 is characterized in that, in the analysis module, the methylation level of each methylation marker is calculated based on the methylation level of each CG site in the methylation marker, wherein the methylation level of the CG site is the ratio of the cytosine detected as methylated at the site to the sum of the cytosine detected as methylated and the cytosine not methylated in all sequence results detected for the site.
[0057] 10. The system according to item 9, characterized in that, in the determination module, the comprehensive methylation level is determined based on the methylation levels of one, two, three, four, or five of the first marker, the second marker, the third marker, the fourth marker, or the fifth marker,
[0058] wherein the threshold is determined based on methylation level detection data of one, two, three, four, or five of the first marker, the second marker, the third marker, the fourth marker, or the fifth marker in a given lung cancer subject or a population of healthy subjects;
[0059] If the comprehensive methylation marker is greater than a specified threshold, the subject is determined to have lung cancer, and if the comprehensive methylation level marker is less than or equal to the specified threshold, the subject is determined to be normal.
[0060] 11. The system according to item 5, wherein the subject sample is a blood sample.
[0061] 12. A method for detecting lung cancer, comprising:
[0062] a sample collection step, which collects a sample from a subject;
[0063] a sample processing step, which extracts cfDNA from the sample and performs a methylation treatment for sequencing; preferably, the methylation treatment is performed using bisulfite;
[0064] a sequencing step, which sequences the sample processed in the sample processing step;
[0065] an analysis step of analyzing the sequencing results to determine the comprehensive methylation levels of the methylation markers in the sample;
[0066] The determination step is to determine whether the subject has lung cancer based on the comprehensive methylation level obtained by analysis.
[0067] 13. The method according to claim 12, characterized in that
[0068] The markers include one or more of the following methylation markers:
[0069] The first marker comprises more than 80%, preferably more than 90%, and preferably all CG sites in the sequence identical to or reverse complementary to the nucleotide sequence at position 151445000-151450000 of chromosome 1.
[0070] The second marker comprises more than 80%, preferably more than 90%, and preferably all CG sites in the sequence identical to or reverse complementary to the nucleotide sequence at positions 191183500-191188500 of chromosome 2.
[0071] The third marker comprises more than 80%, preferably more than 90%, and preferably all CG sites in the sequence identical or reverse complementary to the nucleotide sequence located at position 191184000-191189000 of chromosome 2,
[0072] A fourth marker comprising more than 80%, preferably more than 90%, and preferably all of the CG sites in the sequence identical or reverse complementary to the nucleotide sequence located at position 68566500-68571500 of chromosome 4, and
[0073] The fifth marker includes more than 80%, preferably more than 90%, of all CG sites in the sequence that is identical to or reverse complementary to the nucleotide sequence located at positions 30601500-30606500 of chromosome 11, and preferably includes all CG sites.
[0074] 14. The method according to claim 13, wherein
[0075] The first marker includes more than 80%, preferably more than 90%, and preferably all CG sites in the nucleotide sequence shown in SEQ ID NO: 1 or the reverse complementary sequence of the sequence shown in SEQ ID NO: 1.
[0076] The second marker includes more than 80%, preferably more than 90%, and preferably all CG sites in the nucleotide sequence shown in SEQ ID NO: 2 or the reverse complementary sequence of the sequence shown in SEQ ID NO: 2.
[0077] The third marker includes more than 80%, preferably more than 90%, and preferably all CG sites in the nucleotide sequence shown in SEQ ID NO: 3 or the reverse complementary sequence of the sequence shown in SEQ ID NO: 3.
[0078] The fourth marker includes more than 80%, preferably more than 90%, and preferably all CG sites in the nucleotide sequence shown in SEQ ID NO: 4 or the reverse complementary sequence of the sequence shown in SEQ ID NO: 4, and
[0079] The fifth marker includes more than 80%, preferably more than 90%, of all CG sites in the nucleotide sequence shown in SEQ ID NO: 5 or the reverse complementary sequence of the sequence shown in SEQ ID NO: 5, and preferably includes all CG sites.
[0080] 15. The method according to item 13 or 14 is characterized in that, in the analysis step, the comprehensive methylation level of the methylation marker refers to the one calculated based on the methylation level of one or more selected methylation markers.
[0081] 16. The method according to item 15 is characterized in that, in the analysis step, the methylation level of each methylation marker is calculated based on the methylation level of each CG site in the methylation marker, wherein the methylation level of the CG site is the ratio of the cytosine detected to be methylated at the site to the sum of the cytosine detected to be methylated and the cytosine not methylated in all sequence results detected for the site.
[0082] 17. The method according to item 16, characterized in that, in the determination step, the comprehensive methylation level is determined based on the methylation levels of one, two, three, four, or five of the first marker, the second marker, the third marker, the fourth marker, or the fifth marker,
[0083] wherein the threshold is determined based on methylation level detection data of one, two, three, four, or five of the first marker, the second marker, the third marker, the fourth marker, or the fifth marker in a given lung cancer subject or a population of healthy subjects;
[0084] If the comprehensive methylation marker is greater than a specified threshold, the subject is determined to have lung cancer, and if the comprehensive methylation level marker is less than or equal to the specified threshold, the subject is determined to be normal.
[0085] 18. The method according to item 12, characterized in that the subject sample is a blood sample.
[0086] The markers and systems of the present invention can: 1) be used for early screening of asymptomatic people and prognosis detection of cancer patients in a non-invasive manner, reducing the harm caused by invasive detection; 2) have higher sensitivity and accuracy and can achieve real-time monitoring. DETAILED DESCRIPTION
[0087] The present invention is described in detail below. Although specific embodiments of the present invention are shown, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.
[0088] It should be noted that certain words are used in the specification and claims to refer to specific components. Those skilled in the art should understand that technicians may use different nouns to refer to the same component. This specification and claims do not use the difference in nouns as a way to distinguish components, but use the difference in the functions of the components as the criterion for distinction. For example, "including" or "comprising" mentioned throughout the specification and claims are open-ended terms and should be interpreted as "including but not limited to". The subsequent description of the specification is a preferred embodiment of the present invention, but the description is based on the general principles of the specification and is not intended to limit the scope of the invention. The scope of protection of the present invention shall be as defined in the attached claims.
[0089] definition
[0090] Unless specifically defined elsewhere herein, all other technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs.
[0091] Methylation
[0092] Methylation is a key modification of proteins and nucleic acids that regulates gene expression and inactivation. It is closely related to many diseases, including cancer, aging, and Alzheimer's disease, and is a key area of epigenetic research. The most common methylation modifications are DNA methylation and histone methylation.
[0093] DNA methylation refers to the methylation of the carbon atom at the 5th position of cytosine in CpG dinucleotides. As a stable modification, it can be inherited to new daughter DNA during DNA replication under the action of DNA methyltransferases. It is an important epigenetic mechanism. During DNA methylation, methylation of gene promoter regions can lead to transcriptional silencing of tumor suppressor genes, and is therefore closely linked to the development of tumors. Aberrant methylation includes hypermethylation of tumor suppressor genes and DNA repair genes, hypomethylation of repetitive DNA sequences, and loss of imprinting of certain genes, and is associated with the development of various tumors.
[0094] In this article, the ROC curve can, to a certain extent, reflect the classification performance of a classifier. The AUC is actually the area under the ROC curve. The AUC intuitively reflects the classification ability expressed by the ROC curve.
[0095] To calculate PPV and NPV, we first need to introduce the confusion matrix. During the evaluation of a statistical classification model, the number of observations that the model incorrectly and correctly classified is counted, and the results are then presented in a table called the confusion matrix.
[0096] In the confusion matrix:
[0097] TP represents the number of samples whose true value is positive and the model classifies them as positive.
[0098] FP represents the number of samples whose true value is negative but the model classifies them as positive.
[0099] TN represents the number of samples whose true value is negative and the model classifies them as negative.
[0100] FN represents the number of samples whose true value is positive but the model classifies them as negative.
[0101] PPV (Positive Predictive Value), positive predictive value, is the number of samples classified as positive. The calculation formula PPV = TP / (TP + FP) can be used to evaluate the effect of classification as positive.
[0102] NPV (Negative Predictive Value), negative predictive value, is the number of samples classified as negative. The calculation formula NPV = TN / (TN + FN) can be used to evaluate the effect of classification as negative.
[0103] Accuracy, also known as efficiency, is expressed as the percentage of the sum of TP and TN numbers in the total number of subjects.
[0104] Specificity
[0105] Specificity refers to the proportion of negative test results in samples from patients without a specific clinical disease.
[0106] Sensitivity
[0107] Sensitivity refers to the proportion of samples from patients with established clinical disease that test positive.
[0108] The present invention provides a marker for detecting lung cancer, characterized in that the marker comprises one or more of the following one or more methylation markers:
[0109] The first marker comprises more than 80%, preferably more than 90%, and preferably all CG sites in the sequence identical to or reverse complementary to the nucleotide sequence at position 151445000-151450000 of chromosome 1.
[0110] The second marker comprises more than 80%, preferably more than 90%, and preferably all CG sites in the sequence identical to or reverse complementary to the nucleotide sequence at positions 191183500-191188500 of chromosome 2.
[0111] The third marker comprises more than 80%, preferably more than 90%, and preferably all CG sites in the sequence identical or reverse complementary to the nucleotide sequence located at position 191184000-191189000 of chromosome 2,
[0112] A fourth marker comprising more than 80%, preferably more than 90%, and preferably all of the CG sites in the sequence identical or reverse complementary to the nucleotide sequence located at position 68566500-68571500 of chromosome 4, and
[0113] The fifth marker includes more than 80%, preferably more than 90%, of all CG sites in the sequence that is identical to or reverse complementary to the nucleotide sequence located at positions 30601500-30606500 of chromosome 11, and preferably includes all CG sites.
[0114] Further,
[0115] The first marker includes more than 80%, preferably more than 90%, and preferably all CG sites in the nucleotide sequence shown in SEQ ID NO: 1 or the reverse complementary sequence of the sequence shown in SEQ ID NO: 1.
[0116] The second marker includes more than 80%, preferably more than 90%, and preferably all CG sites in the nucleotide sequence shown in SEQ ID NO: 2 or the reverse complementary sequence of the sequence shown in SEQ ID NO: 2.
[0117] The third marker includes more than 80%, preferably more than 90%, and preferably all CG sites in the nucleotide sequence shown in SEQ ID NO: 3 or the reverse complementary sequence of the sequence shown in SEQ ID NO: 3.
[0118] The fourth marker includes more than 80%, preferably more than 90%, and preferably all CG sites in the nucleotide sequence shown in SEQ ID NO: 4 or the reverse complementary sequence of the sequence shown in SEQ ID NO: 4, and
[0119] The fifth marker includes more than 80%, preferably more than 90%, of all CG sites in the nucleotide sequence shown in SEQ ID NO: 5 or the reverse complementary sequence of the sequence shown in SEQ ID NO: 5, and preferably includes all CG sites.
[0120] The methylation levels of the above five markers are significantly different between lung cancer patients and normal people, so they can be used to detect lung cancer.
[0121] The markers for detecting lung cancer may include any one or more of the above five markers, for example, may include one, two, three, four or five.
[0122] In a specific embodiment, the markers include a first marker, a second marker, a third marker, a fourth marker and a fifth marker.
[0123] In a specific embodiment, the marker includes a first marker and a second marker.
[0124] In a specific embodiment, the markers include a first marker and a third marker.
[0125] In a specific embodiment, the markers include a first marker and a fourth marker.
[0126] In a specific embodiment, the markers include a first marker and a fifth marker.
[0127] In a specific embodiment, the marker includes a second marker and a third marker.
[0128] In a specific embodiment, the markers include a second marker and a fourth marker.
[0129] In a specific embodiment, the markers include a second marker and a fifth marker.
[0130] In a specific embodiment, the markers include a third marker and a fourth marker.
[0131] In a specific embodiment, the markers include a third marker and a fifth marker.
[0132] In a specific embodiment, the markers include a fourth marker and a fifth marker.
[0133] In a specific embodiment, the markers include a first marker, a second marker, and a fourth marker.
[0134] In a specific embodiment, the markers include a first marker, a second marker, and a third marker.
[0135] In a specific embodiment, the markers include a first marker, a second marker, and a fifth marker.
[0136] In a specific embodiment, the markers include a first marker, a third marker, and a fourth marker.
[0137] In a specific embodiment, the markers include a first marker, a third marker, and a fifth marker.
[0138] In a specific embodiment, the markers include a first marker, a fourth marker, and a fifth marker.
[0139] In a specific embodiment, the markers include a second marker, a third marker, and a fourth marker.
[0140] In a specific embodiment, the markers include a second marker, a fourth marker, and a fifth marker.
[0141] In a specific embodiment, the markers include a third marker, a fourth marker, and a fifth marker.
[0142] In a specific embodiment, the markers include a first marker, a second marker, a fourth marker and a fifth marker.
[0143] In a specific embodiment, the markers include a first marker, a third marker, a fourth marker and a fifth marker.
[0144] In a specific embodiment, the markers include a second marker, a third marker, a fourth marker, and a fifth marker.
[0145] In a specific embodiment, the marker comprises a first marker.
[0146] In a specific embodiment, said marker comprises a second marker.
[0147] In a specific embodiment, the marker comprises a third marker.
[0148] In a specific embodiment, said markers include a fourth marker.
[0149] In a specific embodiment, said markers include a fifth marker.
[0150] The present invention also provides a first marker comprising more than 80%, preferably more than 90%, and preferably all CG sites in the sequence identical to or reverse complementary to the nucleotide sequence at positions 151445000-151450000 of chromosome 1, and the use of the first marker for detecting lung cancer or in preparing a kit for detecting lung cancer.
[0151] Use of a second marker comprising more than 80%, preferably more than 90%, and preferably all CG sites in the sequence identical to or reverse complementary to the nucleotide sequence at positions 191183500-191188500 of chromosome 2, for detecting lung cancer or in preparing a kit for detecting lung cancer.
[0152] Use of a third marker comprising more than 80%, preferably more than 90%, and preferably all CG sites in the sequence identical to or reverse complementary to the nucleotide sequence located at positions 191184000-191189000 of chromosome 2, for detecting lung cancer or in preparing a kit for detecting lung cancer.
[0153] Use of a fourth marker comprising more than 80%, preferably more than 90%, and preferably all CG sites in the sequence identical to or reverse complementary to the nucleotide sequence located at position 68566500-68571500 of chromosome 4, for detecting lung cancer or in the preparation of a kit for detecting lung cancer.
[0154] Use of a fifth marker comprising more than 80%, preferably more than 90%, and preferably all CG sites in the sequence identical or reverse complementary to the nucleotide sequence located at positions 30601500-30606500 of chromosome 11, for detecting lung cancer or in the preparation of a kit for detecting lung cancer.
[0155] The kit of the present invention can be any kit for lung cancer diagnosis and early screening based on marker methylation data, as long as it includes reagents for detecting the methylation level of the marker of the present invention.
[0156] The present invention also provides a system for detecting lung cancer, wherein the system comprises:
[0157] A sample collection module, which is used to collect samples from subjects;
[0158] A sample processing module is used to extract cfDNA from the sample and perform methylation treatment for sequencing; preferably, the methylation treatment is performed using bisulfite; a sequencing module is used to sequence the sample after being processed by the sample processing module;
[0159] an analysis module for analyzing the sequencing results to determine the comprehensive methylation levels of the methylation markers in the sample;
[0160] The determination module is used to determine whether the subject has lung cancer based on the comprehensive methylation level obtained by analysis.
[0161] In the sample collection module, it is used to collect samples from the subject, for example, a blood sample.
[0162] In the sample processing module, the cfDNA extraction step can use methods known in the art. Bisulfite treatment can convert C bases to U bases, but methylated C bases do not react with bisulfite and remain C bases during bisulfite treatment. Therefore, in sequencing data comparisons, any detected C base indicates that the C at that position is methylated.
[0163] The sequencing in the sequencing module is preferably performed using the second generation sequencing technology.
[0164] In the analysis module, the comprehensive methylation level of the methylation marker refers to the one calculated based on the methylation level of one or more selected methylation markers. For example, if five methylation markers are selected for a system for detecting lung cancer, the methylation level of each of the five markers can be obtained, and then the corresponding formula is obtained by the algorithm determined after modeling these five methylation markers. The methylation levels of the five markers are input into the corresponding formula to calculate the comprehensive methylation level for these five methylation markers. Similarly, if two methylation markers are selected for a system for detecting lung cancer, the methylation level of each of the two markers can be obtained, and then the comprehensive methylation level for these two methylation markers is calculated by the algorithm determined after modeling these two methylation markers.
[0165] Furthermore, the comprehensive methylation level refers to the methylation level determined based on one, two, three, four or five of the first marker, the second marker, the third marker, the fourth marker or the fifth marker.
[0166] In a specific embodiment, the comprehensive methylation level is determined based on the methylation levels of the first marker, the second marker, the third marker, the fourth marker, and the fifth marker.
[0167] In a specific embodiment, the comprehensive methylation level is determined based on the methylation levels of the first marker and the second marker.
[0168] In a specific embodiment, the comprehensive methylation level is determined based on the methylation levels of the first marker, the second marker, and the fourth marker.
[0169] In a specific embodiment, the comprehensive methylation level is determined based on the methylation levels of the first marker, the second marker, the fourth marker, and the fifth marker.
[0170] In a specific embodiment, the comprehensive methylation level is determined based on the methylation level of a fifth marker.
[0171] The methylation level of each methylation marker is calculated based on the methylation level of each CG site in the methylation marker, wherein the methylation level of the CG site is the ratio of the cytosine detected to be methylated at the site to the sum of the cytosine detected to be methylated and the cytosine not methylated in all sequence results detected for the site.
[0172] For each window, the number of CG sites in each window is counted. Because the depth of the methylated cytosine of each CG site and the total depth of the site are known, the methylation level of the entire window can be calculated, which is the ratio of the sum of the depths of the methylated cytosine of all CG sites divided by the sum of the total depths of all CG sites. Each window will obtain a corresponding methylation level through the above calculation method. Among them, the depth of the methylated cytosine of each CG site is the number of reads whose sequencing test results show that the site is methylated cytosine, that is, the sequencing result shows that the result of the site measured is C (cytosine) reads, and the total depth of the site is the total number of all sequencing reads covering the site, that is, the measured results show that the site is C or T (thymine) reads total. The depth of methylated cytosine and the total depth of the site can be directly provided after analysis by sequencing software.
[0173] Specifically, the sequencing results can be analyzed and the methylation levels calculated by the following methods:
[0174] Use fastp and other quality control software to check the quality of the original sequencing data, and filter, intercept or remove low-quality reads to obtain the corresponding clean data; use Bismark bowtie2 alignment software to align the clean data after quality control to the reference genome (hg19); use deduplicate_bismark to deduplicate the bam file of the initial alignment; use Bismark_methylation_extractor to extract the corresponding methylation site information to obtain the final methylation CG file (including all CG site information files); finally, use the sliding window method to slide the reference genome and calculate the overall methylation level of the CG site in each window interval.
[0175] In the determination module, if the comprehensive methylation level of the marker is greater than a specified threshold, the subject is determined to have lung cancer, and if the comprehensive methylation level of the marker is less than or equal to the specified threshold, the subject is determined to be normal. The specified threshold is determined based on the selected methylation markers, after constructing a model based on a certain amount of sample training sets (including the type and methylation level of each sample), and then based on the cutoff value on the obtained ROC curve. In one specific embodiment, the constructed model is a random forest model constructed using the randomForest package in R.
[0176] In a specific embodiment, the designated threshold is 0.4-0.6. In a specific embodiment, the designated threshold is 0.442-0.501.
[0177] The present invention also provides a method for detecting lung cancer, comprising:
[0178] a sample collection step, which collects a sample from a subject;
[0179] a sample processing step, which extracts cfDNA from the sample and performs a methylation treatment for sequencing; preferably, the methylation treatment is performed using bisulfite;
[0180] a sequencing step, which sequences the sample processed in the sample processing step;
[0181] an analysis step of analyzing the sequencing results to determine the comprehensive methylation levels of the methylation markers in the sample;
[0182] The determination step is to determine whether the subject has lung cancer based on the comprehensive methylation level obtained by analysis.
[0183] The descriptions of the markers, the comprehensive methylation level of the methylation markers, the methylation level of each methylation marker, and the subject having lung cancer are as described above for the system for detecting lung cancer.
[0184] Example
[0185] The present invention provides general and / or specific descriptions of the materials and experimental methods used in the experiments. In the following examples, unless otherwise specified, % represents wt%, i.e., percentage by weight. All reagents or instruments used, without manufacturer's indication, are commercially available conventional reagents.
[0186] Example 1
[0187] 1.1cfDNA extraction and purification
[0188] 1.1.1 Preparation of plasma samples:
[0189] Centrifuge the blood sample at 2000 g for 10 min at 4°C and transfer the plasma to a fresh centrifuge tube. Centrifuge the plasma sample at 16000 g for 10 min at 4°C and proceed to the next step based on the type of collection tube used, as shown in Table 1. The collection tube type used in this experiment is Other.
[0190] Table 1
[0191]
[0192]
[0193] 1.1.2 Cleavage and Binding
[0194] Prepare Binding Solution / Beads Mix according to Table 2 and mix thoroughly.
[0195] Table 2
[0196]
[0197] Add an appropriate volume of plasma sample.
[0198] 1.1.2.2. Thoroughly mix the plasma sample and binding solution / magnetic bead mixture.
[0199] 1.1.2.3. Incubate on a rotary mixer for 10 minutes to allow the cfDNA to bind to the magnetic beads.
[0200] 1.1.2.4. Place the binding tube on the magnetic rack for 5 minutes until the solution becomes clear and the magnetic beads are completely adsorbed on the magnetic rack.
[0201] 1.1.2.5. Carefully discard the supernatant with a pipette. Keep the tube on the magnetic rack for a few minutes and remove the remaining supernatant with a pipette.
[0202] 1.1.3 Washing
[0203] 1.1.3.1. Resuspend the beads in 1 ml of wash buffer.
[0204] 1.1.3.2. Transfer the resuspension to a new 1.5ml centrifuge tube without adsorption. Keep the combined tube.
[0205] 1.1.3.3. Place the centrifuge tube containing the bead resuspension on a magnetic rack for 20 seconds.
[0206] 1.1.3.4. Aspirate the supernatant obtained from the separation into the washing and binding tube, collect the remaining beads after washing into the resuspension solution again, and discard the lysis / binding tube.
[0207] 1.1.3.5. Place the tube on a magnetic rack for 2 min until the solution becomes clear and the beads aggregate on the magnetic rack. Remove the supernatant using a 1 ml pipette.
[0208] 1.1.3.6. While the tube remains on the magnetic stand, use a 200 μL pipette to remove as much residual liquid as possible.
[0209] 1.1.3.7. Remove the tube from the magnetic rack, add 1 ml of wash buffer, and vortex for 30 seconds.
[0210] 1.1.3.8. Place on a magnetic rack for 2 min until the solution becomes clear and the beads aggregate on the magnetic rack. Remove the supernatant using a 1 ml pipette.
[0211] 1.1.3.9. While the tube remains on the magnetic stand, use a 200 μL pipette to completely remove any remaining liquid.
[0212] 1.1.3.10. Remove the tube from the magnetic rack, add 1 ml of 80% ethanol, and vortex for 30 seconds.
[0213] 1.1.3.11. Place on the magnetic stand for 2 minutes. When the solution becomes clear, remove the supernatant with a 1 ml pipette.
[0214] 1.1.3.12. Leave the tube on the magnetic stand and remove any remaining liquid using a 200μL pipette.
[0215] 1.1.3.13. Repeat steps 10-12 above once with 80% ethanol and remove as much supernatant as possible.
[0216] 1.1.3.14. Leave the tube on the magnetic rack and air dry the beads for 3-5 minutes.
[0217] 1.1.4 Elution of cfDNA
[0218] 1.1.4.1. Add dilution solution according to Table 3.
[0219] Table 3
[0220]
[0221] 1.1.4.2. Vortex for 5 minutes and place on a magnetic stand for 2 minutes. The solution becomes clear and the cfDNA in the supernatant is aspirated.
[0222] 1.1.4.3. Use the purified cfDNA immediately, or transfer the supernatant to a new centrifuge tube and store at -20°C.
[0223] 1.2DNA shearing and purification:
[0224] 1.2.1. Based on the Qubit concentration, take 2 μg of DNA, add water to 125 μl, and add it to a Covaris 130 μl interruption tube. Set the program: 50W, 20%, 200 cycles, 250 s.
[0225] 1.2.2 After fragmentation, take 1 μl of sample and use Agilent 2100 for fragment detection. After normal fragmentation, the main peak of the sample detection is about 150bp-200bp.
[0226] For cfDNA samples, Agilent 2100 performed fragment detection and Qubit was used directly for subsequent experiments. 1.3 End repair and 3' end addition of "A":
[0227] 1.3.1. Transfer 50 ng of sheared gDNA or cfDNA to a PCR tube, make up to 50 μl with nuclease-free water, add the reagents in Table 4, and vortex to mix thoroughly:
[0228] Table 4
[0229]
[0230]
[0231] 1.3.2. Set up the following program to perform the reaction on a PCR instrument:
[0232] The specific procedure is shown in Table 5, and the heating cover temperature is 85°C.
[0233] Table 5
[0234] temperature time 20℃ 30min 65℃ 30min 4℃ ∞
[0235] 1.4 Adapter ligation and purification:
[0236] 1.4.1. Refer to Table 6 to dilute the adapter to the appropriate concentration in advance:
[0237] Table 6
[0238] Fragmented DNA per 50 μl ER&AT reaction Connector concentration 1 μg 10 μM 500ng 10 μM 250ng 10 μM 100ng 10 μM 50ng 10 μM 25ng 10 μM 10ng 3μM 5ng 5μM 2.5ng 2.5 μM 1ng 625nM
[0239] Prepare the following reagents according to Table 7, pipette gently to mix, and centrifuge briefly:
[0240] Table 7
[0241]
[0242]
[0243] 1.4.3. Set up the following program shown in Table 8 and run the reaction on a PCR instrument:
[0244] No hot cover.
[0245] Table 8
[0246] temperature time 20℃ 30min 4℃ ∞
[0247] 1.4.4. Add purification magnetic beads according to the system shown in Table 9 and conduct the experiment (Agencourt AMPure XP magnetic beads should be brought to room temperature and shaken to mix evenly before use):
[0248] Table 9
[0249] Components volume Connector connection products 110 μl Agencourt AMPure XP beads 110 μl Total volume 220 μl
[0250] 1.4.4.1. Mix gently by pipetting 6 times.
[0251] 1.4.4.2. Incubate at room temperature for 5-15 minutes. Place the PCR tube on a magnetic rack for 3 minutes to clarify the solution. 1.4.4.3. Remove the supernatant and place the PCR tube on the magnetic rack again. Add 200 μl of 80% ethanol solution to the PCR tube and let it sit for 30 seconds.
[0252] 1.4.4.4. Remove the supernatant and add 200 μl of 80% ethanol solution to the PCR tube. Let it stand for 30 seconds and then completely remove the supernatant (it is recommended to use a 10 μl pipette to remove the residual ethanol solution at the bottom).
[0253] 1.4.4.5. Let stand at room temperature for 3-5 minutes to allow the residual ethanol to evaporate completely.
[0254] 1.4.4.6. Add 22 μl of Nuclease-free water, remove the PCR tube from the magnetic stand, gently pipette to resuspend the magnetic beads to avoid creating bubbles, and let it stand at room temperature for 2 minutes.
[0255] 1.4.4.7. Place the PCR tube on a magnetic rack for 2 minutes to clarify the solution.
[0256] 1.4.4.8. Use a pipette to aspirate 20 μl of supernatant and transfer it to a new PCR tube.
[0257] 1.5 Bisulfite treatment and purification:
[0258] 1.5.1. Prepare the necessary reagents in advance and dissolve them. Add the following reagents according to Table 10:
[0259] Table 10
[0260] Components Volume of high concentration samples (1ng-2μg) Low concentration samples (1-500ng) Adapter ligation purified product 20 μl 40 μl Bisulfite solution 85 μl 85 μl DNA protection buffer 35 μl 15 μl Total volume 140 μl 140 μl
[0261] 1.5.2. When DNA protection buffer is added, the liquid will turn blue. Gently pipette to mix thoroughly, then divide the tubes into two and place them in the PCR instrument.
[0262] 1.5.3. Set up the program shown in Table 11 below and run it:
[0263] Heat lid 105℃.
[0264] Table 11
[0265] temperature time 95℃ 5min 60℃ 10min 95℃ 5min 60℃ 10min 4℃ ∞
[0266] 1.5.4. Briefly centrifuge and combine the two tubes of the same sample into a clean 1.5ml centrifuge tube.
[0267] 1.5.5. Add 310 μl of Buffer BL (1 μl of Carrier RNA (1 μg / μl) if the sample volume is less than 100 ng) to each sample, vortex to mix, and centrifuge briefly.
[0268] 1.5.6. Add 250 μl of absolute ethanol to each sample, vortex mix for 15 seconds, centrifuge briefly, and add the mixture to the corresponding prepared spin column.
[0269] 1.5.7. Let stand for 1 minute, centrifuge for 1 minute, transfer the liquid in the collection tube back to the centrifuge column, centrifuge for 1 minute, and discard the liquid in the centrifuge tube.
[0270] 1.5.8. Add 500 μl of buffer BW (note whether anhydrous ethanol has been added), centrifuge for 1 minute, and discard the waste liquid. 1.5.9. Add 500 μl of buffer BD (note whether anhydrous ethanol has been added), cap the tube, and let it stand at room temperature for 15 minutes. Centrifuge for 1 minute and discard the liquid.
[0271] 1.5.10. Add 500 μl of buffer BW (note whether anhydrous ethanol is added), centrifuge for 1 min, discard the remaining liquid, and repeat once, for a total of 2 times.
[0272] 1.5.11. Add 250 μl of anhydrous ethanol and centrifuge for 1 min. Place the spin column into a new 2 ml collection tube and discard any remaining liquid.
[0273] Place the spin column in a clean 1.5 ml centrifuge tube, add 20 μl of nuclease-free water to the center of the spin column membrane, gently close the tube cap, incubate at room temperature for 1 min, and centrifuge for 1 min.
[0274] 1.5.13. Transfer the liquid in the collection tube back to the centrifuge column, let it stand at room temperature for 1 minute, and centrifuge for 1 minute. 1.6 Amplification and purification:
[0275] Prepare the reaction system as shown in Table 12, mix thoroughly by pipetting, and centrifuge briefly:
[0276] Table 12
[0277] Components volume Purification of the product after bisulfite treatment 20 μl Amplification enzyme 25 μl Upstream primer (10 μM) 2.5 μl Downstream primer (10 μM) 2.5 μl Total volume 50 μl
[0278] Set up the program shown in Table 13 below and start the PCR program:
[0279] Heated cover 105℃
[0280] Table 13
[0281]
[0282]
[0283] 1.6.3. The number of PCR cycles is adjusted according to the amount of DNA input. Reference data are shown in Table 14:
[0284] Table 14
[0285] Input DNA Number of cycles required to generate 100 ng 1 μg 7-8 500ng 8-9 250ng 9-10 100ng 10-11 50ng 11-12 25ng 12-13 10ng 13-14 5ng 14-15 2.5ng 16-17 1ng 18-19
[0286] 1.6.4. Add 50 μl of Agencourt AMPure XP magnetic beads to the PCR tube after the reaction is completed. Mix thoroughly with a pipette to avoid creating bubbles (Agencourt AMPure XP should be mixed and equilibrated at room temperature beforehand). 1.6.5. Incubate at room temperature for 5-15 minutes. Place the PCR tube on a magnetic rack for 3 minutes to allow the solution to clarify.
[0287] 1.6.6. Remove the supernatant and place the PCR tube on the magnetic rack. Add 200 μl of 80% ethanol solution to the PCR tube and let it stand for 30 seconds.
[0288] 1.6.7. Remove the supernatant and add 200 μl of 80% ethanol solution to the PCR tube. Let it stand for 30 seconds and then completely remove the supernatant (it is recommended to use a 10 μl pipette to remove the residual ethanol solution at the bottom).
[0289] 1.6.8. Let stand at room temperature for 5 minutes to allow the residual ethanol to evaporate completely.
[0290] 1.6.9. Add 30 μl of nuclease-free water, remove the centrifuge tube from the magnetic stand, and use a pipette to gently pipette to resuspend the magnetic beads.
[0291] 1.6.10. Let stand at room temperature for 2 min. Place the 200 μl PCR tube on a magnetic rack for 2 min to clarify the solution.
[0292] 1.6.11. Use a pipette to transfer the supernatant to a new 200 μl PCR tube (placed on ice), label the tube with the sample number, and prepare for the next reaction.
[0293] 1.6.12. Take 1 μl of sample and use Qubit to measure the library concentration. Record the library concentration.
[0294] 1.6.13. Take 1 μl of sample and use Agilent 2100 to measure the library fragment length. The library length is approximately between 270 bp and 320 bp.
[0295] 1.6.14. Sequencing was performed using the Illumina high-throughput sequencing platform (Illumina novaseq 6000).
[0296] 1.6.15. Methylation bioinformatics analysis process. The process is as follows: Use quality control software such as fastp to check the quality of the raw sequencing data and filter, truncate, or remove low-quality reads to obtain the corresponding clean data; use Bismarkbowtie2 alignment software to align the clean data after quality control to the reference genome (hg19); use deduplication_bismark to deduplicate the initial aligned bam files; use Bismark_methylation_extractor to extract the corresponding methylation site information to obtain the final methylation CG file (including all individual CG site information files); finally, use the sliding window method to draw a window on the reference genome and calculate the overall methylation level of the CG site in each window interval; for each sample, calculate the methylation level of the corresponding window, and identify differentially methylated windows based on different sample groups.
[0297] Example 2
[0298] Based on the cfDNA training set of 14 lung cancer patients and 22 healthy subjects (see Figure 1 ), the methylation levels of 1583 initial markers in 14 lung cancer patients and 22 healthy subjects were detected using the method described in Example 1, and the five most significant methylation regions that distinguished cfDNA of lung cancer and healthy subjects were screened out as candidate biomarkers for lung cancer detection. As shown in Table 15, the corresponding marker information is as follows: the first marker, which includes all CG sites in the sequence that is identical to or reverse complementary to the nucleotide sequence at positions 151445000-151450000 of chromosome 1; the second marker, which includes all CG sites in the sequence that is identical to or reverse complementary to the nucleotide sequence at positions 191183500-191188500 of chromosome 2; the third marker, which includes all CG sites in the sequence that is identical to or reverse complementary to the nucleotide sequence located at positions 191184000-191189000 of chromosome 2; the fourth marker, which includes all CG sites in the sequence that is identical to or reverse complementary to the nucleotide sequence located at positions 68566500-68571500 of chromosome 4; the fifth marker, which includes all CG sites in the sequence that is identical to or reverse complementary to the nucleotide sequence located at positions 30601500-30606500 of chromosome 11.
[0299] The methylation level data of each candidate marker detected by the method of Example 1 for 14 lung cancer patients and 22 healthy subjects were input into R software, and a random forest model was constructed using the randomForest package of R software for model regression. The regression results showed that in the training set, the cutoff value of the comprehensive methylation level based on the five markers that can be used to predict lung cancer results was 0.442, that is, the specified threshold was 0.442 (greater than 0.442 was interpreted as a lung cancer patient), and the AUC obtained by the final model reached 1, with an accuracy of 100%, a sensitivity of 100%, a specificity of 100%, a PPV of 100%, and an NPV of 100%. Specific information is shown in Tables 15 and Figure 2 .
[0300] Table 15
[0301]
[0302] Example 3
[0303] Based on the five methylation markers in Example 2, using pROC in R software, according to the methylation level of each methylation marker, the cutoff and AUC values of the comprehensive methylation levels of the five methylation markers in the test set (cfDNA of 10 lung cancer patients and cfDNA of 16 healthy people) that can be used to predict lung cancer outcomes were calculated, as shown in Table 16.
[0304] Table 16
[0305]
[0306] Example 4
[0307] Based on the model constructed in Example 2, in the test set of 10 lung cancer patient cfDNA and 16 healthy subjects cfDNA, the cutoff of the comprehensive methylation level based on the five markers that can be used to predict lung cancer outcomes is 0.442, that is, the designated threshold is 0.442 (greater than 0.442 is interpreted as lung cancer patients), the AUC reaches 0.919, the accuracy is 84.62%, the sensitivity is 90%, the specificity is 81.25%, the PPV is 75%, and the NPV is 92.86%. Detailed information is shown in Tables 17 and Figure 3 .
[0308] Table 17
[0309]
[0310] Example 5
[0311] A random forest model was constructed in the training set using two of the five methylation markers (the first marker and the second marker). The cutoff for the combined methylation level of the two markers that could be used to predict lung cancer outcomes was 0.465, with a designated threshold of 0.465 (a value greater than 0.465 was considered a lung cancer patient). In the test set (cfDNA from 10 lung cancer patients and 16 healthy individuals), the AUC reached 0.869, with an accuracy of 84.62%, a sensitivity of 100%, a specificity of 75%, a PPV of 71.43%, and an NPV of 100%. Details are shown in Tables 15 and Figure 4 .
[0312] Example 6
[0313] As in Example 5, a random forest model was constructed in the training set using three of the five methylation markers (the first marker, the second marker, and the fourth marker). The cutoff for the comprehensive methylation level of the three markers that can be used to predict lung cancer outcomes was 0.501, that is, the designated threshold was 0.501 (a value greater than 0.501 indicates a lung cancer patient). In the test set (cfDNA from 10 lung cancer patients and cfDNA from 16 healthy individuals), the AUC reached 0.894, with an accuracy of 80.77%, a sensitivity of 70%, a specificity of 87.5%, a PPV of 77.78%, and an NPV of 82.35%. For details, see Tables 15 and Figure 5 .
[0314] Example 7
[0315] As in Example 5, a random forest model was constructed in the training set using four of the five methylation markers (the first marker, the second marker, the fourth marker, and the fifth marker). The cutoff for the comprehensive methylation level of the four markers that can be used to predict lung cancer outcomes was 0.461, i.e., the designated threshold was 0.461 (a patient with lung cancer was identified as one with a score greater than 0.461). In the test set, the AUC reached 0.906, with an accuracy of 80.77%, a sensitivity of 90%, a specificity of 75%, a PPV of 69.23%, and an NPV of 92.31%. Details are shown in Tables 15 and Figure 6 .
[0316] In summary, it can be seen that the areas screened in cfDNA using this method have a very high correlation with lung cancer screening.
[0317] The technical features of the above-mentioned embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features of the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0318] The above-described embodiments merely illustrate several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the patent for this invention shall be determined by the appended claims.
[0319] Sequence Listing
[0320] SEQ ID NO: 1:
[0321] GCTCCGTTGCCCAAGCTGGAGTGCAGTGGCTTATAAAACATGGCTCACTGCAGCCTTGAACTCCTGGGCTCATGC
[0322] AATCCTCCTGCCTCAGCCTCTGGAGTAGCTGGGACCACAGAAGTGCACCACCACACCCGGCTAGTTTTATTTTTAT
[0323] TTGTAGAGATGGGATCTCCCTATGTTGCCCAGTCTGGTCTCCAACCTGGGGCCTCAAGTGTTCCCCCTGCCTCGGC
[0324] CTCCCAAAGTGCAGGGATTACAGCTTGGGCCACTGTGCCCAGCCCTTTCTTCCATCATTCCGTCCTCTCCTTCCTTC
[0325] TTAATACTGCTTGTCAAACTGATTTTGCCACCCAGTTTGAAAAAGAAGGTCTAGAGGATTATCCCAATCATTAACTA
[0326] TTTCTGCAGGTTAGCTTACCTTTCCAGAATACAGAATAGAATTCAGCAAACTGCCTCAAAGAGCAGTTTCTAGATT
[0327] TGACTTCTTCCTCTAGGGTTTATGAAGTCATCATACTGAAGTGACCTGGATTTCCCTATACGTGAACCTCATGACGT
[0328] GGGAGCCCGGGGTAGGGAGGGACAGATAACATGAGAATCAATAGAAAGAACTGCAAGGTCTGAAGACAAGAAC
[0329] CATTTGTCTAGTGGAGTCTATTTTATATCCATTTTTTAAAAATCTCCATTTAAGATATTTAGCTTCAACGTGCACCTTG
[0330] ACCAATTAATCCAACAAACAAAACTCTGGCTTGGGGGCCCAGTTACTCACTTTCATGGAAAGAGACCTCTGTTTTT
[0331] AGAACCAGGGATCTTGTGAGTCCCTCCCACGTCCCCAGCCCTTCTCTGGCCTCGGGTAACTTTTCCCTCCAGTTTG
[0332] GTGGAAGGTTACATTAGAAGAAAATAGTCCCTACCGTCTAGAAAGTTTTCGGGGGGGGAAAAAAGGGCGTTGGA
[0333] CGTGCGGGTGGGTGGGTATACCGGGAGGGGACAGCTGGAGCCGCGAAGGCGCGGGGGCCCGCGGCGCGTGGGT
[0334] GTGGAGGCCGCGGGGCGGCAGGAGGCACTGGGAAGCGGGGTGAAGAGCACCGAAGTCCAGAAAAGGGCAGTG
[0335] TCCGCCCTGCGCTCCTCTGCTGGCGCCCGCAACCCGGGAGGGAGCGGCGAGGCTCCGAAGGGTGGGATTCGAGG
[0336] TCTGAACCCTCGAGTGGGCTTGGGTTTTCAGGAGCCGCCTTTTCTTATTGATCATTTCTCTTTTGTCATCATAATCTC
[0337] TTGGCCTCAGACTGAGTTAAGGGAGCCATAGTAGGAGTGCAGAAAATGAAGTGGAGGAAAAAAAAATGTGATAT
[0338] TTAAAATATAGTTTTTCCTACTTTCTTTTTCTCTCTCAAATTAATTTAGATGTAAATAACATTTTAACAGGGCCCAACA
[0339] TCAGGCTCCTAGGCTCTGGCTGCTGGTCTCACTTGACACCAGAATAAGATTTTGGTGTGCTAGTAATTTTGCCAAA
[0340] TGGAGGCCTACTCCCCCAAAACATCACCGTGGTGCATGATGAACATTTCTGTAGGTGATGACGTGGAGTTTTTGTTT
[0341] TACTTCATTCTAATACATGTCCCTGCTTCAAGTGGGTAAAATTATACCTAGCTCAGCCCAGGTTTAATTCAGATTATT
[0342] TAAAAAATTTCAGTGGCAAGTAATTTACGAGCTTCTCTTTATTAAGATTTAGGAAATTTTGCTTTATTGCGCATGCAT
[0343] TTTAAGTATTTTTTTGCTGTGTTCCTAGTGTAGATAGAAAACAACCTGTCACTGTTCCTAGTTGGCCCTGTCAACA
[0344] AACATTTGAGAAGTGCTTAGATCAGAATGCTGTTGAGTACTACAAATTAAACAATCTGATGTACTTGTTAATTTCAT
[0345] TACAAACAGCAGTTAGGATTTCATATACATATAGTCTGTACCATGCTATATATCAGAAATCTTTGTAAATATAAACTCC
[0346] TTAAATATAGGGACCCAGTTTTATTCTTCAGCGAAAGGGAGTTAATCTTTTGGTGTTGTGCATTTAGAAGGATCTCG
[0347] GATGTAAGAGGCTCATTAAATGATAGCTGTTAGTATTATTTATCTTGGTATCTTCCAAAGGGCTTTGCACATGGCTAA
[0348] CACTAAGAAATATTGGTTGAATGAATGAAGGTGCCATGCCCATACTGTGATATACACAGAGAAAAGACCAGGATG
[0349] GGGACACTAGCACAGTGATTTCATTTCATCAGTGAAGCTTTGTGTGGAAGTGTCCCTCCTGCAGCCCTTTTACCAG
[0350] ATGCAGGCTCCCTTGACCTTGGATTGTGGACCACTTAAAATTTTCAGGGGAAAAAGTGTGCCTAATTAGCTAGTAG
[0351] GGCCTAACCTCTTGTCAAGATGGAACTTGAGAAGTTTGTCTAAGTATGGAAAAGCTAAGCTCACTTGACACAGAT
[0352] CACGGGAAGAAAGATTTCTGCATGTTATTCAAGTTATTCACGTCAAAAGCGATTAGGGGGCCCCAGCTTCCATTAC
[0353] ATCATTCATTAGGAAAATGTGTGCTGGGAGCTCTGCAAATCACATGATTTGATGGAGGCATTGAGACTGGGTTTCT
[0354] TGTGTGTCTGCCCTTGGCCATATTCAGTTTTTAGTATGTTTGGAAAAGGCTGAAGCAGAACATGAAGATCTAGTTTA
[0355] TTTTAATGCTGTCAGACAGCTAAATGGTGGAAACTTGTTGAAGAGATTTACTGAGCTGTTTCCTTAGATTGAAGAC
[0356] TTTAGAGAATTAATTAACAGGGTTCTGAGAGATACAACATGCCAGACCCCTACTGACTCTAAGAGTGTACTTTCTT
[0357] ACTGATATGATGTGATAGCTGAGCACAGTATATGTGAAACTGCTGGGGAGTAATACAAGAGCAGTGCTGATGTGTT
[0358] CTAAGACATGCAGGTGTTCAGATGAGAACTGGCCATGACAACTGAGTCAGCCAGGCACTGGACTCATTTCCTTTT
[0359] GCTTAAGCACCATGAGCTCAGTAGAAGCACAGACTTTCACAAGGTGCTGAACTACTTGGTGAAAATGAAGTCAC
[0360] AATTTTGGAAGAGCTTCACTGATTTTGATGTGTAGGTATTGTTTCACGTTTTAGCTGTTGCTTTTAGAATGCATGCCT
[0361] TAAGATGGTTCTCAGTTATCATGCTATATCTATGCCACGCCATTTCTAGAGCAAGTAACTTAGCTGAAAGTGACTAT
[0362] ACCTATAAGAAAACAAATATCCCTGTAGAGACATTTGTAGATGAATTGTGGAAGTACTCTAAAATTTTTGGCCAGAT
[0363] TTTGGGGAATGTTTGAGTCTAGCTGGTACAGAAAGTACATTCTGTTTTATGAAAATAGCAACATTGAAGTGAAAAG
[0364] ATATCATTAATGTATACATGTAGTCAGAAATTGAATGTGCCCTGGCTGATTGCAATTTATTTATTCACTTAATGCGTAT
[0365] TTATTGAGTACAAGGCCTTAGAGTTTCCTCAGTTCTTATAGTCTGTAACAGTGCCATTCCATAGGAATGTAATGAGA
[0366] GCCACATATATAATTTACAATTTTCTAGTAGCCACATTAAAAAATTAAAAAGAAACTGATAAAATTAATGTTAATATT
[0367] AATATATTTTATTTCACCAAATATATCTAAAAATATATCTAAAATATTATCCTAATATGTAATTAATATACTTTTATTGGT
[0368] TTTTTTTTTTTGTTTTTTTTTTTTTTTTGGGGACAGGGTCTCACTCTGTCACCCAGGCTGGAGTGCAGTGACACAAT
[0369] CATGGTTCATTGCAGCCTCAACCTTCTGGGCTCAGATGCTCCTCCCACCTCAGCCTCGTGAGTAGCTGAGACCACA
[0370] GGCATTTGCCACCAGCTAATTTTTTGTATTTTTTTGTAGAGACAGGATTTCCCATGTTGCCCAGGCTGGTCTCCAAC
[0371] TCCTGGGCTCACGCGATCCACCCGCCTTGGTCTCCTGAAGTGCAGGGATTACAGGCATGAGCCACTAGAAATTATT
[0372] AACAAATTATTTTACATTCTTCGTTTTGAATTAAGTTTCTCGAATTTACTTACAGCAAATCTCACATTGCACTAGCCA
[0373] CATTTCAGGTACTAAAAAACTACATGTGATTAATGGTTACCATACTGGGCAGTCCATGTCTAGAGAGAAATACAGAT
[0374] GTGTAATTTGTAAACAGGTAAACAGATAATTTCTTTTCTTTTTCTATTTTTCTTTTTTTTTTGAGACAAATTCTCACT
[0375] CTCCCAGGCCGGAGTGCAGAGGCTCAGTGCAGCCTCCACCCCCTGGGCTCAAGCGATCCTCCCAGCTAACCTCCC
[0376] CAACTCCCCACCCCAGCAGCTGGGACTACAGGTGGGTACCACCATGTCCAACTAATTTTTATAGAGATAGGTTTTC
[0377] ATCATGTTACCCAGACTAGATAATTTTAATTAAATGTAATAAAGGCTATGATGAGAAAAGTACACTCATCATACAAT
[0378] GAAAGCATAGAAAAGGGGCTTAGTTCCTCCAGTTTTTCCTCCTAAAATCTCATAATTCTCTCTGCTTCTCTGCATCC
[0379] TGCCTGCCACGACTCTAATCTAAGCCAGATTATCTCTCTCCTAGACTTGCACAACAGCTTCCGAAATTGTCTACCTG
[0380] CATCTATCCCTGCCCCTGTTTGCAGATAAAGTGAGCTTTTAAAAACGCCCATCAGATCATGTCATTCCTCTGTTTAA
[0381] AACCTTAAATATATGATGAACACATTGGACAAGTGTGAGATGATGTACATAAAGCGCCCAGCCCAAAGGTTCTGCT
[0382] CACTGAGTTTAGTTCTCCTCGTTTCTTCCTTTCACTCTCTTTTTTCTTCCTCTTACTTTCCTCCCCAGCTCCTTCCTCT
[0383] TATCACCTATCAGTCTCTTATCATCTGAATTTCTTTTTTCTTTCTTCCTTCTCACTTGCTTTCCATTTTTTTCTTCCATC
[0384] TTTACTTTCTCTCCTTTCCTGCTTTAGTTTATTTTAATATAGAATTTAAATAGATTTAGAAATTCATTCTTTAATTATTTA
[0385] CTTCTGAGTTCCAAGCTCTTCATTTAAAAATCAAAAGAGTACTGGGCACTGTGGCTCACGCCTGTAATCCCAGCAC
[0386] TTTGGGAGGCCGAGGTGGGCGGA
[0387] SEQ ID NO:2:
[0388] TCTCCATGAGCTGGTGCAAGGACAACTAATTTTAAGGTATTAAAAAAACAAACAGAAAAG
[0389] CAACTCTTAAATGTATAACAAAAATTTACTGAGAATTAGGTCTGTGGATACAGTGTTAGA
[0390] CATCTAGACACTTTCCCCTGAAACTATTCTTCTTAAGTCAAGGAAACACCTCCATCCCTT
[0391] TTCAGGAAACCCAGAAGGCAAACTCTAAGGCTTAAAGAATCCAGTAATTTAAGTCGTCCT
[0392] TATTTTAAAGAAATAAGGACGGTTCTCCCTTCAAGTGGGCGCAGCTGCTACCTGGTATGA
[0393] CAGTACAGGGCCAGGCAGTGGGCTGTTTAAGGCGCATGCAGCCTAAAGATGATGGCTACC
[0394] AGATTTCCTAGGCTGAGCTCAAAGCGCAGTGAGGGGTACTTTCAAAGGTTAAACAATGT
[0395] TTGCAAACATGCTCCCCCATGCCACGTTAAATGTCCCTCAACACTAGCCTGAAAGTTCCA
[0396] AATGCCTATACTGACGCTGCAGCCATGCAAGAAAGCTGAATAAAATCACACCACGGGCAG
[0397] AGAGAGGTGCAGCAACTGGCAAAGCACCTGTGTGTGCTGGCTGCACAAAACATGCCGCAC
[0398] AGTTCCGCTAAGCTAAATCTGTCCGTGTCCCCAGACTGCACCCAACTTTTAAGAAAGGAA
[0399] AGAAACTGAGGCCAAAAGAATTAAAGTCGTGTGGGGACAAGCCTGAATTAAGAGCCAGAG
[0400] TGGCAGAAAACCAGAGTGCTGCTGCAAAGAAGATGCTAGAGGGAGAAAGGAAAGACAGGG
[0401] ACGGGGGCTGGGGCTTGGGGAGGGCCAAGGTTGGGGTCTCACAGCGGCGCCTCTGAGCGC
[0402] GACTCGAAACTTCGAGGCCAGCAGTCTCCACCCCCGGCCCTTTTCCCACTGTGGCGAGTG
[0403] CCGCTGAAGAGCCTGCCCTTGGCCCCTCGCTCTCTCACTCACCTCGACATGAGCCTCCAC
[0404] ATCTCGCGCTGCCCCATCGCCAAACACTCCGAAGCTAAAGCAGCAGAGCGAGAATCTCCC
[0405] GGACCGTTCCAGCGCCTCGCGTGAGCCCCGCCCACCGCCGTCCTGCGCCGCCGGGCACAG
[0406] ATGACAAGTCCTCCAGGAAGCCAGAGCGACCGTTTCCGCTACGCGACGGGGAAGGGCGGG
[0407] GCAAGAACGACGCCTGGAGGAAATAGTTGAGGGAGAGGAAGGGATCGGGGACCGGGCCAG
[0408] GGAGAAGGCGGAGAGCGAGCTGAGGCGGAGCGGGAAGAGGGAGGATAGAACGGACTAGGG
[0409] CGGATTCGCAGGAAAGGAGATTGCCCTCTAGAGGCTGTCTTAGCCCAAAGCGCAACCTGT
[0410] TGTCTAGGCTTGGCCTGCCAATTTGACCAGCAGGGACTGATGTGAAAGACTTTTCTGGAA
[0411] TTTTAGTAATTTTTTATTAAAACGTAAACTTAGTATTCTTTGAGACTCAAGTAGAGGAAT
[0412] TACAGTTGATAAAGGACATCGATGTAGTTTATTCATTCCTAACCCAATACAGCCGAATTC
[0413] CTTTAATATTAGGATTGAGACATTTTCATTTTTCTTGGGAAAAGTTACAGATCTCCCACT
[0414] CTTCCTCAAAGACACACAAGAGTTGTCTGATTAAGACATAATAGATATGTCTTCCCAACT
[0415] AAAGTAAAAGTTGCTGTTATCAGTGGAAGCATAATTACATTCAGAAAGTAGCTTACTCTT
[0416] AAATGAAACCAACATTTAACTCAACTACTAGGATCCACTGTTAAATTGAGGAAAACGTTC
[0417] CTCCGCTTCCAGATCCACCCAGATGTGTAGCAGAAGACATTTCATCGCCACTATTAAGTA
[0418] TTGACACATTTATCAGAAAGAAGAAATCCGGTGACAAATTTCAGAGAGGACACATGAAAG
[0419] GAGTTATTCTTAGTACCTTGATATTTAAAGAGATACACTGTTGAGACTTTGATCAAAAAT
[0420] TACGTTGCTAACCTAGGTCAAATGTAATAACCAAATGCCCACTGAGAAGAGGACAGACAA
[0421] CAGACATCACAGAAGGAACTCCTTACCTGCAATTTGGTTTTTAACTGGCTACAGGACATC
[0422] ACCATCTGACATCACCATCATGCCTCATTCACTGCGCACAGAAATGAACCATTCATCTGA
[0423] CCTAGCATGGTTTACTGGAAAAAGCATGGGCTTTCCTGTTAGTCAGACTGAGATTCAAAT
[0424] CTTCCTTGCCGCCCTAGCTTCATGTGTGTGAAACCTGAGTCTCACAAGGCCTCATCGTTA
[0425] GAAGGGCCCTGTGCTTGTTTTAATGCTGTGCTATTAATAATTTTTGAACAAGGAAACACA
[0426] AATTTTCATTTTACACTAGGCCTGCAAATTATGTAATCAGTCCTACCCCTAGCACACTAA
[0427] TAGTTTAATGACCTTGAACAAATTATATAGCTTTTCAGAACCTCCATTTTCTCATCTGAA
[0428] AATGGATGTGCTGATGATTAGAGGTGTTTGTGAAGATCCATTGAAATATAGAAATATGTG
[0429] CATAAATTACTTAGGACTTCGTAGGTATCAAGACCTGTCAGTTATGCCACACCTTCCACC
[0430] AAGCTGTCCTTCCTATTCCTAAATTCTCTATTTCTTCTAATGATGCTATCGTCTTCTGGC
[0431] TTCCCGCTCTAGAAACCTCAGTCCTTTAAAGAAAATTATCTTAACATTCTATTTTGAAGA
[0432] TTTTCAAACATATAGCAAAATTATAAAAATTTAAAGCAAACGCCTATGTGCTCACCACAT
[0433] AGATTCTACCTAGATTTACATTTTATCTTGCTTTATTCCTTATCCCAAATACCTTATTTT
[0434] TAAATGCATTTCAAAGTTAATTACAAGCATCGTTTTACTTCTCCCAAAATATGGCATGTG
[0435] TCTTACTACCTAGAGTTAAATATTTGTTTACATATTTTCCCTTTTGAGACAAAATTTACA
[0436] TACAACAAAATGCACAAATCTTAAGTGTACATTCACTGAGTTTGACAAATGTATACACTT
[0437] GTGTAACCAAAATTCCTTAGATATAAACCATTGTCATCACCCCAAAAAGTTTCCTCTCAC
[0438] GATTTCCCAGACATTCTCTGCCTACCACTGCCTTCCCCACCCCAACCCCCGCCCCAGAAG
[0439] CAACCACTGTTCTGTCTCGAGAACCACTGTTTGTTTTGAGACAGGGTCTCACTCTATCAC
[0440] CCAGGCTAGAGTACAGTGCCACAGTCATAGCTGACTGCAGCTTCAACCTCCCTGGGCTGA
[0441] AGCAATCCTCCCACCTCAGCCTCCTGAGTAGATGGGACTACAGGCACATGCCACCACACC
[0442] CAGCTAATTTTTGTATTTTTTTTTTTTTTGCGAGATAGAGTCTTGCCATGTTACTCAGGC
[0443] TGGCCTCCAACTCCTGGGCTCAAAAGATCTCCCCGCCTCGATCTCCCAAAGTGCTAGGAT
[0444] TACAGGTGTTAAATACTGCACCCAGTCTGATTTTTCTCTACCATCAATTAGTTTTGCATC
[0445] TTCTGAAAATTCATATAAAATGGAATCATACAGTGTGCGCTTTTTTTGTATAGCATTTTT
[0446] CACTTAGCAAAATGTCCATATCGTGTGTATCGGTAGTTTCCTCCATTTTATTACTGAGTA
[0447] GTATTTCATTATATGAATACCCCATAGTTTACTGATCTGTGCTCCTCCTGATAGATTCCT
[0448] GGGCTGTTGCCAGTTGGGGCTATTATGAATAAAGCTGCTGTAAACATTCTTGTGTAAATC
[0449] TATTGTAGACATATGTTTCCTATTCTCTTGGGTAAATATCTATAAATAGAATGTCTTGGT
[0450] TATAGGGTACATCAATGTTTGGTTTTATGTGAAACAGCCAGACCCTTTTTTCAAGAGTTT
[0451] GTATCATTTTGGAAAAGAATTCTGGTGGCATACAATCCACAGAATGAGGAATAATGTTTA
[0452] TAAATGATATATTTGATAAGGGCTATCTAGAATATGTAAAGAACTATCACAATTTAAAT
[0453] AACCAAATTTTAAAATGGATGACAGATTTGAATATACTATACAGCTCTCCAAAAAAGATA
[0454] TATGAATGATCAATACTCATATGAAAAGATGCTCAACATTATTAGCCATCGGGAAATACA
[0455] CATCAAAACCACAATGACTTCACACCCACCATAATTTTTTAAAAAACAAGTGTTAGTGAG
[0456] GAGAATTTGGAATCCTCAGACACCACTGGTGATGTACAACAGTACAGCTATCTTGGGAAAA
[0457] ATGTCTGGCAGTTCCTCAAATGGCTAAACAGAGAGACCATATGACCCAACAATTCCATAC
[0458] CTAGGTATTTACCCAAAATAAATGAAAACATAGTCCTCACAAAACTTGTACATAATGTT
[0459] TATGGCAGCATAATTGCCAAGAAGTAGAAATAACTCAAATGTCCATCAACTGATGAATAG
[0460] GTAAACAAGTTATATATCTGTACAATGAAGTATTATTTGGCAAGAAACAGAAATGA
[0461] AGTATCGATACATGGTACAACATGAGCAAACATGGAAAACACTGTGCTAGATGAAAGAAG
[0462] CCAGTCACAATAGACCACATATTATATGATGCCATTTACATGAGTGTCTAGAATAAGCAA
[0463] ATGTGTACAGATAGAAAGTAGACGAGAGGGTTGCCTAAGGTTGGAGAGGAGGCAGGAGAAT
[0464] TGGAGAGTGACAGCTAAGCGGCACAGGGTTTTCTGGTAGAGTAATGAAAATATTCTAAAA
[0465] TGAATAGTGGTGATGGTTGGACTACACTGAAATGAAAATATTCTAAAATGGATAGTGCTG
[0466] ACAGTTGGCTGAATCATATACTTTATATGGGTGAATTGTGTGGTATGTAAATTGTATCTT
[0467] AATTTTTAAATAAAGAATTCTGGTTGCTCCATATTCTCAGCCACATTTTGTATTATCAGG
[0468] CTTTTAAATTTTAGTCATGCTGGTGTGTGTGTGGTGTTATCTCCTTGTGGTTTTAATTTT
[0469] CATTTTCCTGGTGTCTAATGATGTTGAGCACTTTTTCATATGCTTATTATCTTCTTTTAA
[0470] TTTTCTTGCCCATTTTTAAATTGGGGTTTTTGTTGTTGTTGTTTTGTTTTTTGTTTTTTT
[0471] TTTTAGACAAAGTTTTGCTCT
[0472] SEQ ID NO:3
[0473] AGCCATGCAAGAAAGCTGAATAAAATCACACCACGGGCAGAGAGAGGTGCAGCAACTGGC
[0474] AAAGCACCTGTGTGTGCTGGCTGCACAAAACATGCCGCACAGTTCCGCTAAGCTAAATCT
[0475] GTCCGTGTCCCCAGACTGCACCCAACTTTTAAGAAAGGAAAGAAACTGAGGCCAAAAGAA
[0476] TTAAAGTCGTGTGGGGACAAGCCTGAATTAAGAGCCAGAGTGGCAGAAAACCAGAGTGCT
[0477] GCTGCAAAGAAGATGCTAGAGGGAGAAAGGAAAGACAGGGACGGGGGCTGGGGCTTGGGG
[0478] AGGGCCAAGGTTGGGGTCTCACAGCGGCGCCTCTGAGCGCGACTCGAAACTTCGAGGCCA
[0479] GCAGTCTCCACCCCCGGCCCTTTTCCCACTGTGGCGAGTGCCGCTGAAGAGCCTGCCCTT
[0480] GGCCCCTCGCTCTCTCACTCACCTCGACATGAGCCTCCACATCTCGCGCTGCCCCATCGC
[0481] CAAACACTCCGAAGCTAAAGCAGCAGAGCGAGAATCTCCCGGACCGTTCCAGCGCCTCGC
[0482] GTGAGCCCCGCCCACCGCCGTCCTGCGCCGCCGGGCACAGATGACAAGTCCTCCAGGAAG
[0483] CCAGAGCGACCGTTTCCGCTACGCGACGGGGAAGGGCGGGGCAAGAACGACGCCTGGAGG
[0484] AAATAGTTGAGGGAGAGGAAGGGATCGGGGACCGGGCCAGGGAGAAGGCGGAGAGCGAGC
[0485] TGAGGCGGAGCGGGAAGAGGGAGGATAGAACGGACTAGGGCGGATTCGCAGGAAAGGAGA
[0486] TTGCCCTCTAGAGGCTGTCTTAGCCCAAAGCGCAACCTGTTGTCTAGGCTTGGCCTGCCA
[0487] ATTTGACCAGCAGGGACTGATGTGAAAGACTTTTCTGGAATTTTAGTAATTTTTTATTAA
[0488] AACGTAAACTTAGTATTCTTTGAGACTCAAGTAGAGGAATTACAGTTGATAAAGGACATC
[0489] GATGTAGTTTATTCATTCCTAACCCAATACAGCCGAATTCCTTTAATATTAGGATTGAGA
[0490] CATTTTCATTTTTCTTGGGAAAAGTTACAGATCTCCCACTCTTCCTCAAAGACACACAAG
[0491] AGTTGTCTGATTAAGACATAATAGATATGTCTTCCCAACTAAAGTAAAAGTTGCTGTTAT
[0492] CAGTGGAAGCATAATTACATTCAGAAAGTAGCTTACTCTTAAATGAAACCAACATTTAAC
[0493] TCAACTACTAGGATCCACTGTTAAATTGAGGAAAACGTTCCTCCGCTTCCAGATCCACCC
[0494] AGATGTGTAGCAGAAGACATTTCATCGCCACTATTAAGTATTGACACATTTATCAGAAAG
[0495] AAGAAATCCGGTGACAAATTTCAGAGAGGACACATGAAAGGAGTTATTCTTAGTACCTTG
[0496] ATATTTAAAGAGATACACTGTTGAGACTTTGATCAAAAATTACGTTGCTAACCTAGGTCA
[0497] AATGTAATAACCAAATGCCCACTGAGAAGAGGACAGACAACAGACATCACAGAAGGAACT
[0498] CCTTACCTGCAATTTGGTTTTTAACTGGCTACAGGACATCACCATCTGACATCACCATCA
[0499] TGCCTCATTCACTGCGCACAGAAATGAACCATTCATCTGACCTAGCATGGTTTACTGGAA
[0500] AAAGCATGGGCTTTCCTGTTAGTCAGACTGAGATTCAAATCTTCCTTGCCGCCCTAGCTT
[0501] CATGTGTGTGAAACCTGAGTCTCACAAGGCCTCATCGTTAGAAGGGCCCTGTGCTTGTTT
[0502] TAATGCTGTGCTATTAATAATTTTTGAACAAGGAAACACAAATTTTCATTTTACACTAGG
[0503] CCTGCAAATTATGTAATCAGTCCTACCCCTAGCACACTAATAGTTTAATGACCTTGAACA
[0504] AATTATATAGCTTTTCAGAACCTCCATTTTCTCATCTGAAAATGGATGTGCTGATGATTA
[0505] GAGGTGTTTGTGAAGATCCATTGAAATATAGAAATATGTGCATAAATTACTTAGGACTTC
[0506] GTAGGTATCAAGACCTGTCAGTTATGCCACACCTTCCACCAAGCTGTCCTTCCTATTCCT
[0507] AAATTCTCTATTTCTTCTAATGATGCTATCGTCTTCTGGCTTCCCGCTCTAGAAACCTCA
[0508] GTCCTTTAAAGAAAATTATCTTAACATTCTATTTTGAAGATTTTCAAACATATAGCAAAA
[0509] TTATAAAAATTTAAAGCAAACGCCTATGTGCTCACCACATAGATTCTACCTAGATTTACA
[0510] TTTTATCTTGCTTTATTCCTTATCCCAAATACCTTATTTTTAAATGCATTTCAAAGTTAA
[0511] TTACAAGCATCGTTTTACTTCTCCCAAAATATGGCATGTGTCTTACTACCTAGAGTTAAA
[0512] TATTTGTTTACATATTTTCCCTTTTGAGACAAAATTTACATACAACAAAATGCACAAATC
[0513] TTAAGTGTACATTCACTGAGTTTGACAAATGTATACACTTGTGTAACCAAAATTCCTTAG
[0514] ATATAAACCATTGTCATCACCCCAAAAAGTTTCCTCTCACGATTTCCCAGACATTCTCTG
[0515] CCTACCACTGCCTTCCCCACCCCAACCCCCGCCCCAGAAGCAACCACTGTTCTGTCTCGA
[0516] GAACCACTGTTTGTTTTGAGACAGGGTCTCACTCTATCACCCAGGCTAGAGTACAGTGCC
[0517] ACAGTCATAGCTGACTGCAGCTTCAACCTCCCTGGGCTGAAGCAATCCTCCCACCTCAGC
[0518] CTCCTGAGTAGATGGGACTACAGGCACATGCCACCACACCCAGCTAATTTTTGTATTTTT
[0519] TTTTTTTTTGCGAGATAGAGTCTTGCCATGTTACTCAGGCTGGCCTCCAACTCCTGGGCT
[0520] CAAAAGATCTCCCCGCCTCGATCTCCCAAAGTGCTAGGATTACAGGTGTTAAATACTGCA
[0521] CCCAGTCTGATTTTTCTCTACCATCAATTAGTTTTGCATCTTCTGAAAATTCATATAAAA
[0522] TGGAATCATACAGTGTGCGCTTTTTTTGTATAGCATTTTTCACTTAGCAAAATGTCCATA
[0523] TCGTGTGTATCGGTAGTTTCCTCCATTTTATTACTGAGTAGTATTTCATTATATGAATAC
[0524] CCCATAGTTTACTGATCTGTGCTCCTCCTGATAGATTCCTGGGCTGTTGCCAGTTGGGGC
[0525] TATTATGAATAAAGCTGCTGTAAACATTCTTGTGTAAATCTATTGTAGACATATGTTTCC
[0526] TATTCTCTTGGGTAAATATCTATAAATAGAATGTCTTGGTTATAGGGTACATCAATGTTT
[0527] GGTTTTATGTGAAACAGCCAGACCCTTTTTTCAAGAGTTTGTATCATTTTGGAAAAGAAT
[0528] TCTGGTGGCATACAATCCACAGAATGAGGAATAATGTTTATAAATGATATATTTGATAAA
[0529] AGGCTATCTAGAATATGTAAAGAACTATCACAATTTAAATAACCAAATTTTAAAATGGAT
[0530] GACAGATTTGAATATACTATACAGCTCTCCAAAAAAGATATATGAATGATCAATACTCAT
[0531] ATGAAAAGATGCTCAACATTATTAGCCATCGGGAAATACACATCAAAACCACAATGACTT
[0532] CACACCCACCATAATTTTTTAAAAAACAAGTGTTAGTGAGGAGAATTTGGAATCCTCAGA
[0533] CACCACTGGTGATGTACAACAGTACAGCTATCTTGGAAAAATGTCTGGCAGTTCCTCAAA
[0534] TGGCTAAACAGAGAGACCATATGACCCAACAATTCCATACCTAGGTATTTACCCAAAATA
[0535] AATGAAAACATAGTCCTCACAAAAACTTGTACATAATGTTTATGGCAGCATAATTGCCAA
[0536] GAAGTAGAAATAACTCAAATGTCCATCAACTGATGAATAGGTAAAAACAAAAGTTATATA
[0537] TCTGTACAATGAAGTATTATTTGGCAAGAAACAGAAATGAAGTATCGATACATGGTACAA
[0538] CATGAGCAAACATGGAAAACACTGTGCTAGATGAAAGAAGCCAGTCACAATAGACCACAT
[0539] ATTATATGATGCCATTTACATGAGTGTCTAGAATAAGCAAATGTGTACAGATAGAAAGTA
[0540] GACGAGAGGTTGCCTAAGGTTGGAGAGGAGGCAGGAGAATTGGAGAGTGACAGCTAAGCG
[0541] GCACAGGGTTTTCTGGTAGAGTAATGAAAATATTCTAAAATGAATAGTGGTGATGGTTGG
[0542] ACTACACTGAAATGAAAATATTCTAAAATGGATAGTGCTGACAGTTGGCTGAATCATATA
[0543] CTTTATATGGGTGAATTGTGTGGTATGTAAATTGTATCTTAATTTTTAAATAAAGAATTC
[0544] TGGTTGCTCCATATTCTCAGCCACATTTTGTATTATCAGGCTTTTAAATTTTAGTCATGC
[0545] TGGTGTGTGTGTGGTGTTATCTCCTTGTGGTTTTAATTTTCATTTTCCTGGTGTCTAATG
[0546] ATGTTGAGCACTTTTTCATATGCTTATTATCTTCTTTTAATTTTCTTGCCCATTTTTAAA
[0547] TTGGGGTTTTTGTTGTTGTTGTTTTGTTTTTTGTTTTTTTTTTTAGACAAAGTTTTGCTC
[0548] TTGTTTTCCAGGCTGCAGTGTAATAGCACAATCTGGGCTCACTGCAACCTCCACTTCTGG
[0549] CTAATTTTGTATTTTTAGTAGAGATGGGGTTTCTCCATGTTGGTCAGGCTGGTCTCGACC
[0550] TCCCAATCTCAGGGTGATCCACCTGCCTCGGCCTCCCAAAGTGCTGGGATTATAGGCGTGA
[0551] GTCACCGCACCTGGCCAAATTGGGTTATCTTTTTATTCGTGTTTGATATTGTCTTTTTTT
[0552] AATTATTGTTCAATTATCTATTTCTGTTACAAATTACCCCTAAAATTTAGCAGCTTAAAG
[0553] CAACCAACATTTATTGTTTCTCAGAGATTCCAAGGCTCAGGAATCAGGGAGAGACTTAGC
[0554] CAGGTGGTTCTGGCTCAGGGTCTCTCATGAGGCTGCAAAGCGTTGGCTGGGATTGCAGTC
[0555] ATCTGAATGCTCAGCTTCCAAACTTATGTGGTTGTTGGCAGGTTACAGGTCCTTTCACGT
[0556] GGGCCTCTCCATGGGGCTGCT
[0557] SEQ ID NO:4
[0558] GGAAGGAGGATCCCGAATCCCAGCCAGAACTGAGAAAGCCCTCAAAGGTGCGGAAAGGGG
[0559] ACTTTTCCCTCAGGAAAGCCGGCAACAGCAGAGGCCCTAGCCCACGTCGCCACACCCACT
[0560] CCGCGCGCGCCCCTCGCCTCCCAGGGCCGGCTCTGCTCGGCCCCGCGGCCTCTCGGGCGC
[0561] CCCCAGCCCGCCGGAAAGAAAGAAGCAGAAACCCGGAGCCTGGGTCCCACCCGCGACCCC
[0562] TCACCTTCGCCCTTCTCGTCCTCTCACCTGCCAGTCCCCCAGGAAGAACAGGACGCCTCT
[0563] TCCCCCTGATGGGCGGCCACAGGCTCGGATCCTTCCATTGCCGCCTGAGACACCGCCGCC
[0564] GGCTACTGGAAGGTAGGAAGGGGCGGGACCGTGGGGGGGTCAAGGGGCGGGCGGAGACGT
[0565] CATCAGAGGGGGCGGGCCTGGGGAAGCTAGGAGCCCGGCAGCGCCTTCCCGTCAACCCTA
[0566] GGGGCGTCCTGGTTTCCGGTTTGGGTGTGGCCGCATGGCGTGCTGTGGTGCAGGTGGCCG
[0567] AAGGGGGCGTTACTGTTGCGACTGGCATCCGCATCCGGCAGATGTAGATGGAACCAAAGT
[0568] CCAGAAGTTACGCGTCACCCTTGCTCTACAGCCAAACATGCAGGACTCTAGTAACCCGCG
[0569] AAATGATGGGATAGCGTTGCAAATCCTTAAAAGAGTCTTAACGGTAAGAAGAGGAAACAG
[0570] CTTTATTTTATAAATAAACTGAAGGCTGCTAATAGCTCCAGATTGTCAGTGAGGGGATAC
[0571] CACTTAAAGATGTTATACATTTAATAAATAATTTTGGGAGTCCACTACATGTGAGGCGCT
[0572] GCCTCAGATTTAATAGAATCATGTGACTTTTAGTGCCATAGATGTCGTTAGATCATCCAG
[0573] TTTGGTCCCTTAAATTTGTGGGTAAGGGAAAGGAACTTCCGGTCAGTGCAAGTGACTTGG
[0574] TCAAGGTTATAGAGCTAGTCACTGTTAGATTACCAGAACACATATCTAAAAAGTTCTCAT
[0575] ATAATTCTTTTTCATTTACTTCCGTGTCACCTTTCTTAGAAAGCAAAAATGAGTTGGACA
[0576] CAATAACATTAGGATTTGTCTTTCACTAAATAACACTGTATATTTGAATAAACATCTGCA
[0577] GGATTTTTAAAATTAGCATATACTTTGGGTATTATGTAATTTTCTGTGATGTATCATATT
[0578] TTTAATTAGAAATTTTTCCTCTCATATGGAGATGAGGGTCCAGGCAAAACTGTATTATCC
[0579] TTCGTGCCCTCTATTCTCTTCAAAATTGCACATCTTAAAATACCTGGAAGCTTTATGTTG
[0580] CTCACCTATAGCATTATCCGTGTTCCTCAATCCAGTTGTGAACCTTTTCCACATTAAAAT
[0581] GTTCTTGATCTGTGTTCTATTCTTTCAAGGCTCTGCTCTAGCACTAAGATAATGCTACAG
[0582] GACTGAAGATAATGAAAAAGATTGAAGAAAGCAGAAACAGAACAAAAATCAGATTATTTC
[0583] AAAGTTTCTTTCCTTGTAAAGGTTAAAACAGAGCATACTTCCTTATTATGCCCACTAGGT
[0584] TAACTAGAATCTCCTGTTTTTTTAAAAAAAACTGGCTCATTTTTTCAAAATCCACTTTAA
[0585] TTACATGGTGCTTAGCACAAGCAACTCCATTCTGGTTTGATCTGATGGGGCCTAGTGTAG
[0586] AAGGCTAGGCCAAAACAATGGCCTGCCATTACCCACGCTTCTTCATATTTTAACATCCAT
[0587] GAATAGAGGATGAATTTTTAAAAAATAAATACAGAAAACGATGCATCTTAAAATTGAGGG
[0588] CATGTTAGATTTGATAAAATATGGTGTGAATTATTTCAGGCATGTAACAAGCAACCAGAT
[0589] TTAACAGATTTTAATGTTTTGTCACATATGTAAAATGGGAAAGAAAACTAGGCGATCACA
[0590] GCTTTTTGACAGTGAGGCTTTATGTCATAAGAAAGTAGAATGACATATTGAGGAAACTTA
[0591] AGGGAAAAATATGTAGGCCAAAGATTTTACATTCAGCAAAACTAATATACAAAAGGTACA
[0592] AACTGTTATCAACAATTGATATATCAATTTGGTGAGACTAGTGTTTTCATGAGCCCCTCC
[0593] TTAATAATTAATTGGAGAAAGTGTCTGACAATCAGAATTATTACAGAAACATTGATCTAA
[0594] GGATTGACTCGGTTTCTTTACAGCAAATCATAATCTGTTGCATTTATTAGCTTTTGTGTT
[0595] TCTTGTAGTTTATTCTTTATTACTAATTCTCCTGAGTTCTGTCACCTCATTTCTGAGGTT
[0596] TTTCTAATTCAGAATCAAGTTCTTCCATATCTTTTATCATTTCTCTAATGTCATTTAGCT
[0597] AGTTTCAAAATAATAAATTGTTTTGTAGTGCATCTTTCAGGCATGTTTTTATTATCTTTG
[0598] ATGATGATGTTATATTCCTTATTCTCTTCACATAATAACTGCATGGATTTGAACTTGATG
[0599] CTTTTTTTGTTGCTCTTTTTTTTTTTTTTTTTTTTTTTTTTTTGAGACAGAGCCTTGCTC
[0600] TGTTGCCCAGGTTGGAGTGCAGTGGCGGGATCTTGGCTCACTGCAAGTTCTGCCTCCCGG
[0601] ATTCACGCCATTCTCCTGCCTCAGCCTCCCGAGTAGCTGGAACTACAGGCACCTGCCACC
[0602] ACGCCCGGCTAATTTTTTGTATTTTTAGTAGAGACGGGGTTTCACCGTGTTAGCCAGGAT
[0603] GGTCTCGATCTCCTGACCTCGTGATCCACCCCTTGGCCTCCCAAAGTGCTGGGATTACAG
[0604] GCGTGAGTTATTGCTCATTTTTATGTGAAATTGGTGTTCCTGAACTAATATGAGACAGGC
[0605] TTCTTAGCCCAATAAGCCTAAGTTTATCAGGAACAAAACTGTATTCCAAGAAAGTCAGGT
[0606] ATATTAGTCTGTTTTCACACTGCTATAAAGAAATACCCGAGACTGGGTAATTTATAAAGG
[0607] AAAGAGATTTAGTTGACTCACAGTTCCACATGGCTGGGGAGTCCTCAGGAAACTTAAATC
[0608] ATAGCAGAAGGTGAAGGAGAAGCAAGACATATCTTACACGGTGGCAGGAGAGAAGTGAGA
[0609] TCGTGGGAAAACTGCCACTTTTAAAACCATCAGATCTCCATGAGAACTCCCTCATTATCAC
[0610] AAGAATAGCATGGGGGAACTGCCCCCATAATCTAATCCACCTCCCACCAGGTCTCTCCTTG
[0611] CACACATGGGAATTAAAATTTGAAGTGAGATTTGGGTGGGGACACAGAGCAGACCATATC
[0612] ATCAGGTTTACTGGCTCATTGCAAGGAGGGAGCCCACATTCCAGAAGAACCATGGGGTAC
[0613] CTCACCAACAGAGGAAGAGATAGATTTATTGTAGGAATTTGGGGAAGTATGGAGTTTAGA
[0614] AGGAATTGAAATGAAGTAGTGCTTTGATAGACTCAAAGAAAGCAAGGCTGTTTGTAAAG
[0615] AATTCTACACTACATCTAAACTATTACTAATGGCAAAACTGCAATTACTTTTGCACCAA
[0616] CGTAATACAAGATGGACCCAGGGTTCTGTTTTTTGGAAATTGCAAAGTTCTGCCTCTCAG
[0617] AAGTAGAAACTATTTCTCTGTGTCAATGTGACTTTAGATATTCTAGCCAAGAGTGAGATA
[0618] TTTCAATCTTAATAATAGGATTTAAACAGCAGAGTTTTTGACAATCTTTTGATATTGTCA
[0619] TAAATGTTCTATGATTTTAGAGAAAAGGGCATTGTGTGAGTAAAAAATCAGTTACTCTAA
[0620] GGGTGTGTGTGTGTGTGTGTGTGTGTGTGTGTGTGTGTGTGTGTGTTTGAGACATTTTAAAG
[0621] CATCTTCCAGAGAGAGAAATAGTATTTTCTGTTAATTTCACAGCTGGCTTTATTTATGCCTG
[0622] TTATCCCATAGCTGATGAATGGCAAAGCAGATTTTTACTTTCTTATTCTATGCTTATTTG
[0623] CTTTCTCACAGGTTTTGGGAAAACTTTTCTAATTTCAGAGAATTCCCTTTTTCTGTTGCTT
[0624] TTGTGTAATGTTCAGCAATATTATGGTTTTATTTCCTTTTTAAATTAATACATACATAAT
[0625] AAGCATATATTTATGGGGTATTTTGATATTTGCGTACAAGGTATAATGATAAAACCAAGG
[0626] TAATTGTGATATCCATCACCTCAAACATTGTTTATTTCTTTGTGTTAGGAACATTTCACA
[0627] TCTTGCCTTCTAGCTATTTTGAAATATATGATAAATTATTGTTAACTATAGTCACCCTCT
[0628] AGTGCTATCGAACACTAGAAATTGTTTCCTCTATCTAATTGTATTTTTGTACCCATTACG
[0629] CAACCTTTTTCATCCTCCTTACCCCCAACCTTCCCAGCCTCTGGTAACCATCATTCTACT
[0630] CTCTACTGCCATGAGATCCACTTTTTTAGCTCTCACTTATGAGTGAGAATATGTGCTATT
[0631] TGTCTTTCTATGCCTGGCATTTCACTTAACATAATGACCTCCAGTTCCATCTATGGTGCT
[0632] TCAAACTACATGATTTCATGTTTTTGTGTAGATGACACATTTTTAAAATCCATTCATCTG
[0633] TTGATGGACACTTAGATTGATTCCATATCTTGGCTATTGTGAATAGTACTGTAATAAACA
[0634] TGGGCATGCAGATACCTTTTAAATTTATTGAATTCCGTTTTTGGATGTATACTCAGGAAT
[0635] GGGATTGTTGGATCATGTGGTAGATCTATTTTTAGATTTTTGAGAAACTTCCATACTGTT
[0636] TTTCATAGTTTCTGTACTAATTTACGTTCCCACCAGCAATGCACTAGTGTTCTCCTTTTG
[0637] CCACATCCTCACCAGCATCTGTTATTTTCTGCCTTTTTGGTAATAGTAATCTTAACTGAG
[0638] GTGAGATGATGTCTTATTGTGGTTTTGATTTGCATTTCCCTGGTGACTGGTGATGTTGAG
[0639] CATTTTTTCATATGTCTGTTGGCCATTTATATGTCTTCTTTTGATAAATGCCTGTTTAGA
[0640] TCATTTGCCCATTTTAAAATTGGATTATTTGGGGGTTTTTGCTATGGAGTTGTTTGAGTT
[0641] CCTTATATATTCTGATTATTA
[0642] SEQ ID NO:5
[0643] TAACCTCTCTGTTCTTTTCCGTTCTTATCTATAAATAGGGGCAACCAACCACCAGTCCTC
[0644] AGAGGCTCTGAAAATCTAATTTCAATTGTGGTTGAAAGTGCTCTGAAAATTTGAAATGTT
[0645] ACACAGATGTATGGGATTCCTCCTGAGCTGTATTTGAAACCACAGATATCCATCTGATTT
[0646] GTAAATCATTTACCTTTTAGTCTCCCTAGCAGAAGTTAATAGTGAATGCTTTTCTTTTTT
[0647] TTCTTTTTATTTTCTTGATTGCTAATTTTGTTAAGTTCATAAGCGATACTTACATATGTA
[0648] CATGTGGAGGCTGGAATCTGCTCTGGTTGATGTTGTAGTGCGTGAATGCCTGGGTGGGGT
[0649] TTGAGCTGTACTCATCCACCGTTATGGTAACTTTGCCTTGAGAAGGAATCCCATGTGCCA
[0650] TCCTTCCTCCCTATAGGCATGAGCAATTCACAACTTTACAGAGCAAAAGATGGAAGCAGG
[0651] CATGGTGCGTTTCAGCCAACCTCTGAATCCAAGGATCTACTCATCAATACCGCTGTGCAT
[0652] TTCTGAAAGAAAAAAAAAATGTATATAAAATGAGAACAAAGAGAAAAAGGGTGTTCTAAG
[0653] AGACATAAATTTAACATCTAGACATGAAACTAGTTCAGGGAAAATCTGAATATGTGTTGT
[0654] TGGCTCATGGATACAACTGTTCCCAACTCTCTTCCCTTCATGCACTTCAAAGTGGATATT
[0655] ACACTCTATACATATTGCTGGGTGCTAGCAAATCTTCTCAATGGAGTTCAAACTAAGATT
[0656] TCCTCTTTAATTATAATTATGGCACTTAGCGTGAATTACGGAAGGGGCATATCCACCGTT
[0657] GCCATGGCATCTGGCTATACATGGTTTGTGCTCCAAGTGAAGAATCTTCCAAACCTGCCC
[0658] TGCTAATTGGAATGTTTGGTCTTTACAAAGTTGGTGTCACGTAATCATTTGAAAGCTCCA
[0659] TGTTTAAAGCAGCTATATAACACAATGACAGATAAAGAATTTAAAATGTTGTGCTTATAAC
[0660] AAATACATGTGTATTACTATAACATGCACATAATAACCACACATTACACATATTTTTTTC
[0661] TTTCACAAGCAGCTCCCAAAATACACAGGTAGGCAATGCTTGGCATTCCTCACTGGCCCA
[0662] GAACCTTATGGTAAGGCAAATCTTTCTTGTTAAATGTAGTGTCATTTTTTTCAACACTCC
[0663] AGTCACTCAGGGGAAAACCAAACACCTCCAGGTCTGGACCACCACTTTGAATGCATTCAT
[0664] TACCTCCCTTCTTCAAGCCCCGATGGAGCGGTAAAGAAGGGCGGCCATTACATCTGCCAC
[0665] CTCATTCTTGGTGCAACGTAATTCTCAGTAATTGGTTTCATTTTTTAAAATTTCCATTAG
[0666] TTCTGATTTTTTCATAAATGTCAATATATTATTTGCTGCACTGGGTTGAGAAATGAAGAT
[0667] GTCTATTTTTATGGTAGGTGTGCAAAGGCTGAGGGAAGCAAATTCTCATGTCCAAGTTGT
[0668] TGCAATTGATCAGTGCGGTAGAGAAGAAGCCGTAGTACACTCCTTGGGTAACCCTTAGGT
[0669] AACCCATGCTCTTACTTTAAGTGCACAACGTCAGAATAGGTTTGCCATGTGCACCATTAA
[0670] GCTCTGAGAATCTTATCTGCTATTTCAAATAAAAATATTTAGAATAATAAGATGAGATTT
[0671] AAAATTGGCATGAAACACCTTAAATATCGAGCAAAAGGCAGATGGCCTGCCATAGGCCAT
[0672] TTATTTTATTATTTTACAACCCGGGGCTTGAGTAATCCCCCATATGCTAGCTGGGCTAAG
[0673] AAATTCCAGAGTGTTCTACTCAAATCATTGCCTTCTTCATATTAGCAAAGCAAGCTAAAT
[0674] AACACCCACCTTGCCTGCCTTCCTAATGTCAAGTGCACTGCCAAGAGATGAATCCCCGTG
[0675] CTCTGGCATTTTAAACAAAGTGTCACTTTGGGTTCTCTTTCTTTTTCCCTGACTTCATGGG
[0676] AGGTATGAGTTCTTTTAAGCGCCAGATGATAGACAACCAGGCCAGTAGGATTAACTAAATT
[0677] AAAATCTTTCCCTGCTCCTCTGATGTAAGATATCTATAAGAACTGCTTACAAATCATAAT
[0678] GTAGCAGTTTTCATACCAAGTTTGCAGCTGTTTCCTTAATTTCTCCCACAATTATATTTT
[0679] ACTGTTATTTTGCATAACAGTAACACACAACACCTGCTTCAAATTGCAGTCCTGTTCTAA
[0680] GTACTGTAATTCTCTTTTGAGACACAAAGTACCGAGACACTACTGAAGCGCAAAACTGCA
[0681] TAATTGCGTCGCGCTGAAAAGTAAATACAAATGAAAACATCTTCCTGTTATAGAAATGAC
[0682] TGTCATTAAAAGAATGTTAGCCATCAAAGCACAGAGGGGGAAAAATTCACATTTTCAGA
[0683] CTGTGAGAAGTTAGTGTTTGCATAAATTAGAGAGTGGTGTAAAAATTACTTGACAATTCT
[0684] TATGTTCATTCTTCCATCAATGTCTATTGTGTGAAATACTGAACATTAAGAAGCAAAGGA
[0685] AATCAAATTCATGTCTAGACCGCTGTGACCTCTGAGTAACAATCTCTGCACCCAGTACTA
[0686] GAAAGATTTTAAATGTTTCTGGAGTAAACACACTAATTAAACCAAGCAAATGTGGCCTAG
[0687] AAATAAAATATACATACCTAATATTTAAACATCTAGGTAATGACTTGTAGCAGTTAGCAT
[0688] GATACCCACCCAAGTTCTACTATTTGAAGCTAAGAAGAGACAGAATTTAGCAGCTGATCA
[0689] ATTTGCACTTCTACATGGTCCGCTTGATCCTAATGTCAACCTTCTCTATCACTGAACATC
[0690] TATGCCTCTTTGCACCTTTGCCCCAGTAAGACAAATAGAACATCTTTTCATCAGCAAGCA
[0691] CAGGCCCCTGTAACTCTGAAAGGCTAAATGGGTGGCTGGCACCTCATTTGAATTCATCTT
[0692] TGCCATGTTCCCTTTCAAATGTTTCATCGAAACAGTGCACCTGCAGCGTCTATCTTCTTT
[0693] AATCTGGAGTCCTTTGAAGAAAATGTCAGACTGAGGCAGAGAGCTTCAAAACTTTCACAG
[0694] CACCTTTACAGACAAGTCAAGAAAAGGCTTATTTTTTAAAATGCTATGTTCATCCCCACC
[0695] CGTATCTCAAAATGAACACATCTTAAGAATTTTAAACACAAACACCTGGAAAAGACTTTT
[0696] TCTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTGGTAAACTCTCAGAGACAAGGGACAAT
[0697] AGACAAAACCACTTCCTATAACCTGAGAAATTTGGAGAGAGCTTGAAACAAGACAGAAGG
[0698] ATACACTGCTTACAACTATCAAGAAGATGTATATTTAGGCAGTAAAACCTTCCAAATGAT
[0699] CAGATTTTCGTGCATCTCACCCTCTTCACACTTACTTTAACTTTTTACACACAATATTTA
[0700] AAGACAGCACTATTCTTTTCATAAATTTTCATTTTATACTTCAACACACTTTCCACAAAA
[0701] CAAAGGAACCCAGGACGAAATATCAAGAAGCCGAAAGAAAAATCTAAAAGGGACTGACAC
[0702] TGCCAAACATTCATGCTGGTGCTAGTTTGATTTTTCTCTTTAAAAAGAAAATCATGTGAG
[0703] TACATTTTGCATTTCCTAGACAGCAATCACTTCTGTCCTTTAAGCTGTACAACTAGCAAT
[0704] ATTCTTTTCCAGGAATCTAGTTGGAAAACCTTTTAAAGTCGGCTGGCAATGAAGGGATTA
[0705] AAAGCTGAGGAATTAAATAACTGTCTCCACCTCCTAGACACCAGAAGTCCAATTTACCAA
[0706] AACTTTCAAAGCAACAACTCATGACCTCTCTACCCCTGCACTGCCCAGCACAGCCTATGA
[0707] GAAAAGAGCCCGCGGTTTTCTCGTTTCAACAAGGATGCTCATTTTAATCACTTGCCCCAA
[0708] GATCTCCAGACCGAGCCTGGCGGGACTTTGGAACTTGCCTGCTCCAAATCCGCCAAATTT
[0709] CGCTCTCTCCTCCCTCTGCCGCGCAGATCCAGTGCGTTCTTAATTCACGCCCCTACTCTT
[0710] TTATGCCCACTTCCAGTTACTCATCCCTGTATCTGGGAGCGATGCCCACACGCCTTCCCG
[0711] GCGGGCTTCTCCGCCAGGGTTTTTTTGGGGGAGGCTGACATAAAGTTCCACGTGCTTCTG
[0712] TCCACCGAGCGCCCTCCCTCCCACCGCACTTACACCTGAACTTGTCTCCAGCACTGCGGA
[0713] CACCCGGGTGACACATTCTTTCGGAGAGGGAGGGCAAGACTTTTCTCCGGAGCCTCTGGC
[0714] AACTCGTGGAGTCTATATAAGGTGGGGCGGTGGGGGGCGCGGAATTCTCTACTTTGGATT
[0715] GGTTTGTTTGTTTTAAATCTAGGGTGGTGGGCAGTTGCGCGAGGCGGGGTGGGGGGTGCT
[0716] TTATGAACTACACACACACACACACACACACACACACACACACACCACATACACTCGCCC
[0717] TGCCCTTCTAGGCTCACGGCTGGCTCTTCCCGGCTACAACTTCCCCTTCCTTCCTGAAAC
[0718] ACCAGCGCCTTACTCGCTCAGTCCCCCGTTGCCACCTCCGTCACCTCCTCGGCCGCCGCT
[0719] CCTGCTGGCACTGCTGGCGCTAGCGCTGGCGGCCCCGGGCTCCGCTGGACGCCTATCTTTT
[0720] TTTCTCTCCCTCTTCCCTAGCTCTGGATTGATTGATGACCACCGGCGGCCAATGGCTCTG
[0721] ACACCACGCCATCCTCGGTCGGGCATCACAGCCCCACGTCCCGACCTCAGCCTCCCGGG
[0722] CGCGTTGGCATCCCCGAACCCACAGCGCTGAGCCCGAGCCCGGGGCCGCCGCCTCCTCCC
[0723] TGCTCTACGCCTGGGCGCTGCCAGGTTCCTTCCCTGGAACGGCGGAGAAAGGAAGGGAGG
[0724] GAGCCAGCGAGCGTTGGAAGGGGGTGTGGAGGCCACGGGATCTCCCCCTCTTCTGACATC
[0725] TTTTTTCACTTAAACTTTTGTTTGTCCCACCCCCCCGACCTGTTTGTGTTTGTTTGCAGC
[0726] ATCTCAAATCAATTATACATT SEQUENCE LISTING <110> Bolcheng (Beijing) Technology Co., Ltd. <120> Markers for detecting lung cancer and their uses and systems <130> PE01733 <160> 5 <170> PatentIn version 3.5 <210> 1 <211> 5001 <212> DNA <213> Artificial sequence <220> <223> Artificial Sequence Description: Synthetic sequence <400> 1 gctccgttgc ccaagctgga gtgcagtggc ttataaaaca tggctcactg cagccttgaa 60 ctcctgggct catgcaatcc tcctgcctca gcctctggag tagctgggac cacagaagtg 120 caccaccaca cccggctagt tttattttta tttgtagaga tgggatctcc ctatgttgcc 180 cagtctggtc tccaacctgg ggcctcaagt gttccccctg cctcggcctc ccaaagtgca 240 gggattacag cttgggccac tgtgcccagc cctttcttcc atcattccgt cctctccttc 300 cttcttaata ctgcttgtca aactgatttt gccacccagt ttgaaaaaga aggtctagag 3'to 360 gattatccca atcattaact atttctgcag gttagcttac ctttccagaa tacagaatag 420 aattcagcaa actgcctcaa agagcagttt ctagatttga cttcttcctc tagggtttat 480 gaagtcatca tactgaagtg acctggattt ccctatacgt gaacctcatg acgtgggagc 540 ccggggtagg gagggacaga taacatgaga atcaatagaa agaactgcaa ggtctgaaga 600 caagaaccat ttgtctagtg gagtctattt tatatccatt ttttaaaaat ctccatttaa 660 It should be noted that in the original text, there seems to be an error in line where "3'to 360" might be incorrect or an incomplete annotation. The translation is done as accurately as possible based on the provided text.gatattagc ttcaacgtgc accttgacca attaatccaa caaacaaaac tctggcttgg 720 gggcccagtt actcactttc atggaaagag acctctgttt ttagaaccag ggatcttgtg 780 agtccctccc acgtccccag cccttctctg gcctcgggta acttttccct ccagtttggt 840 ggaaggttac attagaagaa aatagtccct accgtctaga aagttttcgg gggggggaaaa 900 aagggcgttg gacgtgcggg tgggtgggta taccgggagg ggacagctgg agccgcgaag 960 gcgcggggc ccgcggcgcg tgggtgtgga ggccgcgggg cggcaggagg cactgggaag 1020 cggggtgaag agcaccgaag tccagaaaag ggcagtgtcc gccctgcgct cctctgctgg 1080 cgcccgcaac ccgggaggga gcggcgaggc tccgaagggt gggattcgag gtctgaaccc 1140 tcgagtgggc ttgggttttc aggagccgcc ttttcttatt gatcatttct cttttgtcat 1200 cataatctct tggcctcaga ctgagttaag ggagccatag taggagtgca gaaaatgaag 1260 tggaggaaaaa aaaaatgtga tatttaaaat atagtttttc ctactttctt tttctctctc 1320 aaattaattt agatgtaaat aacattttaa cagggcccaa catcaggctc ctaggctctg 1380 gctgctggtc tcacttgaca ccagaataag atttggtgt gctagtaatt ttgccaaatg 1440 gaggcctact ccccaaaaca tcaccgtggt gcatgatgaa catttctgta ggtgatgacg 1500 tggagttttt gttttacttc attctaatac atgtccctgc ttcaagtggg taaaattata 1560 cctagctcag cccaggttta attcagatta tttaaaaaat ttcagtggca agtaatttac 1620 gagcttctct ttattaagat ttaggaaatt ttgctttatt gcgcatgcat tttaagtatt 1680 ttttgctgt gttcctagtg tagatagaaa acaacctgtc actgtctcct agttggccct 1740 gtcaacaaac atttgagaag tgcttagatc agaatgctgt tgagtactac aaattaaaca 1800 atctgatgta cttgttaatt tcattacaaa cagcagttag gatttcatat acatatagtc 1860 tgtaccatgc tatatatcag aaatctttgt aatataaac tccttaaata tagggaccca 1920 gttttattct tcagcgaaag ggagttaatc ttttggtgtt gtgcatttag aaggatctcg 1980 gatgtaagag gctcattaaa tgatagctgt tagtattatt tatcttggta tcttccaaag 2040 ggctttgcac atggctaaca ctaagaaata ttggttgaat gaatgaaggt gccatgccca 2100 tactgtgata tacacagaga aaagaccagg atggggacac tagcacagtg atttcatttc 2160 atcagtgaag ctttgtgtgg aagtgtccct cctgcagccc ttttaccaga tgcaggctcc 2220 cttgaccttg gattgtggac cacttaaaat tttcagggga aaaagtgtgc ctaattagct 2280 agtagggcct aacctcttgt caagatggaa cttgagaagt ttgtctaagt atggaaaagc 2340 taagctcact tgacacagat cacgggaaga aagatttctg catgttattc aagttattca 2400 cgtcaaaagc gattaggggg ccccagcttc cattacatca ttcattagga aaatgtgtgc 2460 tgggagctct gcaaatcaca tgatttgatg gaggcattga gactgggttt cttgtgtgtc 2520 tgcccttggc catattcagt ttttagtatg tttggaaaag gctgaagcag aacatgaaga 2580 tctagtttat tttaatgctg tcagacagct aaatggtgga aacttgttga agagatttac 2640 tgagctgttt ccttagattg aagactttag agaattaatt aacagggttc tgagagatac 2700 aacatgccag acccctactg actctaagag tgtactttct tactgatatg atgtgatagc 2760 tgagcacagt atatgtgaaa ctgctgggga gtaatacaag agcagtgctg atgtgttcta 2820 agacatgcag gtgttcagat gagaactggc catgacaact gagtcagcca ggcactggac 2880 tcattcctt ttgcttaagc accatgagct cagtagaagc acagactttc acaggtgct 2940 gaactacttg gtgaaaatga agtcacaattt ttggaagc ttcactgatt ttgatgtgta 3000 ggtattgttt cacgttttag ctgttgcttt tagatgcat gccttagat ggttctcagt 3060 tatcatgcta tatctatgcc acgccattc tagagcaagt aacttagctg aaagtgacta 3120 tacctataag aaaaaaata tccctgtaga gatattgta gatgaattgt ggaagtactc 3180 taaaattttt ggccagattt tgggaatgt tgagtctag ctggtacaga aagtacattc 3240 tgttttatga aaatagcaac attgaagtga aagatatca ttaatgtata catgtagtca 3300 gaaattgaat gtgccctgqgc tgattgcaat tattattc acttaatgcg tattattga 3360 gtacaaggcc ttagagtttc ctcagttctt atagtctgta acagtgccat tccataggaa 3420 tgtaatgaga gccacatata taatttacaa ttttctagta gccacattaa aaaattaaaa 3480 agaaactgat aaaatttaatg ttattatta tattttat ttcaccaat atatctaaaa 3540 atatatctaa aatattatcc tatatgtaa ttaatatact tttattggtt tttttttttt 3600 gtttttttt ttttttggg gacagggtct cactctgtca cccaggctgg agtgcagtga 3660 cacaatcatg gttcattgca gcctcaacct tctgggctca gatgctcctc ccacctcagc 3720 ctcgtgagta gctgagacca caggcatttg ccaccagcta atttttgta ttttttgta 3780 gagacaggat ttcccatgtt gcccaggctg gtctccaact cctgggctca cgcgatccac 3840 ccgccttggt ctcctgaagt gcagggatta caggcatgag ccactagaaa ttattaacaa 3900 attattttac attcttcgtt ttgaattaag tttctcgaat ttacttacag caaatctcac 3960 attgcactag ccacatttca ggtactaaaa aactacatgt gattaatggt taccatactg 4020 ggcagtccat gtctagagag aaatacagat gtgtaatttg taaacaggta aacagataat 4080 ttctttcttt ttcttatttt tcttttttt ttgagacaaa ttctcactct cccaggccgg 4140 agtgcagagg ctcagtgcag cctccacccc ctgggctcaa gcgatcctcc cagctaacct 4200 ccccaactcc ccaccccagc agctgggact acaggtgggt accaccatgt ccaactaatt 4260 tttatagaga taggttttca tcatgttacc cagactagat aattttaatt aaatgtaata 4320 aaggctatga tgagaaaagt acactcatca tacaatgaaa gcatagaaaa ggggcttagt 4380 tcctccagtt tttcctccta aaatctcata attctctctg cttctctgca tcctgcctgc 4440 cacgactcta atctaagcca gattatctct ctcctagact tgcacaacag cttccgaaat 4500 tgctacctg catctatccc tgcccctgtt tgcagataaa gtgagctttt aaaaacgccc 4560 atcagatcat gtcattcctc tgtttaaaac cttaaatata tgatgaacac attggacaag 4620 tgtgagatga tgtacataaa gcgcccagcc caaaggttct gctcactgag tttagttctc 4680 ctcgtttctt cctttcactc tctttttct tctcttact ttcctcccca gctccttcct 4740 cttatcacct atcagtctct tatcatctga atttcttttt tctttcttcc ttctcacttg 4800 ctttccatttt ttttcttcca tctttacttt ctctcctttc tgctttagtt tattttaata 4860 tagaatttaa atagatttag aaattcattc tttaattatt tacttctgag ttccaagctc 4920 ttcatttaaa aatcaaaaga gtactgggca ctgtggctca cgcctgtaat cccagcactt 4980 tgggaggccg aggtgggcgg a 5001 <210> 2 <211> 5001 <212> DNA <213> Artificial Sequence <220> <223> Artificial Sequence Description: Synthetic sequence <400> 2 tctccatgag ctggtgcaag gacaactaat tttaaggtat taaaaaaaca aacagaaaag 60 caactcttaa atgtataaca aaaatttact gagaattagg tctgtggata cagtgttaga 120 catctagaca ctttcccctg aaactattct tcttaagtca aggaaacacc tccatccctt 180 ttcaggaaac ccagaaggca aactctaagg cttaaagaat ccagtaattt aagtcgtcct 240 tattttaaag aaataaggac ggttctccct tcaagtgggc gcagctgcta cctggtatga 300 cagtacaggg ccaggcagtg ggctgtttaa ggcgcatgca gcctaaagat gatggctacc 360[[ID=?]] agatttccta ggctgagctc aaagcgcagt gaggggtact ttcaaaaagt taaacaatgt 420 ttgcaaacat gctcccccat gccacgttaa atgtccctca acactagcct gaaagttcca 480 aatgcctata ctgacgctgc agccatgcaa gaaagctgaa taaaatcaca ccacgggcag 540 It seems there is a tag number error in the original text where is shown as <000169?>. I've noted it as in the translation for consistency with the rest of the tags. If this was a specific error in the input, you may want to correct it in the original text.agagaggtgc agcaactggc aaagcacctg tgtgtgctgg ctgcacaaaa catgccgcac 600 agttccgcta agctaaatct gtccgtgtcc ccagactgca cccaactttt aagaaaggaa 660 agaaactgag gccaaaagaa ttaaagtcgt gtggggacaa gcctgaatta agagccagag 720 tggcagaaaa ccagagtgct gctgcaaaga agatgctaga gggagaaagg aaagacaggg 780 acgggggctg gggcttgggg agggccaagg ttggggtctc acagcggcgc ctctgagcgc 840 gactcgaaac ttcgaggcca gcagtctcca cccccggccc ttttcccact gtggcgagtg 900 ccgctgaaga gcctgccctt ggcccctcgc tctctcactc acctcgacat gagcctccac 960 atctcgcgct gccccatcgc caaacactcc gaagctaaag cagcagagcg agaatctccc 1020 ggaccgttcc agcgcctcgc gtgagccccg cccaccgccg tcctgcgccg ccgggcacag 1080 atgacaagtc ctccaggaag ccagagcgac cgtttccgct acgcgacggg gaagggcggg 1140 gcaagaacga cgcctggagg aaatagttga gggagaggaa gggatcgggg accgggccag 1200 ggagaaggcg gagagcgagc tgaggcggag cgggaagagg gaggatagaa cggactaggg 1260 cggattcgca ggaaggaga ttgcctcta gaggctgtct tagcccaag cgcaacctgt 1320 tgtctaggct tggcctgcca atttgaccag tgtgaagc ttttctga ttttagtaat ttttatta aacgtaaact tagtattctt tgagactca gtagaggaat 1440 tacagttgat aaggacatc gatgtagttt attcattcct aacccaatac agccgaatttc 1500 ctttaatatt aggattgaga cattttcatt tttctggga aaagttacag atctcccact 1560 cttcctcaaa gacacaag agttgtctga ttagacata atagatatgt cttcccaact 1620 aaagtaaag ttgctgttat cagtggaagc attackac tcagaagta gcttactct 1680 aaatgaaacc aacatttaac tcaacta ggatccactg ttaattgag gaaacgttc 1740 ctccgcttcc agatccacc agatgtgtag cagaagacat ttcatcgcca ctattaagta 1800 ttgacacatt tatcagaag aagaatccg gtgacaatt tcagagagga cacatgaag 1860 gagttattct tagtaccttg attattaaag agatacactg ttgagacttt gatcaaaat 1920 tacgttgcta acctaggtca aatgtaataa ccaatgccc actgagaga ggacagacaa 1980 2040. 2040. 2040. 2040. 2040. 2040. 2040. 2040 accatctgac atcaccatca tgcctcattc actgcgcaca gaaatgaacc attcatctga cctagcatgg tttactgga aaagcatggg ctttcctgtt agtcagactg agattcaaat cttccttgcc gccctagctt catgtgtgtg aaacctgagt ctcacaaggc ctcatcgtta 2220. gaagggccct gtgcttgttt taatgctgtg ctattaata tttttgaaca aggaaacaca aattttcatt ttacactagg cctgcaaatt atgtaatcag tcctacccct agcacactaa tagtttaatg accttgaaca aattatatag cttttcagaa cctccatttt ctcatctgaa aatggatgtg ctgatgatta gaggtgtttg tgaagatcca ttgaaatata gaaatatgtg cataattac ttaggacttc gtaggtatca agacctgtca gttatgccac accttccacc aagctgtcct tcctattcct aaattctcta tttcttctaa tgatgctatc gtcttctggc 2580 ttcccgctct agaaacctca gtcctttaaa gaaattatc ttaacattct attttgaaga ttttcaaaca tatagcaaaa ttataaaaat ttaaagcaaa cgcctatgtg ctcaccacat agattctacc tagatttaca ttttatcttg ctttattcct tatcccaaat accttatttt 2760 taaatgcatt tcaaagttaa ttacaagcat cgttttactt ctcccaaaat atggcatgtg 2820 tcttactacc tagagttaaa tatttgttta catattttcc cttttgagac aaaattttaca 2880 tacaacaaaa tgcacaaatc ttaagtgtac attcactgag tttgacaaat gtatacactt 2940 gtgtaaccaa aattccttag atataaacca ttgtcatcac cccaaaaagt ttcctctcac 3000 gatttcccag acattctctg cctaccactg ccttccccac cccaaccccc gccccagaag 3060 caaccactgt tctgtctcga gaaccactgt ttgttttgag acagggtctc actctatcac 3120 ccaggctaga gtacagtgcc acagtcatag ctgactgcag cttcaacctc cctgggctga 3180 agcaatcctc ccacctcagc ctcctgagta gatgggacta caggcacatg ccaccacacc 3240 cagctaattt ttgtattttt tttttttttg cgagatagag tcttgccatg ttactcaggc 3300 tggcctccaa ctcctgggct caaaagatct ccccgcctcg atctcccaaa gtgctaggat 3360 tacaggtgtt aaatactgca cccagtctga tttttctcta ccatcaatta gttttgcatc 3420 ttctgaaaat tcatataaaa tggaatcata cagtgtgcgc tttttgta tagcatttt 3480 cacttagcaa aatgtccata tcgtgtgtat cggtagtttc ctccatttta ttactgagta 3540 gtatttcatt atatgaatac cccatagttt actgatctgt gctcctcctg atagattcct 3600 gggctgttgc cagttggggc tattatgaat aaagctgctg taacattct tgtgtaaatc 3660 tattgtagac atatgtttcc tattctctg ggtaaatatc tataaataga atgtcttggt 3720 tatagggtac atcaatgttt ggttttgt gaacagcca gacctttt tcaagagttt 3780 gtatcatttt ggaaaagaat tctggtggca tacaatccac agaatgagga ataatgttta 3840 taatgatat atttgataa aggctatcta gatatgtaa agactatca caatttaat 3900 aaccaattt taaatggat gagatttg atatactat acagctctcc aaaaaagata 3960 tatgaatgat attactcat atgaaat gctcacatt attagccatc gggaatca 4020 catcaaaacc acatgactt cacacccacc atattttt aaaaaaacc tgttagtgag 4080 gagaatttgg aatcctcaga caccactggt gatgtacaac agtacacta tctggaaaa 4140 atgtctggca gttcctcaa tggctaaaca gagagaccat atgacccac aattccatac 4200 ctaggtattt acccaaaata atgaaaaca tagtcctcac aaaaacttgt acataatgtt 4260 tatggcagca taattgccaa gaagtagaaa taactcaaat gtccatcaac tgatgaatag 4320 gtaaaaaaa aagttatata tctgtacaat gaagttatta ttggcaagaa acagaatga 4380 agtatcgata catggtacaa catgagcaa catggaaaac actgtgctag atgaagaag 4440 ccagtcaca taggaccat attatgat gccatttaca tgagtgtcta gaatagcaa 4500 atgtgtacag atagaagta gacgagaggt tgcctaggt tggaggag gcaggagaat 4560 tggagtga cagctaagcg gcacaggtt ttctggtaga gtaatgaaaa tattctaaaa 4620 tgaatagtgg tgatggttgg actacactga atgaaaata ttctaaaatg gatagtgctg 4680 acagttggct gatcatata ctttatatgg gtgaattgtg tggtatgtaa attgtatctt 4740 aatttttaaa taaagaatttc tggttgctc atattctcag ccacattttg tattatcagg 4800 cttttaaatt ttagtcatgc tggtgtgtgt gtggtgttat ctccttgtgg ttttaatttt 4860 cattttcctg gtgtctaatg atgttgagca ctttttcata tgcttattat cttcttttaa 4920 ttttcttgcc catttttaaa ttggggtttt tgttgttgtt gttttgtttt ttgttttttt 4980 ttttagacaa agttttgctc t 5001 <210> 3 <211> 5001<A <212> DNA <213> Artificial sequence <220> <223> Artificial sequence description: Synthetic sequence <400> 3 agccatgcaa gaaagctgaa taaaatcaca ccacgggcag agagaggtgc agcaactggc 60 aaagcacctg tgtgtgctgg ctgcacaaaa catgccgcac agttccgcta agctaaatct 120 gtccgtgtcc ccagactgca cccaactttt aagaaaggaa agaaactgag gccaaaagaa 180 ttaaagtcgt gtggggacaa gcctgaatta agagccagag tggcagaaaa ccagagtgct 240 gctgcaaaga agatgctaga gggagaaagg aaagacaggg acgggggctg gggcttgggg 300 agggccaagg ttggggtctc acagcggcgc ctctgagcgc gactcgaaac ttcgaggcca 360 gcagtctcca cccccggccc ttttcccact gtggcgagtg ccgctgaaga gcctgccctt 420 ggccctcgc tctctcactc acctcgacat gagcctccac atctcgcgct gccccatcgc 480 caaacactcc gaagctaaag cagcagagcg agaatctccc ggaccgttcc agcgcctcgc 540 gtgagccccg cccaccgccg tctgcgccg ccgggcacag atgacaagtc ctccaggaag 600 ccagagcgac cgtttccgct acgcgacggg gaagggcggg gcaagaacga cgcctggagg 660 aatagttga gggagaggaa gggatcgggg accgggccag ggagaaggcg gagacgagc 720 tgaggcggag cgggaagagg gaggatagaa cggactaggg cggattcgca ggaaaggaga 780 ttgccctcta gaggctgtct tagcccaaag cgcaacctgt tgtctaggct tgcctgcca 840 atttgaccag cagggactga tgtgaaagac ttttctggaa ttttagtaat ttttattaa 900 aacgtaaact tagtattctt tgagactcaa gtagaggaat tacagttgat aaaggacatc 960 gatgtagttt attcattcct aacccaatac agccgaattc ctttaatatt aggattgaga 1020 cattttcatt tttcttggga aaagttacag atctcccact cttcctcaaa gacacacaag 1080 agttgtctga ttaagacata atagatatgt cttcccaact aaagtaaaag ttgctgttat 1140 cagtggaagc ataattacat tcagaaagta gcttactctt aaatgaaacc aacatttaac 1200 tcaactacta ggatccactg ttaaattgag gaaaacgttc ctccgcttcc agatccaccc 1260 agatgtgtag cagaagacat ttcatcgcca ctattaagta ttgacacatt tatcagaaag 1320 aagaaatccg gtgacaaatt tcagagagga cacatgaaag gagttattct tagtaccttg 1380 atatttaaag agatacactg ttgagacttt gatcaaaaat tacgttgcta acctaggtca 1440 aatgtaataa ccaaatgccc actgagaaga ggacagacaa cagacatcac agaaggaact 1500 ccttacctgc aatttggttt ttaactggct acaggacatc accatctgac atcaccatca 1560 tgcctcattc actgcgcaca gaaatgaacc attcatctga cctagcatgg tttactggaa 1620 aaagcatggg ctttcctgtt agtcagactg agattcaaat cttccttgcc gccctagctt 1680 catgtgtgtg aaacctgagt ctcacaaggc ctcatcgtta gaagggccct gtgcttgttt 1740 taatgctgtg ctattaataa tttttgaaca aggaaacaca aattttcatt ttacactagg 1800 cctgcaaatt atgtaatcag tcctacccct agcacactaa tagtttaatg accttgaaca 1860 aattatatag cttttcagaa cctccatttt ctcatctgaa aatggatgtg ctgatgatta 1920 gaggtgtttg tgaagatcca ttgaaatata gaaatatgtg cataaattac ttaggacttc 1980 gtaggtatca agacctgtca gttatgccac accttccacc aagctgtcct tcctattcct 2040 aaattctcta tttcttctaa tgatgctatc gtcttctggc ttcccgctct agaaacctca 2100 gtcctttaaa gaaaattatc ttaacattct attttgaaga ttttcaaaca tatagcaaaa 2160 ttataaaaaat ttaaagcaaa cgcctatgtg ctcaccacat agattctacc tagatttaca 2220 ttttatcttg ctttattcct tatcccaaat accttatttt taaatgcatt tcaaagttaa 2280 ttacaagcat cgttttactt ctcccaaaat atggcatgtg tcttactacc tagagttaaa 2340 tatttgttta catattttcc cttttgagac aaaattttaca tacaacaaaa tgcacaaatc 2400 ttaagtgtac attcactgag tttgacaaat gtatacactt gtgtaaccaa aattccttag 2460 atataaacca ttgtcatcac cccaaaaagt ttcctctcac gatttcccag acattctctg 2520 cctaccactg ccttccccac cccaaccccc gccccagaag caaccactgt tctgtctcga 2580 gaaccactgt ttgttttgag acagggtctc actctatcac ccaggctaga gtacagtgcc 2640 acagtcatag ctgactgcag cttcaacctc cctgggctga agcaatcctc ccacctcagc 2700 ctcctgagta gatgggacta caggcacatg ccaccacacc cagctaattt ttgtattttt 2760 ttttttttg cgagatagag tcttgccatg ttactcaggc tggcctccaa ctcctgggct 2820 caaaagatct ccccgcctcg atctcccaaa gtgctaggat tacaggtgtt aaatactgca 2880 cccagtctga ttttctcta ccatcaatta gttttgcatc ttctgaaaat tcatataaaa 2940 tggaatcata cagtgtgcgc ttttttgta tagcattttt cacttagcaa aatgtccata 3000 tcgtgtgtat cggtagtttc ctccatttta ttactgagta gtatttcatt atatgaatac 3060 cccatagttt actgatctgt gctcctcctg atagattcct gggctgttgc cagttggggc 3120 tattatgaat aaagctgctg taaacattct tgtgtaaatc tattgtagac atatgtttcc 3180 tattctcttg ggtaaatatc tataataga atgtcttggt tatagggtac atcaatgttt 3240 ggtttatgt gaaacagcca gaccctttt tcaagagttt gtatcatttt ggaaaagaat 3300 tctggtggca tacaatccac agaatgagga atatgttta taatgatat atttgataaa 3360 aggctatcta gatatgtaa agactatca catttaat aaccaattt taaatggat 3420 gagagatttg atatactat acagctctcc aaaaagata tatgaatgat caatactcat 3480 atgaaaagat gctcaacatt attagccatc gggaaataca catcaaacc acatgactt 3540 cacacccacc atatttttt aaaaaacaag tgttagtgag gagaatttgg aatcctcaga 3600 caccactggt gatgtacaac agtacagcta tcttggaaa atgtctggca gttcctcaa 3660 tggctaaaca gagagaccat atgacccac aattccatac ctaggtattt acccaaaata 3720 aaaaacttgt acataatgtt tatggcagca taattgccaa 3780 gagtagaa taactcaat gtccatcaac tgatgatag gtaaaaaa aagttatta3840 tctgtacaat gaagtattat ttggcaagaa acagaatga agtatcgata catggtacaa 3900 catgagcaaa catggaaaac actgtgctag atgaagaag ccagtcaca taggaccacat 3960 attatatgat gccatttaca tgagtgtcta gaataagcaa atgtgtacag atagaaagta 4020 gacgagaggt tgcctaaggt tggagaggag gcaggagaat tggagagtga cagctaagcg 4080 gcacagggtt ttctggtaga gtaatgaaaa tattctaaaa tgaatagtgg tgatggttgg 4140 actacactga aatgaaaata ttctaaaatg gatagtgctg acagttggct gaatcatata 4200 ctttatatgg gtgaattgtg tggtatgtaa attgtatctt aatttttaaa taaagaattc 4260 tggttgctcc atattctcag ccacattttg tattatcagg cttttaaatt ttagtcatgc 4320 tggtgtgtgt gtggtgttat ctccttgtgg ttttaatttt cattttcctg gtgtctaatg 4380 atgttgagca ctttttcata tgcttattat cttcttttaa ttttcttgcc catttttaaa 4440 ttggggtttt tgttgttgtt gttttgtttt ttgttttttt ttttagacaa agttttgctc 4500 ttgttttcca ggctgcagtg taatagcaca atctgggctc actgcaacct ccacttctgg 4560 ctaattttgt atttttagta gagatggggt ttctccatgt tggtcaggct ggtctcgacc 4620 tcccaatctc aggtgatcca cctgcctcgg cctcccaaag tgctgggatt ataggcgtga 4680 gtcaccgcac ctggccaaat tgggttatct ttttattcgt gtttgatatt gtcttttttt 4740 aattattgtt caattatcta tttctgttac aaattacccc taaaatttag cagcttaaag 4800 caaccaacat ttattgtttc tcagagattc caaggctcag gaatcaggga gagacttagc 4860 caggtggttc tggctcaggg tctctcatga ggctgcaaag cgttggctgg gattgcagtc 4920 atctgaatgc tcagcttcca aacttatgtg gttgttggca ggttacaggt cctttcacgt 4980 gggcctctcc atggggctgc t 5001 <210> 4 <211> 5001 <212> DNA <213> Artificial Sequence <220> <223> Artificial Sequence Description: Synthetic sequence <400> 4 ggaaggagga tcccgaatcc cagccagaac tgagaaagcc ctcaaaggtg cggaaagggg 60 acttttccct caggaaagcc ggcaacagca gaggccctag cccacgtcgc cacacccact 120 ccgcgcgcgc ccctcgcctc ccagggccgg ctctgctcgg ccccgcggcc tctcgggcgc 180 ccccagcccg ccggaaagaa agaagcagaa acccggagcc tgggtcccac ccgcgacccc 240 tcaccttcgc ccttctcgtc ctctcacctg ccagtccccc aggaagaaca ggacgcctct 300 tccccctgat gggcggccac aggctcggat ccttccattg ccgcctgaga caccgccgcc 360 ggctactgga aggtaggaag gggcgggacc gtggggggtt caggggcgg gcggagacgt 420 catcagagggg ggcgggcctg gggaagctag gagcccggca gcgccttccc gtcaacccta 480 ggggcgtcct gtttccggt tggtgtgg ccgcatggcg tgctgtggtg caggtggccg 540 aagggggcgt tactgttgcg actggcatcc gcatccggca gatgtagatg gaaccaagt 600 ccagaagtta cgcgtcaccc ttgcttaca gccaaacatg caggactcta gtaacccgcg 660 aaatgatggg atagcgttgc aaatccttaa agagtctta acggtaagaa gaggaacag 720 ctttatttta taaataact gaaggctgct atagctcca gattgtcagt gaggggatac 780 cacttaaga tgttatacat ttataata atttggg tccactacat gtgaggcgct 840 gcctcagatt taatagaatc atgtgacttt tagtgccata gatgtcgtta gatcatccag 900 tttggtccct taaatttgtg ggtaagggaa aggaacttcc ggtcagtgca agtgacttgg 960 tcaggttat agagctagtc actgttagat taccagaca catatctaaa aagttctcat 1020 ataattcttt ttcatttact tccgtgtcac ctttcttaga aagcaaaaat gagttggaca 1140. 1140. 1140. 1140. 1140. 1140. 1140. 1140. 1140. 1140. 1140. 1140 ggatttttaa aattagcata tactttggggt attgtaat tttctgtgat gtatcatatt ttttattaga aatttttcct ctcatatgga gatgagggtc caggcaaaac tgtattatcc 1320. ttcgtgccct ctattctctt caaattgca catcttaaaa tacctggaag ctttatgttg ctcacctata gcattatccg tgttcctcaa tccagttgtg aaccttttcc acattaaaat gttcttgatc tgtgttctat tctttcaagg ctctgctcta gcactaagat aatgctacag 1500. 1500. 1500. 1500. 1500. 1500. 1500. 1500. 1500 aaagtttctt tccttgtaaa ggttaaaaca gagcatactt ccttattatg cccactaggt 1560. 1620. ctcctgtttt tttaaaaaaa actggctcat tttttcaaaa tccactttaa ttacatggtg cttagcacaa gcaactccat tctggtttga tctgatgggg cctagtgtag 1680 aaggctaggc caaaacaatg gcctgccatt acccacgctt cttcatattt taacatccat gagagagga tgaattttta aaaaat acagaaacg atgcatctta aaattgaggg 1800 catgttagat ttgataaat atggtgtgaa tattttcagg catgtaaca gcaaccagat 1860 ttaacagatt ttaatgtttt gtcacatatg taaaatggga aagaaacta ggcgatcaca 1920 gctttttgac agtgaggctt tatgtcataa gaagtagaa tgacatattg aggaactta 1980 agggaaaaat atgtaggcca aagattttac attcagcaa actaatac aaaaggtaca 2040 aactgttatc aacaatgat atatcaattt ggtgagacta gtgttttcat gagcccctcc 2100 ttaatatta attggagaaa gtgtctgaca atcagaatta ttacagaac attgatctaa 2160 ggattgactc ggttcttta cagcaaatca taatctgttg cattattag cttttgtgtt 2220 tcttgtagtt tattctttat tactattct cctgagttct gtcacct tctgaggtt 2280 tttctaattc agaatcaagt tcttccatat ctttatcat ttctctaatg tcatttagct 2340 agtttcaaaa taataaatttg tttgtagtg catctttcag gcatgttttt attatctttg 2400 atgatgatgt tatattcctt attctcttca cataataact gcatggattt gaacttgatg 2460 ctttttttgt tgctcttttt tttttttttt ttttttttt tttgagacag agccttgctc 2520 tgttgcccag gttggagtgc agtggcggga tcttggctca ctgcaagttc tgcctcccgg 2580 attcacgcca ttctcctgcc tcagcctccc gagtagctgg aactacaggc acctgccacc 2640 acgccccggct aatttttgt atttttagta gagacggggt ttcaccgtgt tagccaggat 2700 ggtctcgatc tcctgacctc gtgatccacc ccttggcctc ccaaagtgct gggattacag 2760 gcgtgagtta ttgctcattt ttatgtgaaa ttggtgttcc tgaactaata tgagacaggc 2820 ttcttagccc aataagccta agtttatcag gaacaaaact gtattccaag aaagtcaggt 2880 atattagtct gttttcacac tgctataaag aaatacccga gactgggtaa tttataaagg 2940 aaagagattt agttgactca cagttccaca tggctggggga gtcctcagga aacttaaatc 3000 atagcagaag gtgaaggaga agcaagacat atcttacacg gtggcaggag agaagtgaga 3060 tcgtgggaaa actgccactt ttaaaaccat cagatctcat gagaactccc tcattatcac 3120 aagaatagca tgggggaact gcccccataa tctaatcacc tcccaccagg tctctccttg 3180 cacacatggg aattaaaatt tgaagtgaga ttgggtgggg gandacagagc agaccatatc 3240 atcaggttta ctggctcatt gcaggagggg agcccacatt ccagagaac catggggtac 3300 ctcaccaca gaggagaga taggttatt gtaggaattt gggaagtat ggaggtttaga 3360 aggaattgaa atgagtagt gctttgatag actcaaaga aagcaggct gtttgtaag 3420 aattctacac tacatctaaa ctattacta tggcaaaac tgcaattact tttgcaccaa 3480 cgtaatacaa gatggaccca gggttctgtt ttttggaat tgcaagttc tgcctctcag 3540 aagtagaaac tattctctg tgtcaatgtg actttagata ttctagccaa gagtgagata 3600 ttcaatctt aaatagga ttaacagc agagttttg acaatcttt gatattgtca 3660 taaatgttct atgattttag agaaaagggc attgtgtg taaaaaatca gttactctaa 3720 gggtgtgtgt gtgtgtgtgt gtgtgtgtgt gtgtgtgt gtgtttgaga cattttaaag 3780 catcttccag agagaatag tattttctgt taatttcaca gctggcttta tttagcctg 3840 ttatcccata gctgatgaat ggcaaagcag atttttactt tcttattcta tgcttatttg 3900 ctttctcaca ggttttggga aaacttttct aatttcagag aattcccttt tctgttgctt 3960 ttgtgtaatg ttcagcaata ttatggtttt atttcctttt taaattaata catacataat 4020 aagcatatat ttatggggta ttttgatatt tgcgtacaag gtataatgat aaaaccaagg 4080 taattgtgat atccatcacc tcaaacattg tttatttctt tgtgttagga acatttcaca 4140 tcttgccttc tagctatttt gaaatatatg ataaattatt gttaactata gtcaccctct 4200 agtgctatcg aacactagaa attgtttcct ctatctaatt gtatttttgt acccattacg 4260 caaccttttt catcctcctt acccccaacc ttcccagcct ctggtaacca tcattctact 4320 ctctactgcc atgagatcca cttttttagc tctcacttat gagtgagaat atgtgctatt 4380 tgtctttcta tgcctggcat ttcacttaac ataatgacct ccagttccat ctatggtgct 4440 tcaaactaca tgatttcatg tttttgtgta gatgacacat ttttaaaatc cattcatctg 4500 ttgatggaca cttagattga ttccatatct tggctattgt gaatagtact gtaataaaca 4560 tgggcatgca gatacctttt aaatttattg aattccgttt ttggatgtat actcaggaat 4620 gggattgttg gatcatgtgg tagatctatt tttagatttt tgagaaactt ccatactgtt 4680 tttcatagtt tctgtactaa tttacgttcc caccagcaat gcactagtgt tctccttttg 4740 ccacatcctc accagcatct gttattttct gcctttttgg taatagtaat cttaactgag 4800 gtgagatgat gtcttattgt ggttttgatt tgcatttccc tggtgactgg tgatgttgag 4860 cattttttca tatgtctgtt ggccatttat atgtcttctt ttgataaatg cctgtttaga 4920 tcatttgccc attttaaaat tggattattt gggggttttt gctatggagt tgtttgagtt 4980 ccttatatat tctgattatt a 5001 <210> 5 <211> 5001 <212> DNA <213> Artificial Sequence <220> <223> Description of artificial sequence: Artificially synthesized sequence <400> 5 taacctctct gttcttttcc gttcttatct ataaataggg gcaaccaacc accagtcctc 60 agaggctctg aaaatctaat ttcaattgtg gttgaaagtg ctctgaaaat ttgaaatgtt 120 acacagatgt atgggattcc tcctgagctg tatttgaaac cacagatatc catctgattt 180 gtaaatcatt taccttttag tctccctagc agtagttaat agtgaatgct tttctttttt ttctttttat tttcttgatt gctaattttg ttagttcat aagcgatact tacatatgta catgtggagg ctggaatctg ctctggttga tgttgtagtg cgtgaatgcc tgggtggggt 360 ttgagctgta ctcatccacc gttatggtaa ctttgccttg agaaggaatc ccatgtgcca tccttcctcc ctataggcat gagcaattca caactttaca gagcaaaaga tggaagcagg catggtgcgt ttcagccaac ctctgaatcc aaggatctac tcatcaatac cgctgtgcat ttctgaaaga aaaaaaaaat gtatataaaa tgagaacaaa gagaaaagg gtgttctaag 660. the sword of the sword of the sword of the sword of the sword of the sword of the sword of the sword of the sword of the sword of the sword of the sword of the sword of the sword of the sword of the sword tggctcatgg atacaactgt tcccaactct cttcccttca tgcacttcaa agtggatatt acactctata catattgctg ggtgctagca aatcttctca atggagttca aactaagatt tcctctttaa tttaattat ggcacttagc gtgaattacg gaaggggcat atccaccgtt gccatggcat ctggctatac atggtttgtg ctccaagtga agaatcttcc aaaacctgccc 900. tgctaattgg aatgtttggt cttacaaag ttggtgtcac gtaatcattt gaaagctcca 960 tgtttaaagc agctatataa cacaatgaca gataagaatt taaaatgttg tgcttataac 1020 aaatacatgt gtattactat aacatgcaca tataaccac acattacaca tatttttttc 1080 tttcacaagc agctcccaaa atacacaggt aggcaatgct tggcattcct cactggccca 1140 gaaccttatg gtaaggcaaa tctttcttgt taaatgtagt gtcatttttt tcaacactcc 1200 agtcactcag gggaaaacca aacacctcca ggtctggacc accactttga atgcattcat 1260 tacctccctt cttcaagccc cgatggagcg gtaaagaagg gcggccatta catctgccac 1320 ctcattcttg gtgcaacgta attctcagta attggtttca ttttttaaaa tttccattag 1380 ttctgattt ttcataaatg tcaatatatt atttgctgca ctgggttgag aaatgaagat 1440 gtctattttt atggtaggtg tgcaaaggct gagggaagca aattctcatg tccaagttgt 1500 tgcaattgat cagtgcggta gagaagaagc cgtagtacac tccttgggta acccttaggt 1560 aacccatgct cttactttaa gtgcacaacg tcagaatagg tttgccatgt gcaccattaa 1620 gctctgagaa tcttatctgc tatttcaaat aaaaatattt agaataataa gatgagattt 1680 aaaattggca tgaaacacct taaatatcga gcaaaaggca gatggcctgc cataggccat 1740 ttttttatt attttacaac ccggggcttg agtaatcccc catatgctag ctgggctaag 1800 aaattccaga gtgttctact caaatcattg ccttcttcat attagcaaag caagctaaat 1860 aacacccacc ttgcctgcct tcctaatgtc aagtgcactg ccaagagatg aatccccgtg 1920 ctctggcatt ttaaacaaag tgtcactttg ggttctcttt cttttcctg acttcatggg 1980 aggtatgagt cttttaagcg ccagatgata gacaaccagg ccagtaggat taactaaatt 2040 aaaatctttc cctgctcctc tgatgtaaga tatctataag aactgcttac aaatcataat 2100 gtagcagttt tcataccaag tttgcagctg tttccttaat ttctcccaca attatatttt 2160 actgttattt tgcataacag taaaacacaa cacctgcttc aaattgcagt cctgttctaa 2220 gtactgtaat tctcttttga gacacaaagt accgagacac tactgaagcg caaaactgca 2280 taattgcgtc gcgctgaaaa gtaaatacaa atgaaaacat cttcctgtta tagaaatgac 2340 tgtcattaaa aagaatgtta gccatcaaag cacagagggg gaaaattca cattttcaga 2400 ctgtgagaag ttagtgtttg cataaattag agagtggtgt aaaaaattact tgacaattct 2460 tatgttcatt cttccatcaa tgtctattgt gtgaaatact gaacattaag aagcaaagga 2520 aatcaaattc atgtctagac cgctgtgacc tctgagtaac aatctctgca cccagtacta 2580 gaaagatttt aaatgtttct ggagtaaaca cactaattaa accaagcaaa tgtggcctag 2640 aaataaaata tacataccta atatttaaac atctaggtaa tgacttgtag cagttagcat 2700 gatacccacc caagttctac tatttgaagc taagaagaga cagaatttag cagctgatca 2760 atttgcactt ctacatggtc cgcttgatcc taatgtcaac cttctctatc actgaacatc 2820 tatgcctctt tgcacctttg ccccagtaag acaaatagaa catcttttca tcagcaagca 2880 caggcccctg taactctgaa aggctaaatg ggtggctggc acctcatttg aattcatctt 2940 tgccatgttc cctttcaaat gtttcatcga aacagtgcac ctgcagcgtc tatcttcttt 3000 aatctggagt cctttgaaga aaatgtcaga ctgaggcaga gagcttcaaa actttcacag 3060 cacctttaca gandaagtca gaaaaggctt atttttaaa atgctatgtt catccccacc 3120 cgtatctcaa atgacaca tcttaagaat tttaacaca aacacctgga aagactttt 3180 tctttttttttttttttttttttttttttggtaac tctcagagac aagggacaat 3240 agacaaaacc acttcctata acctgagaaa tttggagaga gcttgaaca agacagaagg 3300 atacactgct tacaactatc aagagatgt atacactgc agtaaaacct tccaaatgat 3360 cagattttcg tgcatctcac cctctcaca cttactttaa ctttcac acaatattta 3420 aagacagcac tattctttc aaattttc attttatact tcacacact ttccacaaaa 3480 aaggaacc spy spy tatcaagg spraaggaaaatctaaag gggactgaagg 3540 tgccaaacat tcatgctgt gctagtttga tttctct taaaaagaaa atcatgtgt 3600 tacattttgc atttcctaga cagcaatcac ttctgtcctt taagctgtac aactagcaat 3660 attctttttcc aggaatctag ttggaaacc ttttaaagtc ggctggcaat gaagggatta 3720 aaagctgagg aattaaataa ctgtctccac ctcctagaca ccagaagtcc aatttacca 3780 aactttcaaa gcaacaactc atgacctctc tacccctgca ctgcccagca cagcctatga 3840 gaaaagagcc cgcggttttc tcgtttcaac aggatgctc attttaatca cttgccccaa 3900 gatctccaga ccgagcctgg cgggactttg gaacttgcct gctccaaatc cgccaaattt 3960 cgctctctcc tccctctgcc gcgcagatcc agtgcgttct taattcacgc ccctactctt 4020 ttatgcccac ttccagttac tcatccctgt atctgggagc gatgcccaca cgcttcccg 4080 gcgggcttct ccgccagggt ttttttgggg gaggctgaca taaagttcca cgtgcttctg 4140 tccaccgagc gccctccctc ccaccgcact tacacctgaa cttgtctcca gcactgcgga 4200 cacccgggtg acacattctt tcggagaggg agggcaagac ttttctccgg agcctctggc 4260 aactcgtgga gtctatataa ggtggggcgg tggggggcgc ggaattctct actttggatt 4320 ggtttgtttg ttttaaatct agggtggtgg gcagttgcgc gaggcggggt ggggggtgct 4380 ttatgaacta cacacaca cacacaca cacacaca cacacacat acactcgccc 4440 tgcccttcta ggctcacggc tggctcttcc cggctacaac ttccccttcc ttcctgaaac 4500 accagcgcct tactcgctca gtcccccgtt gccacctccg tcacctcctc ggccgccgct 4560 cctgctggca ctgctggcgc tagcgctggc ggccccgggc tccgctggac gcctatcttt 4620 tttctctccc tcttccctag ctctggattg attgatgacc accggcggcc aatggctctg 4680 acaccacgcc atcctcggtc gggcatcaca gccccacgtc cccgacctca gcctcccggg 4740 cgcgttggca tccccgaacc cacagcgctg agcccgagcc cggggccgcc gcctcctccc 4800 tgctctacgc ctgggcgctg ccaggttcct tccctggaac ggcggagaaa ggaagggagg 4860 gagccagcga gcgttggaag ggggtgtgga ggccacggga tctccccctc ttctgacatc 4920 ttttttcact taaacttttg tttgtcccac ccccccgacc tgtttgtgtt tgtttgcagc 4980 atctcaaatc aattatacat t 5001
Claims
1. a first marker that is identical to or reverse complementary to the nucleotide sequence at positions 151445000-151450000 on chromosome 1; or A second marker that is identical to or reverse complementary to the nucleotide sequence at positions 191183500-191188500 of chromosome 2; or a third marker that is identical to or reverse complementary to the nucleotide sequence located at positions 191184000-191189000 of chromosome 2; or a fourth marker that is identical to or reverse complementary to the nucleotide sequence located at positions 68566500-68571500 of chromosome 4; or a fifth marker that is identical to or reverse complementary to the nucleotide sequence located at positions 30601500-30606500 on chromosome 11, Use in preparing a kit for detecting lung cancer.
2. A system for detecting lung cancer, characterized in that: The system comprises: A sample collection module, which is used to collect samples from subjects; A sample processing module is used to extract cfDNA from the sample and perform methylation treatment for sequencing; A sequencing module, which sequences the samples processed by the sample processing module; an analysis module for analyzing the sequencing results to determine the comprehensive methylation levels of the methylation markers in the sample; a determination module, which is used to determine whether the subject has lung cancer based on the comprehensive methylation level obtained by analysis; The markers include one or more of the following methylation markers: The first marker is identical to or reverse complementary to the nucleotide sequence at position 151445000-151450000 of chromosome 1, The second marker is identical to or reverse complementary to the nucleotide sequence at positions 191183500-191188500 on chromosome 2, The third marker is identical to or reverse complementary to the nucleotide sequence located at position 191184000-191189000 on chromosome 2, a fourth marker that is identical to or reverse complementary to the nucleotide sequence located at positions 68566500-68571500 on chromosome 4, and The fifth marker is identical to or reverse complementary to the nucleotide sequence located at positions 30601500-30606500 of chromosome 11.
3. The system according to claim 2, characterized in that In the analysis module, the comprehensive methylation level of the methylation marker refers to the comprehensive methylation level calculated based on the methylation levels of one or more selected methylation markers.
4. The system according to claim 3, characterized in that In the analysis module, the methylation level of each methylation marker is calculated based on the methylation level of each CG site in the methylation marker, wherein the methylation level of the CG site is the ratio of the cytosine detected to be methylated at the site to the sum of the cytosine detected to be methylated and the cytosine not methylated in all sequence results detected for the site.
5. The system according to claim 4, characterized in that In the determination module, the comprehensive methylation level refers to the methylation level determined based on one, two, three, four or five of the first marker, the second marker, the third marker, the fourth marker or the fifth marker. wherein the threshold is determined based on methylation level detection data of one, two, three, four, or five of the first marker, the second marker, the third marker, the fourth marker, or the fifth marker in a given lung cancer subject or a population of healthy subjects; If the comprehensive methylation marker is greater than a specified threshold, the subject is determined to have lung cancer, and if the comprehensive methylation level marker is less than or equal to the specified threshold, the subject is determined to be normal.
6. The system according to claim 2, wherein: The subject sample is a blood sample.
7. The system according to claim 2, wherein: The methylation treatment is carried out using bisulfite.
Citation Information
Patent Citations
Detection of lung neoplasia by analysis of methylated DNA
CN109563546A
Diagnostic and prognostic methods for cancer
WO2009105154A2