Methods for detecting Parkinson's disease

By analyzing the expression levels of specific genes in skin surface lipids, the method provides an early and accurate non-invasive detection of Parkinson's disease, addressing the limitations of current diagnostic methods.

JP7849812B2Active Publication Date: 2026-04-22KAO CORP +1
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
KAO CORP
Filing Date
2021-05-14
Publication Date
2026-04-22

AI Technical Summary

Technical Problem

Current methods for detecting Parkinson's disease are limited in their ability to diagnose the condition early and accurately, often relying on symptoms that appear in later stages, and there is a lack of non-invasive biomarkers for early detection.

Method used

The method involves analyzing the expression levels of specific genes, such as SNORA16A, SNORA24, SNORA50, and REXO1L2P, in skin surface lipids (SSL) to detect Parkinson's disease through RNA sequencing and using these genes as markers for early detection.

Benefits of technology

This approach allows for early, accurate, and non-invasive detection of Parkinson's disease with high sensitivity and specificity, utilizing novel biomarkers that have not been previously associated with the condition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007849812000074
    Figure 0007849812000074
  • Figure 0007849812000001
    Figure 0007849812000001
  • Figure 0007849812000002
    Figure 0007849812000002
Patent Text Reader

Abstract

To provide a marker gene for detecting Parkinson disease, and a method for detecting Parkinson disease using the marker gene.SOLUTION: A method for detecting Parkinson disease in a subject including a process of measuring at least one gene selected from four types of gene groups consisting of SNORA16A, SNORA24, SNORA50 and REXO1L2P or an expression level of its expression product regarding a biological sample collected from the subject.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for detecting Parkinson's disease using Parkinson's disease markers.

Background Art

[0002] Pathologically, Parkinson's disease is a progressive neurodegenerative disease mainly characterized by the formation of Lewy bodies mainly composed of α-synuclein aggregates and the degeneration and cell death of dopaminergic neurons in the substantia nigra of the midbrain. Clinically, it is a disease mainly characterized by movement disorders such as muscle rigidity, tremors, akinesia, and gait disturbances. Parkinson's disease is the second most common neurodegenerative disease after Alzheimer's disease, with an incidence rate of 120-130 per 100,000 people. In Japan, it is estimated that there are about 140,000 patients.

[0003] Currently, there is no radical cure for Parkinson's disease, and symptomatic treatment such as supplementing L-DOPA is considered important for maintaining QOL by controlling symptoms. However, the subjective symptoms of movement disorders appear in the mid to late stages, and it is required to diagnose the disease early and intervene early.

[0004] As biomarkers for detecting Parkinson's disease, in addition to those that detect α-synuclein accumulation, detecting circulating serum-derived microRNA (Patent Literature 1), measuring the concentration ratio of tyrosine to phenylalanine in blood (Patent Literature 2), etc. have been proposed. In addition, in the skin of Parkinson's disease patients, the formation of α-synuclein aggregates is observed as in the brain (Non-Patent Literature 1), and it has been reported that skin diseases and symptoms such as seborrheic dermatitis, melanoma, bullous pemphigoid, and rosacea appear in Parkinson's disease patients (Non-Patent Literature 2). It is considered that there is some relationship between Parkinson's disease and the skin condition, but the scientific relevance is completely unknown.

[0005] Meanwhile, technologies are being developed to investigate the current and future physiological state of the human body by analyzing nucleic acids such as DNA and RNA in biological samples. Analysis using nucleic acids has advantages such as the establishment of comprehensive analytical methods that allow for obtaining abundant information in a single analysis, and the ease of functionally linking the analytical results based on numerous research reports on single nucleotide polymorphisms and RNA function. Nucleic acids derived from living organisms can be extracted from bodily fluids such as blood, secretions, and tissues, but recently it has been reported that RNA contained in skin surface lipids (SSL) can be used as a sample for biological analysis, and that marker genes of the epidermis, sweat glands, hair follicles, and sebaceous glands can be detected from SSL (Patent Document 3). [Prior art documents] [Patent Documents]

[0006] [Patent Document 1] Special Publication No. 2019-506183 [Patent Document 2] Japanese Patent Publication No. 2016-75644 [Patent Document 3] International Public Gazette No. 2018 / 008319 [Non-patent literature]

[0007] [Non-Patent Document 1] Rodriguez-Leyva I et al. Ann Clin Transl Neurol. 2014 (modified) [Non-Patent Document 2] Ravn AH et al. Clin Cosmet Investig Dermatol. 2017 [Overview of the Initiative] [Problems that the invention aims to solve]

[0008] The present invention relates to providing a marker for detecting Parkinson's disease and a method for detecting Parkinson's disease using the marker. [Means for solving the problem]

[0009] The inventors collected SSL from the skin of Parkinson's disease patients and healthy individuals, and comprehensively analyzed the RNA expression status contained in SSL as sequencing information. As a result, they found that the expression levels of certain genes differed significantly between the two groups, and that this could be used as an indicator to detect Parkinson's disease.

[0010] In other words, the present invention relates to the following 1) to 3). 1) A method for detecting Parkinson's disease in a subject, comprising the step of measuring the expression level of at least one gene or its expression product selected from a group of four genes consisting of SNORA16A, SNORA24, SNORA50, and REXO1L2P in a biological sample taken from the subject. 2) A test kit for detecting Parkinson's disease used in the method of 1), comprising an oligonucleotide that specifically hybridizes with the gene, or an antibody that recognizes the expression product of the gene. 3) A detection marker for Parkinson's disease, comprising at least one gene or its expression product selected from the gene groups shown in Tables 3-1 to 3-4 and Tables 6-1 to 6-2. [Effects of the Invention]

[0011] According to the present invention, it is possible to detect Parkinson's disease early, with high accuracy, sensitivity, and specificity, in a simple and non-invasive manner. [Brief explanation of the drawing]

[0012] [Figure 1] A confusion matrix plotting the predicted and actual values ​​for the optimal prediction model on the test data. [Modes for carrying out the invention]

[0013] All patent documents, non-patent documents, and other publications cited in this specification are hereby incorporated by reference in their entirety.

[0014] In the present invention, the term "nucleic acid" or "polynucleotide" means DNA or RNA. DNA includes cDNA, genomic DNA, and synthetic DNA, and "RNA" includes total RNA, mRNA, rRNA, tRNA, non-coding RNA, and synthetic RNA.

[0015] In the present invention, the "gene" includes double-stranded DNA containing human genomic DNA, single-stranded DNA (sense strand) containing cDNA, single-stranded DNA (complementary strand) having a sequence complementary to the sense strand, and fragments thereof, which means that some biological information is contained in the sequence information of the bases constituting the DNA. In addition, the "gene" includes not only the "gene" represented by a specific base sequence, but also nucleic acids encoding homologs (i.e., homologs or orthologs) thereof, mutants such as gene polymorphisms, and derivatives. The names of the genes disclosed in this specification follow the Official Symbol described in NCBI ([www.ncbi.nlm.nih.gov / ]). On the other hand, regarding Gene Ontology (GO), it follows the Pathway ID. described in String ([string-db.org / ]).

[0016] In the present invention, the "expression product" of a gene is a concept that includes the transcriptional product and translational product of the gene. The "transcriptional product" is RNA generated by transcription from a gene (DNA), and the "translational product" means a protein encoded by the gene that is translationally synthesized based on RNA.

[0017] In the present invention, "Parkinson's disease" means a sporadic, progressive disease with the main lesion being the degeneration of dopaminergic neurons in the substantia nigra pars compacta, and expressing three motor symptoms (rest tremor, rigidity, bradykinesia, and akinesia) in a slowly progressive manner.

[0018] In the present invention, the "detection" of Parkinson's disease means to clarify the presence or absence of Parkinson's disease, and can also be paraphrased by terms such as examination, measurement, determination, evaluation, or evaluation support. In addition, the terms "determination" or "evaluation" in this specification do not include determination or evaluation by a doctor.

[0019] The four genes consisting of SNORA16A, SNORA24, SNORA50, and REXO1L2P of the present invention are genes selected from the 33 genes described in Table A below, in which it has been found that the expression level of SSL-derived RNA is significantly increased (UP) or decreased (DOWN) in Parkinson's disease patients compared to healthy subjects as shown in the examples described later, and genes that have not been previously known to be associated with Parkinson's disease (shown in bold in the table).

[0020]

Table A

[0021] The 33 genes shown in Table A are from data on the expression levels (read count values) of RNA extracted from SSL of subjects in two tests (Test 1: 15 healthy subjects / Parkinson's disease patients each, Test 2: 50 healthy subjects / Parkinson's disease patients each), which were converted to RPM values corrected for differences in the total number of reads between samples, and based on the values obtained by converting the RPM values to the logarithm base 2 (Log2RPM values), RNAs with a p-value of 0.05 or less in Parkinson's disease patients compared to healthy subjects were identified by Student's t-test (Test 1: 111 up-regulated genes, 68 down-regulated genes (total 179 genes, Tables 1-1 to 1-5), Test 2: 565 up-regulated genes, 294 down-regulated genes (total 859 genes, Tables 1-6 to 1-27)). From these, genes that were commonly up-regulated (18 genes) and down-regulated (15 genes) in Test 1 and Test 2 were selected. ​​Therefore, genes or expression products selected from the group consisting of 179 genes and 859 genes (a total of 1005 genes after removing duplicates) can serve as Parkinson's disease markers for detecting Parkinson's disease, and among these, genes or expression products selected from the group consisting of 33 genes shown in Table A are preferred Parkinson's disease markers. In Table A and Table 1 below, "p-value" refers to the probability in statistical testing that a statistic more extreme than the statistic actually calculated from the data under the null hypothesis is observed. Therefore, the smaller the "p-value," the more significant the difference between the comparison groups can be considered. Genes indicated by "UP" are those whose expression levels increase in Parkinson's disease patients, while genes indicated by "DOWN" are those whose expression levels decrease in Parkinson's disease patients.

[0022] The group of genes whose expression was altered as described above included those associated with Parkinson's disease (hsa05012) through the exploration of the biological process (BP) and KEGG pathway using gene ontology (GO) enrichment analysis (see Table 2 below). On the other hand, among the group of genes whose expression was altered as described above, the genes shown in Tables 3-1 to 3-4 below are genes that have not been reported to be associated with Parkinson's disease at all to date. Therefore, at least one gene or its expression product selected from these gene groups is a novel Parkinson's disease marker for detecting Parkinson's disease, and at least one gene or its expression product selected from the group consisting of SNORA16A, SNORA24, SNORA50, and REXO1L2P, which are common to both Test 1 and Test 2, is particularly preferred as a novel Parkinson's disease marker. More preferably, two or more genes selected from this group are chosen, even more preferably three or more genes are chosen, and even more preferably all four genes are chosen. It is also preferable to include at least SNORA24, which is common to Table A above and Table B below.

[0023] Furthermore, the identification of differentially expressed RNA can also be performed using RNA expression level data (read count values), specifically using normalized count values ​​corrected with DESeq2 (Love MI et al. Genome Biol. 2014), or the base-2 logarithm (Log2(count+1) value) obtained by adding an integer 1. For example, using normalized count values ​​as the RNA expression levels extracted from the SSL of subjects in the two trials described above, and identifying RNAs in Parkinson's disease patients with a corrected p-value (FDR) of 0.25 or less in the likelihood ratio test compared to healthy individuals, we obtained 74 up-expression genes and 209 down-expression genes, totaling 283 genes (Tables 4-1 to 4-8) in Trial 1, and 151 up-expression genes and 308 down-expression genes, totaling 459 genes (Tables 4-9 to 4-20) in Trial 2. In both Test 1 and Test 2, seven genes were elevated (ANXA1, AQP3, EMP1, KRT16, POLR2L, SERPINB4, SNORA24), and ten genes were decreased (ATP6V0C, BHLHE40, CCL3, CCNI, CXCR4, EGR2, GABARAPL1, RHOA, RNASEK, SERINC1), for a total of 17 genes (Table B). Therefore, genes or expression products selected from the group consisting of 283 genes and 459 genes (a total of 725 genes after removing duplicates) can serve as Parkinson's disease markers for detecting Parkinson's disease. Among these, genes or expression products selected from the group consisting of 17 genes shown in Table B are preferred Parkinson's disease markers, and genes or expression products selected from the group consisting of 11 genes shown in Table C, which are common to the genes shown in Table A above, are more preferred Parkinson's disease markers.

[0024] Furthermore, among the genes whose expression was altered as described above, the genes shown in Tables 6-1 to 6-2 below are genes that have not been reported to be associated with Parkinson's disease at all to date. Therefore, at least one gene selected from this group of genes or its expression product is a novel Parkinson's disease marker for detecting Parkinson's disease, and in particular, SNORA24 (shown in bold in the table), which is common to both Test 1 and Test 2, or its expression product is preferred as a novel Parkinson's disease marker.

[0025] [Table B]

[0026] Furthermore, the genes that can serve as Parkinson's disease markers (hereinafter also referred to as "target genes") include genes that have a base sequence substantially identical to the base sequence of the DNA constituting the said gene, insofar as they can serve as biomarkers for detecting Parkinson's disease. Here, substantially identical base sequences mean, for example, that when searching using the homology calculation algorithm NCBI BLAST with the conditions expected value = 10; gap allowed; filtering = ON; match score = 1; mismatch score = -3, the base sequence is 90% or more, preferably 95% or more, more preferably 98% or more, and even more preferably 99% or more identical to the base sequence of the DNA constituting the said gene.

[0027] The present invention provides a method for detecting Parkinson's disease, which includes measuring the expression level of a target gene, or, in one embodiment, at least one gene selected from the group consisting of SNORA16A, SNORA24, SNORA50, and REXO1L2P, or its expression product, in a biological sample taken from a subject.

[0028] In the Parkinson's disease detection method of the present invention, the subjects from whom biological samples are collected include mammals, including humans and non-human mammals, and are preferably humans. When the subject is human, their sex, age, and race are not particularly limited and may include infants to the elderly. Preferably, the subject is a person who needs or desires detection of Parkinson's disease. For example, the subject is a person suspected of developing Parkinson's disease or a person with a genetic predisposition to Parkinson's disease.

[0029] The biological samples used in the present invention may be any tissues and biomaterials in which the expression of the gene of the present invention changes with the onset and progression of Parkinson's disease. Specifically, examples include organs, skin, blood, urine, saliva, sweat, stratum corneum, surface lipids (SSL), tissue exudates and other bodily fluids, serum and plasma prepared from blood, and others such as feces and hair. Preferably, these are skin, stratum corneum, or surface lipids (SSL), and more preferably, surface lipids (SSL). The site of skin from which SSL is collected is not particularly limited and may include skin from any part of the body such as the head, face, neck, trunk, hands and feet. Areas with high sebum secretion, such as the skin of the head or face, are preferred, and facial skin is more preferred.

[0030] Here, "superficial lipids (SSL)" refers to the lipid-soluble fraction present on the surface of the skin, and is sometimes called sebum. Generally, SSL mainly consists of secretions from exocrine glands such as sebaceous glands in the skin, and exists on the skin surface in the form of a thin layer covering the skin surface. SSL contains RNA expressed in skin cells (see Patent Document 3 above). Furthermore, in this specification, unless otherwise specified, "skin" is a general term for the region including the stratum corneum, epidermis, dermis, hair follicles, and tissues such as sweat glands, sebaceous glands, and other glands.

[0031] Any means used for the collection or removal of SSL from the skin can be employed to collect SSL from the subject's skin. Preferably, an SSL absorbent material, an SSL adhesive material, or an instrument for scraping off SSL from the skin, as described later, can be used. The SSL absorbent material or SSL adhesive material is not particularly limited as long as it is a material that has an affinity for SSL, and examples include polypropylene and pulp. More detailed examples of procedures for collecting SSL from the skin include methods of absorbing SSL onto a sheet material such as oil-blotting paper or oil-blotting film, methods of adhering SSL to a glass plate or tape, and methods of scraping off and collecting SSL with a spatula, scraper, etc. To improve the adsorption of SSL, an SSL absorbent material containing a highly lipid-soluble solvent beforehand may be used. On the other hand, since the adsorption of SSL is inhibited if the SSL absorbent material contains a highly water-soluble solvent or water, it is preferable that the content of highly water-soluble solvents or water is low. It is preferable to use the SSL absorbent material in a dry state. The skin from which SSL is collected is not particularly limited and can be any part of the body, such as the head, face, neck, trunk, hands, or feet. Areas with high sebum secretion, such as the skin of the face, are preferred.

[0032] RNA-containing SSLs collected from subjects may be stored for a certain period of time. To minimize the degradation of the contained RNA, it is preferable to store the collected SSLs under low-temperature conditions as quickly as possible after collection. The storage temperature conditions for the RNA-containing SSLs in this invention may be 0°C or lower, preferably -20±20°C to -80±20°C, more preferably -20±10°C to -80±10°C, even more preferably -20±20°C to -40±20°C, even more preferably -20±10°C to -40±10°C, even more preferably -20±10°C, and even more preferably -20±5°C. The storage period for the RNA-containing SSLs under these low-temperature conditions is not particularly limited, but is preferably 12 months or less, for example, 6 hours to 12 months, more preferably 6 months or less, for example, 1 day to 6 months, and even more preferably 3 months or less, for example, 3 days to 3 months.

[0033] In the present invention, the objects to be measured for the expression level of a target gene or its expression product include cDNA artificially synthesized from RNA, the DNA encoding that RNA, the protein encoded by that RNA, molecules that interact with that protein, molecules that interact with that RNA, or molecules that interact with that DNA. Here, molecules that interact with RNA, DNA, or proteins include DNA, RNA, proteins, polysaccharides, oligosaccharides, monosaccharides, lipids, fatty acids, and their phosphorylated, alkylated, and glycosidic compounds, as well as complexes of any of the above. Furthermore, the expression level comprehensively refers to the amount of expression or activity of the gene or expression product in question.

[0034] In the method of the present invention, in a preferred embodiment, SSL is used as the biological sample. In this case, the expression level of RNA contained in the SSL is analyzed. Specifically, the RNA is converted to cDNA by reverse transcription, and then the cDNA or its amplified product is measured. For RNA extraction from SSL, methods commonly used for RNA extraction or purification from biological samples can be employed, such as the phenol / chloroform method, the AGPC (acid guanidinium thiocyanate-phenol-chloroform extraction) method, or methods using columns such as TRIzol®, RNeasy®, or QIAzol®, or methods using special silica-coated magnetic particles, methods using Solid Phase Reversible Immobilization magnetic particles, or extraction using commercially available RNA extraction reagents such as ISOGEN.

[0035] For the reverse transcription, primers targeting a specific RNA to be analyzed may be used, but for more comprehensive nucleic acid preservation and analysis, random primers are preferable. A general reverse transcriptase or reverse transcription reagent kit can be used for the reverse transcription. Preferably, a reverse transcriptase or reverse transcription reagent kit with high accuracy and efficiency is used, such as M-MLV Reverse Transcriptase and its variants, or commercially available reverse transcriptase or reverse transcription reagent kits, such as the PrimeScript® Reverse Transcriptase series (Takara Bio Inc.) and the SuperScript® Reverse Transcriptase series (Thermo Scientific Inc.). SuperScript® III Reverse Transcriptase and SuperScript® VILO cDNA Synthesis kit (both from Thermo Scientific Inc.) are preferably used. In the reverse transcription extension reaction, it is preferable to adjust the temperature to preferably 42°C ± 1°C, more preferably 42°C ± 0.5°C, and even more preferably 42°C ± 0.25°C, while adjusting the reaction time to preferably 60 minutes or more, more preferably 80 to 120 minutes.

[0036] Methods for measuring expression levels can be selected from nucleic acid amplification methods such as PCR, real-time RT-PCR, multiplex PCR, SmartAmp, and LAMP, which use DNA that hybridizes to RNA, cDNA, or DNA as primers; hybridization methods (DNA chips, DNA microarrays, dot blot hybridization, slot blot hybridization, Northern blot hybridization, etc.) which use nucleic acids that hybridize to these as probes; methods for determining the base sequence (sequencing); or methods combining these.

[0037] In PCR, a specific DNA to be analyzed may be amplified using a primer pair that targets that specific DNA, or multiple DNAs may be amplified using multiple primer pairs. Preferably, the PCR is multiplex PCR. Multiplex PCR is a method of simultaneously amplifying multiple gene regions by using multiple primer pairs simultaneously in the PCR reaction system. Multiplex PCR can be performed using commercially available kits (for example, the Ion AmpliSeqTranscriptome Human Gene Expression Kit; Life Technologies Japan Co., Ltd., etc.). The temperature for the annealing and extension reactions in the PCR cannot be generalized as it depends on the primers used, but when using the above-mentioned multiplex PCR kit, it is preferably 62°C ± 1°C, more preferably 62°C ± 0.5°C, and even more preferably 62°C ± 0.25°C. Therefore, in the PCR, the annealing and extension reactions are preferably performed in one step. The duration of the annealing and extension reaction steps can be adjusted depending on the size of the DNA to be amplified, but is preferably 14 to 18 minutes. The conditions for the denaturation reaction in the PCR can be adjusted depending on the DNA to be amplified, but is preferably 95 to 99°C for 10 to 60 seconds. Reverse transcription and PCR at the above temperatures and times can be performed using a thermal cycler commonly used for PCR.

[0038] The purification of the reaction product obtained by the PCR is preferably carried out by size separation of the reaction product. Size separation allows the target PCR reaction product to be separated from primers and other impurities contained in the PCR reaction mixture. DNA size separation can be carried out, for example, by a size separation column, a size separation chip, or magnetic beads that can be used for size separation. Preferred examples of magnetic beads that can be used for size separation include Solid Phase Reversible Immobilization (SPRI) magnetic beads such as Ampure XP.

[0039] The purified PCR reaction product may be subjected to further processing necessary for subsequent quantitative analysis. For example, the purified PCR reaction product may be prepared into a suitable buffer solution for DNA sequencing, the PCR primer region in the PCR-amplified DNA may be cleaved, or adapter sequences may be further added to the amplified DNA. For instance, the purified PCR reaction product can be prepared into a buffer solution, the amplified DNA can be subjected to removal of PCR primer sequences and adapter ligation, and the resulting reaction product can be amplified as needed to prepare a library for quantitative analysis. These operations can be performed, for example, using the 5×VILO RT Reaction Mix included with the SuperScript® VILO cDNA Synthesis kit (Life Technologies Japan Co., Ltd.), the 5×Ion AmpliSeq HiFi Mix included with the Ion AmpliSeq Transcriptome Human Gene Expression Kit (Life Technologies Japan Co., Ltd.), and the Ion AmpliSeq Transcriptome Human Gene Expression Core Panel, according to the protocols included with each kit.

[0040] When measuring the expression level of a target gene or nucleic acid derived therefrom using Northern blot hybridization, for example, a probe DNA is first labeled with a radioisotope, a fluorescent substance, etc. Then, the resulting labeled DNA is hybridized with RNA derived from a biological sample transferred to a nylon membrane, etc., according to a conventional method. Subsequently, the double helix formed between the labeled DNA and RNA is measured by detecting the signal originating from the label.

[0041] When measuring the expression level of a target gene or nucleic acid derived therefrom using RT-PCR, for example, cDNA is first prepared from RNA derived from a biological sample according to a conventional method, and a pair of primers (a positive strand that binds to the cDNA (- strand), and a reverse strand that binds to the + strand) prepared to amplify the target gene of the present invention are hybridized with this cDNA as a template. Then, PCR is performed according to a conventional method, and the resulting amplified double-stranded DNA is detected. For the detection of the amplified double-stranded DNA, a method can be used to detect labeled double-stranded DNA produced by performing the above PCR using primers that have been previously labeled with an RI, a fluorescent substance, etc.

[0042] When measuring the expression level of a target gene or nucleic acid derived therefrom using a DNA microarray, for example, an array in which at least one nucleic acid (cDNA or DNA) derived from the target gene of the present invention is immobilized on a support is used, labeled cDNA or cRNA prepared from mRNA is bound to the microarray, and the expression level of mRNA can be measured by detecting the label on the microarray. The nucleic acid immobilized on the array can be any nucleic acid that hybridizes specifically (i.e., substantially only to the target nucleic acid) under stringent conditions. For example, it may be a nucleic acid having the entire sequence of the target gene of the present invention, or it may be a nucleic acid consisting of a partial sequence. Here, "partial sequence" refers to a nucleic acid consisting of at least 15 to 25 bases. Here, stringent conditions can typically be washing conditions of about "1×SSC, 0.1%SDS, 37°C", more stringent hybridization conditions can be about "0.5×SSC, 0.1%SDS, 42°C", and even more stringent hybridization conditions can be about "0.1×SSC, 0.1%SDS, 65°C". Hybridization conditions are described in J. Sambrook et al., Molecular Cloning: A Laboratory Manual, Third Edition, Cold Spring Harbor Laboratory Press (2001), etc.

[0043] When measuring the expression level of a target gene or nucleic acid derived therefrom by sequencing, for example, analysis can be performed using a next-generation sequencer (e.g., the Ion S5 / XL system, Life Technologies Japan Co., Ltd.). RNA expression can be quantified based on the number of reads generated by sequencing (read count).

[0044] The probes or primers used in the above measurements, namely primers for specifically recognizing and amplifying the target gene of the present invention or nucleic acids derived therefrom, or probes for specifically detecting the RNA or nucleic acids derived therefrom, can be designed based on the base sequence constituting the target gene. Here, "specifically recognizing" means, for example, in the Northern blotting method, that substantially only the target gene of the present invention or nucleic acids derived therefrom can be detected, or for example, in the RT-PCR method, that substantially only the nucleic acids in question are amplified, so that the detected substance or product can be determined to be the gene or nucleic acids derived therefrom. Specifically, the present invention can utilize DNA consisting of the base sequence constituting the target gene, or oligonucleotides containing a certain number of nucleotides complementary to its complementary strand. Here, "complementary strand" refers to the other strand of a double-stranded DNA consisting of A:T (U in the case of RNA) and G:C base pairs. Furthermore, "complementary" is not limited to cases where the sequence is perfectly complementary in the given number of consecutive nucleotide regions, but is preferably 80% or more, more preferably 90% or more, and even more preferably 95% or more of the base sequence identity. The base sequence identity can be determined by algorithms such as BLAST. When used as a primer, such oligonucleotides only need to be able to perform specific annealing and chain extension, and typically have a chain length of, for example, 10 bases or more, preferably 15 bases or more, more preferably 20 bases or more, and for example 100 bases or less, preferably 50 bases or less, and more preferably 35 bases or less. When used as a probe, it is sufficient that specific hybridization can be performed, and oligonucleotides that have at least a part or all of the sequence of DNA (or its complementary strand) consisting of the base sequence constituting the target gene of the present invention are used, and for example, oligonucleotides with a chain length of 10 bases or more, preferably 15 bases or more, and for example 100 bases or less, preferably 50 bases or less, and more preferably 25 bases or less are used. Here, "oligonucleotides" can be DNA or RNA, and may be synthetic or natural. Alternatively, the probes used for hybridization are usually labeled.

[0045] Furthermore, when measuring the translation product (protein) of the target gene of the present invention, molecules that interact with said protein, molecules that interact with RNA, or molecules that interact with DNA, methods such as protein chip analysis, immunoassay (e.g., ELISA), mass spectrometry (e.g., LC-MS / MS, MALDI-TOF / MS), 1-hybrid methods (PNAS 100, 12271-12276 (2003)), and 2-hybrid methods (Biol. Reprod. 58, 302-311 (1998)) can be used and can be appropriately selected depending on the target. For example, when a protein is used as the target of measurement, the procedure is carried out by contacting a biological sample with an antibody against the expression product of the present invention, detecting the polypeptide in the sample bound to the antibody, and measuring its level. For example, using the Western blotting method, the above antibody is used as the primary antibody, and then the primary antibody is labeled with a radioisotope, a fluorescent substance, or an enzyme, and the signal derived from these labeling substances is measured using a radiation detector, fluorescence detector, etc. Furthermore, the antibodies against the above-mentioned translation products may be polyclonal or monoclonal antibodies. These antibodies can be manufactured according to known methods. Specifically, polyclonal antibodies can be obtained by using proteins expressed and purified in E. coli or other bacteria according to conventional methods, or by synthesizing partial polypeptides of such proteins according to conventional methods, immunizing non-human animals such as rabbits, and then obtaining them from the serum of the immunized animals according to conventional methods. On the other hand, monoclonal antibodies can be obtained from hybridoma cells prepared by immunizing non-human animals such as mice with proteins expressed and purified in E. coli or other bacteria according to conventional methods, or with partial polypeptides of said proteins, and then fusing the resulting spleen cells with myeloma cells. Monoclonal antibodies may also be produced using phage display (Griffiths, AD; Duncan, AR, Current Opinion in Biotechnology, Volume 9, Number 1, February 1998, pp. 102-108(7)).

[0046] Thus, the expression level of the target gene of the present invention or its expression product in a biological sample taken from a subject is measured, and Parkinson's disease is detected based on the expression level. Specifically, the detection is performed by comparing the measured expression level of the target gene of the present invention or its expression product with a control level. When analyzing the expression levels of multiple target genes by sequencing, it is preferable to use the read count value, which is expression data, the RPM value obtained by correcting the difference in the total number of reads between samples, the value obtained by converting the RPM value to a base-2 logarithm (Log2RPM value), or the count value corrected using DESeq2 (Normalized count value) or the base-2 logarithm with an integer 1 added (Log2(count+1) value) as an indicator, as described above. Alternatively, values ​​calculated by fragments per kilobase of exon per million reads mapped (FPKM), reads per kilobase of exon per million reads mapped (RPKM), transcripts per million (TPM), etc., which are common quantitative values ​​for RNA-seq, may also be used. Furthermore, signal values ​​obtained by microarray methods and their corrected values ​​may also be used. In addition, when analyzing only specific target genes by RT-PCR, etc., it is preferable to analyze the expression level of the target gene by converting it to a relative expression level based on the expression level of housekeeping genes, or to analyze the copy number obtained by absolute quantification using a plasmid containing the region of the target gene. The copy number obtained by the digital PCR method may also be used. Here, "control level" refers, for example, to the expression level of the target gene or its expression product in healthy individuals. The expression level in healthy individuals may be a statistical value (e.g., the mean) of the expression level of the gene or its expression product measured from a healthy population. If there are multiple target genes, it is preferable to determine the reference expression level for each gene or its expression product.

[0047] Furthermore, the detection of Parkinson's disease in this invention can also be performed by increasing or decreasing the expression level of the target gene or its expression product. In this case, the expression level of the target gene or its expression product in a biological sample derived from a subject is compared with the cutoff value (reference value) of each gene or its expression product. The cutoff value can be appropriately determined based on statistical values ​​such as the mean value and standard deviation of the expression level, obtained in advance as reference data from the expression level of the target gene or its expression product in healthy individuals.

[0048] Furthermore, by using measured values ​​of the expression levels of target genes or their expression products derived from Parkinson's disease patients and those derived from healthy individuals, a discriminant formula (predictive model) can be constructed to distinguish between Parkinson's disease patients and healthy individuals, and Parkinson's disease can be detected using this discriminant formula. In other words, using measured values ​​of the expression levels of target genes or their expression products derived from Parkinson's disease patients and those derived from healthy individuals as training samples, a discriminant formula (predictive model) can be constructed to distinguish between Parkinson's disease patients and healthy individuals, and a cutoff value (reference value) for distinguishing between Parkinson's disease patients and healthy individuals can be determined based on this discriminant formula. Note that in creating the discriminant formula, dimensionality reduction can be performed using principal component analysis (PCA), and the major components can be used as explanatory variables. Then, by similarly measuring the level of the target gene or its expression product from a biological sample taken from the subject, substituting the obtained measurement values ​​into the discriminant formula, and comparing the results obtained from the discriminant formula with a reference value, the presence or absence of Parkinson's disease in the subject can be evaluated.

[0049] Here, the algorithm used to construct the discriminant can be a publicly known algorithm, such as one used in machine learning. Examples of machine learning algorithms include Random Forest, Support Vector Machine (SVM linear), Support Vector Machine (SVM rbf), Neural Network, Generalized Linear Model, Regularized Linear Discriminant Analysis, and Regularized Logistic Regression. By inputting validation data into the constructed prediction model and calculating the predicted values, the model that best fits the predicted values ​​to the observed values ​​can be selected. For example, the model with the largest F-score (recall, precision, and harmonic mean of these values) can be calculated from the predicted and observed values, and the model with the largest F-score can be selected as the optimal prediction model.

[0050] The method for determining the cutoff value (reference value) is not particularly restricted and can be determined according to known methods. For example, it can be obtained from an ROC (Receiver Operating Characteristic Curve) curve created using a discriminant. In an ROC curve, the vertical axis plots the probability of a positive result in a positive patient (sensitivity), and the horizontal axis plots the value obtained by subtracting the probability of a negative result in a negative patient (specificity) from 1 (false positive rate). Regarding the "true positive (sensitivity)" and "false positive (1-specificity)" shown in the ROC curve, the value (Youden index) at which "true positive (sensitivity)" - "false positive (1-specificity)" is maximized can be used as the cutoff value (reference value).

[0051] As shown in the examples described later, when predictive models were constructed using machine learning algorithms with the values ​​of each principal component obtained from the expression level data (Log2RPM values) of the target genes (33 genes, or 4 genes selected from them) shown in Table A as explanatory variables, and healthy individuals and Parkinson's disease patients as dependent variables, it was shown that Parkinson's disease could be predicted using models with four genes: SNORA16A, SNORA24, SNORA50, and REXO1L2P. Furthermore, it was shown that Parkinson's disease could be predicted with higher accuracy using models with all 33 genes.

[0052] Therefore, when creating a discriminant formula to distinguish between the Parkinson's disease patient group and the healthy control group, in addition to the four target genes SNORA16A, SNORA24, SNORA50, and REXO1L2P, it is advisable to appropriately add expression data of at least one gene or its expression product selected from the other 29 genes shown in Table A. Preferably, by adding an appropriate number of genes with high variable importance based on the variable importance shown in Table 8, a discriminant formula exhibiting a high detection rate and accuracy can be created, enabling more accurate detection of Parkinson's disease. Specifically, eight types are preferred: EGR2, RHOA, CCNI, RNASEK, CSF2RB, SERP1, ANKRD12, and SLC25A3. Twelve types are preferred, further including four types: CD83, CXCR4, ITGAX, and UQCRH. Eighteen types are preferred, further including six types: KCNQ1OT1, CCL3, C10orf116, SERPINB4, LCE3D, and CNFN. Even more preferably, all 29 types are included.

[0053] Alternatively, in addition to the four target genes SNORA16A, SNORA24, SNORA50, and REXO1L2P, it is also possible to appropriately add expression data for at least one gene other than SNORA24, or its expression product, selected from the 11 genes shown in Table C below, which are listed as differentially expressed genes in both Tables A and B mentioned above.

[0054] [Table C]

[0055] Furthermore, when creating a discriminant formula to separate the Parkinson's disease patient group from the healthy control group, it is possible to use the expression data of at least one gene or its expression product selected from the genes shown in Table B as the target gene, preferably including SNORA24, preferably using at least one other, more preferably using the expression data of the genes or their expression products shown in Table C, and even more preferably using the expression data of all the genes or their expression products shown in Table B.

[0056] The Parkinson's disease detection kit of the present invention contains a test reagent for measuring the expression level of the target gene or its expression product in a biological sample isolated from a patient. Specifically, this includes reagents for nucleic acid amplification and hybridization, including oligonucleotides (e.g., primers for PCR) that specifically bind (hybridize) to the target gene or nucleic acid derived therefrom, or reagents for immunological measurement, including antibodies that recognize the expression product (protein) of the target gene of the present invention. The oligonucleotides, antibodies, etc., included in the kit can be obtained by known methods as described above. Furthermore, the test kit may include, in addition to the antibodies and nucleic acids mentioned above, labeling reagents, buffers, chromogenic substrates, secondary antibodies, blocking agents, instruments and controls necessary for the test, and tools for collecting biological samples (for example, oil-absorbing film for collecting SSL).

[0057] Aspects and preferred embodiments of the present invention are shown below. <1> A method for detecting Parkinson's disease in a subject, comprising the step of measuring the expression level of at least one gene or its expression product selected from a group of four genes consisting of SNORA16A, SNORA24, SNORA50, and REXO1L2P in a biological sample taken from the subject. <2> This includes at least measuring the expression level of the SNORA24 gene or its expression product, <1> A method for detecting Parkinson's disease. <3> The expression level of a gene or its expression product is measured by the amount of mRNA expression. <1> or <2> The method. <4> The gene or its expression product is RNA contained in the lipids on the skin surface of the subject. <1> ~ <3> One of the following methods. <5> The measured expression levels are compared with reference values ​​for each of the aforementioned genes or their expression products to assess the presence or absence of Parkinson's disease. <1> ~ <4> One of the following methods. <6> Using the expression levels of the gene or its expression product derived from Parkinson's disease patients and the measured expression levels of the same gene or its expression product derived from healthy individuals as training samples, a discriminant formula is created to distinguish between Parkinson's disease patients and healthy individuals. The measured expression levels of the same gene or its expression product obtained from biological samples collected from subjects are substituted into this discriminant formula, and the obtained results are compared with a reference value to evaluate the presence or absence of Parkinson's disease in the subject. <1> ~ <4> One of the following methods. <7> The expression levels of all genes or their expression products from the aforementioned four gene groups are measured. <6> The method. <8> In addition to at least one gene selected from the four gene groups mentioned above, the expression level of at least one gene or its expression product selected from the following 29 gene groups is measured. <6> or <7> The method, ANKRD12, C10orf116, CCL3, CCNI, CD83, CNFN, CNN2, CSF2RB, CXCR4, EGR2, EMP1, ITGAX, KCNQ1OT1, LCE3D, LITAF, N DUFA4L2, NDUFS5, POLR2L, RHOA, RNASEK, RPL7A, RPS26, SERINC1, SERP1, SERPINB4, SLC25A3, SNRPG, SRRM2, UQCRH. <9> In addition to at least one gene selected from the four gene groups mentioned above, the expression level of at least one gene or its expression product selected from the following ten gene groups will be measured. <8> The method, CCL3, CCNI, CXCR4, EGR2, EMP1, POLR2L, RHOA, RNASEK, SERINC1, SERPINB4. <10> In addition to at least one gene selected from the four gene groups mentioned above, the expression level of at least one gene or its expression product selected from the following 16 gene groups will be measured. <6> or <7> The method, ANXA1, AQP3, ATP6V0C, BHLHE40, CCL3, CCNI, CXCR4, EGR2, EMP1, GABARAPL1, KRT16, POLR2L, RHOA, RNASEK, SERINC1, SERPINB4. <11> In addition to at least one gene selected from the four gene groups mentioned above, the expression level of at least one gene selected from the gene groups shown in Tables 3-1 to 3-4 and Tables 6-1 to 6-2 below (excluding the four genes mentioned above) or its expression product is measured. <6> or <7> The method. <12> In addition to at least one gene selected from the four gene groups mentioned above, the expression level of at least one gene or its expression product selected from the 1005 gene groups shown in Tables 1-1 to 1-27 and the 725 gene groups shown in Tables 4-1 to 4-20, excluding the four genes mentioned above, is measured. <6> or <7> The method. <13> The system contains an oligonucleotide that specifically hybridizes with the gene or nucleic acid derived therefrom, or an antibody that recognizes the expression product of the gene. <1> ~ <10> A test kit for detecting Parkinson's disease, used in one of the following methods. <14> Use of at least one gene or its expression product, selected from the gene groups shown in Tables 3-1 to 3-4 and Tables 6-1 to 6-2 below, as a detection marker for Parkinson's disease. <15> Use of at least one gene or its expression product selected from a group of four genes consisting of SNORA16A, SNORA24, SNORA50, and REXO1L2P as a detection marker for Parkinson's disease. <16> A detection marker for Parkinson's disease, comprising at least one gene or its expression product selected from the gene groups shown in Tables 3-1 to 3-4 and Tables 6-1 to 6-2 below. <17> It consists of at least one gene selected from a group of four genes consisting of SNORA16A, SNORA24, SNORA50, and REXO1L2P, or its expression product. <16> A marker for detecting Parkinson's disease. [Examples]

[0058] The present invention will be described in more detail below based on examples, but the present invention is not limited thereto.

[0059] Example 1: Detection of Parkinson's disease using RNA extracted from SSL 1) SSL collection The examination was conducted twice, as follows: Exam 1 and Exam 2. Study 1: The subjects included 15 healthy individuals (men and women aged 40-89) and 15 Parkinson's disease (PD) patients (men and women aged 40-89). Study 2: The subjects included 50 healthy individuals (men aged 40-89) and 50 individuals with Parkinson's disease (PD) (men aged 40-89). Each participant had previously been diagnosed with Parkinson's disease (Hoehn & Yahr stage I or II) by a neurologist. Sebum was collected from each subject's entire face using an oil-absorbing film (5 x 8 cm, made of polypropylene, 3M). The oil-absorbing film was then transferred to a vial and stored at -80°C for approximately one month until it was ready for RNA extraction.

[0060] 2) RNA preparation and sequencing The oil-absorbing film described in 1) above was cut to an appropriate size, and RNA was extracted using QIAzol Lysis Reagent (Qiagen) according to the included protocol. Based on the extracted RNA, cDNA was synthesized by reverse transcription at 42°C for 90 minutes using the SuperScript VILO cDNA Synthesis kit (Life Technologies Japan Co., Ltd.). Random primers included in the kit were used as primers for the reverse transcription reaction. From the obtained cDNA, a library containing DNA derived from the 20802 gene was prepared by multiplex PCR. Multiplex PCR was performed using the Ion AmpliSeqTranscriptome Human Gene Expression Kit (Life Technologies Japan Co., Ltd.) under the conditions [99°C, 2 min → (99°C, 15 sec → 62°C, 16 min) × 20 cycles → 4°C, Hold]. The obtained PCR products were purified with Ampure XP (Beckman Coulter, Inc.), followed by buffer reconstitution, primer sequence digestion, adapter ligation and purification, and amplification to prepare the library. The prepared library was loaded onto an Ion 540 Chip and sequenced using the Ion S5 / XL system (Life Technologies Japan Co., Ltd.).

[0061] 3) Data Analysis i) RNA expression analysis 1 In the RNA expression data (read count values) from subjects measured in 2) above, data with a read count of less than 10 were treated as missing values. These were converted to RPM values ​​corrected for differences in the total number of reads between samples, and the missing values ​​were then imputed using a method called Singular Value Decomposition (SVD) inputting. However, only genes for which expression data without missing values ​​was obtained for more than 80% of the subjects in all samples were used for the following analysis. For the analysis, in order to approximate the RPM values, which follow a negative binomial distribution, to a normal distribution, the RPM values ​​obtained by converting the RPM values ​​of the read counts to base 2 logarithms (Log2RPM values) were used.

[0062] Based on the SSL-derived RNA expression levels (Log2RPM values) obtained above for both healthy individuals and PD patients, we identified differentially expressed RNAs in PD patients with p-values ​​of 0.05 or less on the Student's t-test compared to healthy individuals. In Experiment 1, 111 RNAs were upexpressed and 68 RNAs were downexpressed in PD patients compared to healthy individuals (Tables 1-1 to 1-3) (Tables 1-4 to 1-5). In Experiment 2, 565 RNAs were upexpressed (Tables 1-6 to 1-19) and 294 RNAs were downexpressed (Tables 1-20 to 1-27). In both Experiment 1 and Experiment 2, 18 RNAs were upexpressed and 15 RNAs were downexpressed (genes shown in bold in the tables).

[0063] [Table 1-1]

[0064] [Table 1-2]

[0065] [Table 1-3]

[0066] [Table 1-4]

[0067] [Table 1-5]

[0068] [Table 1-6]

[0069] [Table 1-7]

[0070] Table 1-8

[0071] Table 1-9

[0072] Table 1-10

[0073] Table 1-11

[0074] Table 1-12

[0075] Table 1-13

[0076] Table 1-14

[0077] Table 1-15

[0078] Table 1-16

[0079] Table 1-17

[0080] Table 1-18

[0081] Table 1-19

[0082] Table 1-20

[0083] Table 1-21

[0084] Table 1-22

[0085] Table 1-23

[0086] Table 1-24

[0087] Table 1-25

[0088] Table 1-26

[0089] Table 1-27

[0090] Using the public database STRING, we explored the biological process (BP) and KEGG pathway through gene ontology (GO) enrichment analysis. As a result, 30 KEGG pathways associated with increased or decreased expression of genes were identified in Study 1 and 39 in Study 2, and both studies showed that the term hsa05012 (Parkinson's disease), which indicates Parkinson's disease, was included (Tables 2-1 to 2-2).

[0091] [Table 2-1]

[0092] [Table 2-2]

[0093] Regarding the genes whose expression was altered in at least one of Test 1 or Test 2, as shown in Tables 1-1 to 1-27, we reviewed previously published literature on their association with Parkinson's disease. Of the genes whose expression was altered in Test 1, the 21 shown in Table 3-1, and of the genes whose expression was altered in Test 2, the 92 shown in Tables 3-2 to 3-4, have not been previously reported to be associated with Parkinson's disease. Therefore, it became clear that these could serve as novel detection markers for Parkinson's disease. Note that genes shown in bold in the tables are common to both Test 1 and Test 2.

[0094] [Table 3-1]

[0095] [Table 3-2]

[0096] [Table 3-3]

[0097] [Table 3-4]

[0098] ii) RNA expression analysis 2 The RNA expression data (read count values) from subjects measured in 2) above were corrected using the DESeq2 method. However, samples in which 4161 or more genes were not detected were excluded, and only genes for which expression data without missing values ​​was obtained for 90% or more of the subjects in the remaining samples were used for the following analysis. The count values ​​corrected using the DESeq2 method (normalized count values) were used for the analysis.

[0099] Based on the SSL-derived RNA expression levels (normalized count values) obtained above for both healthy individuals and PD patients, we identified differentially expressed RNAs in PD patients whose p-value correction ratio (FDR) was 0.25 or less compared to healthy individuals. In Experiment 1, 74 RNAs were increased in expression in PD patients compared to healthy individuals (Tables 4-1 to 4-2), and 209 RNAs were decreased in expression (Tables 4-3 to 4-8). In Experiment 2, 151 RNAs were increased in expression (Tables 4-9 to 4-12), and 308 RNAs were decreased in expression (Tables 4-13 to 4-20). In both Experiment 1 and Experiment 2, 7 RNAs were commonly increased in expression, and 10 RNAs were commonly decreased in expression (genes shown in bold in the tables).

[0100] [Table 4-1]

[0101] [Table 4-2]

[0102] [Table 4-3]

[0103] Table 4-4

[0104] Table 4-5

[0105] Table 4-6

[0106] Table 4-7

[0107] Table 4-8

[0108] Table 4-9

[0109] Table 4-10

[0110] Table 4-11

[0111] Table 4-12

[0112] Table 4-13

[0113] [Table 4-14]

[0114] [Table 4-15]

[0115] [Table 4-16]

[0116] [Table 4-17]

[0117] [Table 4-18]

[0118] [Table 4-19]

[0119] [Table 4-20]

[0120] Using the public database STRING, we explored the biological process (BP) and KEGG pathway through gene ontology (GO) enrichment analysis. As a result, 30 KEGG pathways associated with increased or decreased expression of genes were identified in Study 1 and 28 in Study 2, and both studies showed that the term hsa05012 (Parkinson's disease), which indicates Parkinson's disease, was included (Tables 5-1 to 5-2).

[0121] [Table 5-1]

[0122] [Table 5-2]

[0123] Regarding the genes whose expression was altered in at least one of Test 1 or Test 2, as shown in Tables 4-1 to 4-20, we reviewed previously published literature on their association with Parkinson's disease. It was found that 19 genes showing altered expression in Test 1 (shown in Table 6-1) and 30 genes showing altered expression in Test 2 (shown in Table 6-2) have not been previously reported to be associated with Parkinson's disease. Therefore, these genes could potentially serve as novel detection markers for Parkinson's disease. Note that genes shown in bold in the tables are common to both Test 1 and Test 2.

[0124] [Table 6-1]

[0125] [Table 6-2]

[0126] Example 2: Creation and Verification of a Discriminant Model 1 1) Usage data Similar to RNA expression analysis 1 in Example 1, in the expression level data (read count values) of SSL-derived RNA from subjects, data with a read count of less than 10 were treated as missing values. After converting to RPM values ​​corrected for differences in the total number of reads between samples, the missing values ​​were imputed using a method called Singular Value Decomposition (SVD) inputting. However, only genes for which expression level data without missing values ​​was obtained in more than 80% of all samples were used for the following analysis. To construct the machine learning model, RPM values ​​that follow a negative binomial distribution were converted to base-2 logarithmic values ​​(Log2RPM values) to approximate a normal distribution.

[0127] 2) Dataset splitting From the RNA profile dataset obtained from the subjects of Study 1, RNA profile data from 20 individuals (10 healthy individuals and 10 individuals with PD) were used as training data for the PD prediction model, and the RNA profile data from the remaining 10 individuals were used as test data to evaluate the model's accuracy. From the RNA profile dataset obtained from the subjects of Study 2, RNA profile data from 80 individuals (40 healthy individuals and 40 individuals with PD) were used as training data for the PD prediction model, and the RNA profile data from the remaining 20 individuals were used as test data to evaluate the model's accuracy.

[0128] 3) Selection of feature genes In RNA expression analysis 1 of Example 1, 18 RNAs that were commonly elevated and 15 RNAs that were commonly decreased in expression in PD patients compared to healthy individuals in both Test 1 and Test 2 (genes shown in bold in Tables 1-1 to 1-27) were selected as feature genes. After converting their expression levels into principal components using principal component analysis, the first to tenth principal components were used as explanatory variables. Furthermore, from the 18 RNAs that were commonly elevated and 15 RNAs that were commonly decreased in expression in PD patients in both Test 1 and Test 2, four RNAs—SNORA16A, SNORA24, SNORA50, and REXO1L2P—were selected as feature genes. After converting their expression levels into principal components using principal component analysis, the first to fourth principal components were used as explanatory variables.

[0129] 4) Model construction A predictive model was constructed using the values ​​of each principal component obtained from the expression levels (Log2RPM values) of the feature genes selected from SSL-derived RNA, which served as the training data, as explanatory variables, and healthy individuals (HL) and PD as the dependent variables. For each item to be predicted, the predictive model was trained using seven algorithms: Random Forest, Linear Kernel Support Vector Machine (SVM linear), rbf Kernel Support Vector Machine (SVM rbf), Neural Network, Generalized Linear Model, Regularized Linear Discriminant Analysis, and Regularized Logistic Regression, with 10-fold cross-validation. For each algorithm, the values ​​of each principal component obtained from the feature gene expression levels (Log2RPM values) of the test data were input into the trained model, and the target predicted value for each item was calculated. The detection rate (Recall), accuracy (Precision), and its harmonic mean (F-score) were calculated from the predicted and actual values, and the model with the largest F-score was selected as the optimal prediction model.

[0130] 5) Results Table 7 shows the algorithm used, detection rate, accuracy, and F-score for each item being predicted. Figure 1 shows the confusion matrix plotting the predicted and actual values ​​for the optimal prediction model on the test data. The numbers in the figure indicate the sample size for each quadrant. Table 8 shows the results of calculating the variable importance of each feature gene when a model was constructed using random forest.

[0131] The F1 of the model using four RNAs—SNORA16A, SNORA24, SNORA50, and REXO1L2P—was 0.67 in Trial 1, 0.75 in Trial 2, and 0.76 when Trial 1 and Trial 2 were combined, demonstrating the ability to predict PD. The F1 of the model using a total of 33 RNAs—18 elevated and 15 decreased in PD patients—was 0.91 in Trial 1, 0.80 in Trial 2, and 0.82 when Trial 1 and Trial 2 were combined, demonstrating the ability to predict PD with higher accuracy.

[0132] [Table 7]

[0133] [Table 8]

[0134] Example 3: Creation and Verification of a Discriminant Model 2 1) Usage data Similar to RNA expression analysis 2 in Example 1, the expression level data (read count values) of SSL-derived RNA from subjects were corrected using the DESeq2 method. However, samples in which 4161 or more genes were not detected were excluded, and only genes for which expression level data was obtained without missing values ​​for 90% or more of the subjects in the remaining samples were used for the following analysis. The count values ​​corrected using the DESeq2 method (normalized count values) were used for the analysis.

[0135] 2) Dataset splitting From the RNA profile dataset obtained from the subjects of Study 1, RNA profile data from 15 individuals (9 healthy individuals and 6 individuals with PD) were used as training data for the PD prediction model, and RNA profile data from the remaining 5 individuals (4 healthy individuals and 1 individual with PD) were used as test data to evaluate the model's accuracy. From the RNA profile dataset obtained from the subjects of Study 2, RNA profile data from 72 individuals (37 healthy individuals and 35 individuals with PD) were used as training data for the PD prediction model, and RNA profile data from the remaining 24 individuals (13 healthy individuals and 11 individuals with PD) were used as test data to evaluate the model's accuracy.

[0136] 3) Selection of feature genes In RNA expression analysis 2 of Example 1, 17 RNAs (genes shown in bold in Tables 4-1 to 4-20) that showed increased or decreased expression in both PD patients compared to healthy individuals in Test 1 and Test 2 were selected as characteristic genes. After converting the expression data of these genes into principal components using principal component analysis, the first to fourth principal components were used as explanatory variables.

[0137] 4) Model construction A predictive model was constructed using the values ​​of each principal component obtained from the expression levels of the feature genes selected from SSL-derived RNA (normalized count value plus 1 and then the base-2 logarithm) as explanatory variables, and healthy individuals (HL) and PD as dependent variables. For each item to be predicted, the predictive model was trained using seven algorithms: Random Forest, Linear Kernel Support Vector Machine (SVM linear), rbf Kernel Support Vector Machine (SVM rbf), Neural Network, Generalized Linear Model, Regularized Linear Discriminant Analysis, and Regularized Logistic Regression, with 10-fold cross-validation. For each algorithm, the values ​​of each principal component obtained from the feature gene expression levels of the test data (normalized count value plus 1 and then the base-2 logarithm) were input into the trained model, and the target predicted value for each item was calculated. The detection rate (Recall), accuracy (Precision), and its harmonic mean (F-score) were calculated from the predicted and actual values, and the model with the largest F-score was selected as the optimal prediction model.

[0138] 5) Results Table 9 shows the algorithm used, detection rate, accuracy, and F-score for each item being predicted.

[0139] The results of the DESeq2-adjusted likelihood ratio tests for Test 1 and Test 2 showed that the F-scores for the model using 17 RNAs that were commonly elevated or decreased in expression were 1 for Test 1 and 0.87 for Test 2, indicating that PD can be predicted.

[0140] [Table 9]

[0141] Example 4: Creation and Verification of a Discriminant Model 3 1) Usage data Similar to RNA expression analysis 2 in Example 1, the expression level data (read count values) of SSL-derived RNA from subjects were corrected using the DESeq2 method. However, samples in which 4161 or more genes were not detected were excluded, and only genes for which expression level data was obtained without missing values ​​for 90% or more of the subjects in the remaining samples were used for the following analysis. The count values ​​corrected using the DESeq2 method (normalized count values) were used for the analysis.

[0142] 2) Dataset splitting From the RNA profile dataset obtained from the subjects of Study 1, RNA profile data from 15 individuals (9 healthy individuals and 6 individuals with PD) were used as training data for the PD prediction model, and RNA profile data from the remaining 5 individuals (4 healthy individuals and 1 individual with PD) were used as test data to evaluate the model's accuracy. From the RNA profile dataset obtained from the subjects of Study 2, RNA profile data from 72 individuals (37 healthy individuals and 35 individuals with PD) were used as training data for the PD prediction model, and RNA profile data from the remaining 24 individuals (13 healthy individuals and 11 individuals with PD) were used as test data to evaluate the model's accuracy.

[0143] 3) Selection of feature genes In RNA expression analysis 2 of Example 1, 19 RNAs (genes shown in Table 6-1) that showed increased or decreased expression in PD patients compared to healthy individuals in Study 1, or 30 RNAs (genes shown in Table 6-2) that showed increased or decreased expression in PD patients compared to healthy individuals in Study 2, were selected as feature genes. After converting the expression data of these genes into principal components using principal component analysis, the first to fourth principal components were used as explanatory variables.

[0144] 4) Model Construction A predictive model was constructed using the values ​​of each principal component obtained from the expression levels of the feature genes selected from SSL-derived RNA (normalized count value plus 1 and then the base-2 logarithm) as explanatory variables, and healthy individuals (HL) and PD as dependent variables. For each item to be predicted, the predictive model was trained using seven algorithms: Random Forest, Linear Kernel Support Vector Machine (SVM linear), rbf Kernel Support Vector Machine (SVM rbf), Neural Network, Generalized Linear Model, Regularized Linear Discriminant Analysis, and Regularized Logistic Regression, with 10-fold cross-validation. For each algorithm, the values ​​of each principal component obtained from the feature gene expression levels of the test data (normalized count value plus 1 and then the base-2 logarithm) were input into the trained model, and the target predicted value for each item was calculated. The detection rate (Recall), accuracy (Precision), and its harmonic mean (F-score) were calculated from the predicted and actual values, and the model with the largest F-score was selected as the optimal prediction model.

[0145] 5) Results Tables 10 and 11 show the algorithm used, detection rate, accuracy, and F-score for each item being predicted.

[0146] In the DESeq2-adjusted likelihood ratio test results for Study 1, the F-value of the model using 19 RNAs that showed increased or decreased expression but had not previously been reported to be associated with Parkinson's disease was 1, indicating that PD can be predicted. In the DESeq2-adjusted likelihood ratio test results for Study 2, the F-value of the model using 30 RNAs that showed increased or decreased expression but had not previously been reported to be associated with Parkinson's disease was 0.87, indicating that PD can be predicted.

[0147] [Table 10]

[0148] [Table 11]

[0149] Example 5: Creation and Verification of a Discriminant Model 4 1) Usage data Similar to RNA expression analysis 1 in Example 1, in the expression level data (read count values) of SSL-derived RNA from subjects, data with a read count of less than 10 were treated as missing values. After converting to RPM values ​​corrected for differences in the total number of reads between samples, the missing values ​​were imputed using a method called Singular Value Decomposition (SVD) inputting. However, only genes for which expression level data without missing values ​​was obtained in more than 80% of all samples were used for the following analysis. To construct the machine learning model, RPM values ​​that follow a negative binomial distribution were converted to base-2 logarithmic values ​​(Log2RPM values) to approximate a normal distribution.

[0150] 2) Dataset splitting From the RNA profile dataset obtained from the subjects of Study 1, RNA profile data from 20 individuals (10 healthy individuals and 10 individuals with PD) were used as training data for the PD prediction model, and the RNA profile data from the remaining 10 individuals were used as test data to evaluate the model's accuracy. From the RNA profile dataset obtained from the subjects of Study 2, RNA profile data from 80 individuals (40 healthy individuals and 40 individuals with PD) were used as training data for the PD prediction model, and the RNA profile data from the remaining 20 individuals were used as test data to evaluate the model's accuracy.

[0151] 3) Selection of feature genes In RNA expression analysis 1 of Example 1, 21 RNAs (genes shown in Table 3-1) that showed increased or decreased expression in PD patients compared to healthy individuals in Test 1, or 92 RNAs (genes shown in Tables 3-2 to 3-4) that showed increased or decreased expression in PD patients compared to healthy individuals in Test 2, were selected as feature genes. After converting the expression data of these genes into principal components using principal component analysis, the first to fourth principal components were used as explanatory variables.

[0152] 4) Model construction A predictive model was constructed using the values ​​of each principal component obtained from the expression levels (Log2RPM values) of the feature genes selected from SSL-derived RNA, which served as the training data, as explanatory variables, and healthy individuals (HL) and PD as the dependent variables. For each item to be predicted, the predictive model was trained using seven algorithms: Random Forest, Linear Kernel Support Vector Machine (SVM linear), rbf Kernel Support Vector Machine (SVM rbf), Neural Network, Generalized Linear Model, Regularized Linear Discriminant Analysis, and Regularized Logistic Regression, with 10-fold cross-validation. For each algorithm, the values ​​of each principal component obtained from the feature gene expression levels (Log2RPM values) of the test data were input into the trained model, and the target predicted value for each item was calculated. The detection rate (Recall), accuracy (Precision), and its harmonic mean (F-score) were calculated from the predicted and actual values, and the model with the largest F-score was selected as the optimal prediction model.

[0153] 5) Results Tables 12 and 13 show the algorithms used, detection rates, accuracy, and F-scores for the items to be predicted.

[0154] In the Log2RPM-corrected test results of Test 1, the F-score of a model using 21 RNAs that showed increased or decreased expression but had not previously been reported to be associated with Parkinson's disease was 0.91, indicating that PD can be predicted. In the Log2RPM-corrected test results of Test 2, the F-score of a model using 92 RNAs that showed increased or decreased expression but had not previously been reported to be associated with Parkinson's disease was 0.9, indicating that PD can be predicted.

[0155] [Table 12]

[0156] [Table 13]

Claims

1. A method for supporting the assessment of Parkinson's disease in a subject, which involves measuring the mRNA expression levels of all genes from a group of four genes consisting of SNORA16A, SNORA24, SNORA50, and REXO1L2P in surface lipids collected from the subject's skin, and predicting the presence or absence of Parkinson's disease in the subject based on these expression levels. This includes predicting the presence or absence of Parkinson's disease using a predictive model based on the mRNA expression level of the aforementioned gene, The method is characterized in that the predictive model is a predictive model constructed by machine learning using pre-trained data, with the measured expression level of the mRNA of the gene as the explanatory variable and the presence or absence of Parkinson's disease as the dependent variable.

2. A method for supporting the assessment of Parkinson's disease in a subject, which involves measuring the mRNA expression levels of all 33 genes in the gene group shown in Table 1 in skin surface lipids collected from the subject's skin, and predicting the presence or absence of Parkinson's disease in the subject based on said expression levels. This includes predicting the presence or absence of Parkinson's disease using a predictive model based on the mRNA expression level of the aforementioned gene, The method is characterized in that the predictive model is a predictive model constructed by machine learning using pre-trained data, with the measured expression level of the mRNA of the gene as the explanatory variable and the presence or absence of Parkinson's disease as the dependent variable. (Table 1)

3. A method for supporting the assessment of Parkinson's disease in a subject, which involves measuring the mRNA expression levels of all genes in the 17 gene groups shown in Table 2 in skin surface lipids collected from the subject's skin, and predicting the presence or absence of Parkinson's disease in the subject based on said expression levels. This includes predicting the presence or absence of Parkinson's disease using a predictive model based on the mRNA expression level of the aforementioned gene, The method is characterized in that the predictive model is a predictive model constructed by machine learning using pre-trained data, with the measured expression level of the mRNA of the gene as the explanatory variable and the presence or absence of Parkinson's disease as the dependent variable. (Table 2)

4. A method for supporting the assessment of Parkinson's disease in a subject, which involves measuring the mRNA expression levels of all genes in the 19 gene groups shown in Table 3 in skin surface lipids collected from the subject's skin, and predicting the presence or absence of Parkinson's disease in the subject based on said expression levels. This includes predicting the presence or absence of Parkinson's disease using a predictive model based on the mRNA expression level of the aforementioned gene, The method is characterized in that the predictive model is a predictive model constructed by machine learning using pre-trained data, with the measured expression level of the mRNA of the gene as the explanatory variable and the presence or absence of Parkinson's disease as the dependent variable. (Table 3)

5. A method for supporting the assessment of Parkinson's disease in a subject, which involves measuring the mRNA expression levels of all 30 genes in the gene group shown in Table 4 in skin surface lipids collected from the subject's skin, and predicting the presence or absence of Parkinson's disease in the subject based on said expression levels. This includes predicting the presence or absence of Parkinson's disease using a predictive model based on the mRNA expression level of the aforementioned gene, The method is characterized in that the predictive model is a predictive model constructed by machine learning using pre-trained data, with the measured expression level of the mRNA of the gene as the explanatory variable and the presence or absence of Parkinson's disease as the dependent variable. (Table 4)

6. A method for supporting the assessment of Parkinson's disease in a subject, which involves measuring the mRNA expression levels of all genes in the 21 gene groups shown in Table 5 in skin surface lipids collected from the subject's skin, and predicting the presence or absence of Parkinson's disease in the subject based on said expression levels. This includes predicting the presence or absence of Parkinson's disease using a predictive model based on the mRNA expression level of the aforementioned gene, The method is characterized in that the predictive model is a predictive model constructed by machine learning using pre-trained data, with the measured expression level of the mRNA of the gene as the explanatory variable and the presence or absence of Parkinson's disease as the dependent variable. (Table 5)

7. A method for supporting the assessment of Parkinson's disease in a subject, which involves measuring the mRNA expression levels of all genes from the 92 gene groups shown in Table 6 in skin surface lipids collected from the subject's skin, and predicting the presence or absence of Parkinson's disease in the subject based on said expression levels. This includes predicting the presence or absence of Parkinson's disease using a predictive model based on the mRNA expression level of the aforementioned gene, The method is characterized in that the predictive model is a predictive model constructed by machine learning using pre-trained data, with the measured expression level of the mRNA of the gene as the explanatory variable and the presence or absence of Parkinson's disease as the dependent variable. (Table 6)

Citation Information

Patent Citations

  • Method for early diagnosis of parkinson disease

    JP2016075644A

  • Circulating serum microRNA biomarkers and methods

    JP2019506183A

  • Method for preparing nucleic acid sample

    WO2018008319A1

  • Biomarkers and uses thereof

    WO2020025967A1