Methods for detecting the severity of premenstrual syndrome

By analyzing RNA expression in skin surface lipids from specific genes, PMS severity is objectively detected, addressing the limitations of current symptom-based diagnosis and enabling targeted interventions.

JP2026090049APending Publication Date: 2026-06-02KAO CORP

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
KAO CORP
Filing Date
2024-11-21
Publication Date
2026-06-02

Smart Images

  • Figure 2026090049000001
    Figure 2026090049000001
  • Figure 2026090049000002
    Figure 2026090049000002
  • Figure 2026090049000003
    Figure 2026090049000003
Patent Text Reader

Abstract

To provide a PMS severity marker that enables the detection of PMS severity, and a method for detecting PMS severity using the said detection marker. [Solution] A method for detecting the severity of premenstrual syndrome in a subject, comprising measuring the expression levels of at least 10 genes or their expression products selected from the 30 gene groups shown in Tables 1 and 2, in a biological sample taken from the subject.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for detecting the severity of premenstrual syndrome, a test kit for detecting the severity of premenstrual syndrome, and a detection marker for detecting the severity of premenstrual syndrome.

Background Art

[0002] Physical and mental discomfort in women before and / or during menstruation is called menstrual-associated symptoms, and is typified by premenstrual syndrome (Premenstrual Syndrome; hereinafter also referred to as PMS), and dysmenorrhea. Since it has been pointed out that PMS is also related to the labor productivity of women, alleviating or reducing menstrual-associated symptoms typified by PMS and dysmenorrhea and improving the QOL (Quality of Life) of women is not only hygienically important but also socially important. However, in many cases of PMS, the menstrual function itself is normal and no abnormalities are found in blood tests or imaging tests, so the only way to diagnose PMS is to observe the relationship between symptoms and the menstrual cycle.

[0003] On the other hand, techniques for examining the current and even future physiological states in the human body by analyzing nucleic acids such as DNA and RNA in biological samples have been developed. Analysis using nucleic acids has advantages such as a comprehensive analysis method being established and abundant information being obtained by a single analysis, and functional association of analysis results being easy based on many research reports on single nucleotide polymorphisms and RNA functions. Nucleic acids derived from living organisms can be extracted from body fluids such as blood, secretions, tissues, etc. Recently, it has been reported that RNA contained in skin surface lipids (SSL) can be used as a sample for biological analysis, and marker genes of the epidermis, sweat glands, hair follicles, and sebaceous glands can be detected from SSL (Patent Document 1). Furthermore, Patent Document 2 reports that 51 marker genes for evaluating PMS severity can be detected from SSL.

Prior Art Documents

Patent Documents

[0004] [Patent Document 1] International Public Gazette No. 2018 / 008319 [Patent Document 2] Japanese Patent Publication No. 2021-175395 [Overview of the Initiative] [Problems that the invention aims to solve]

[0005] The present invention relates to a PMS severity marker that enables the detection of the severity of PMS, and a method for detecting the severity of PMS using the detection marker. [Means for solving the problem]

[0006] The inventors collected SSL from the skin of subjects with varying degrees of PMS severity and comprehensively analyzed the RNA expression status contained in SSL as sequencing information. As a result, they found that the expression levels of certain genes correlate with the severity of PMS, and that the severity of PMS can be detected using the expression levels of at least 10 of these specific genes as indicators.

[0007] In other words, the present invention relates to the following 1) to 3). 1) A method for detecting the severity of premenstrual syndrome in a subject, comprising measuring the expression levels of at least 10 genes or their expression products selected from the 30 gene groups shown in Tables 1 and 2 below, in a biological sample taken from the subject. 2) A test kit for detecting the severity of premenstrual syndrome used in the method described in 1), comprising an oligonucleotide that specifically hybridizes with the gene or nucleic acid derived therefrom, or an antibody that recognizes the expression product of the gene. 3) A detection marker for detecting the severity of premenstrual syndrome, comprising at least 10 genes or their expression products selected from the 30 gene groups shown in Tables 1 and 2 below.

[0008] [Table 1]

[0009] [Table 2] [Effects of the Invention]

[0010] According to the present invention, it is possible to detect the severity of PMS simply and non-invasively. This allows for an objective assessment of PMS severity based on indicators, enabling appropriate measures to be taken against PMS. [Modes for carrying out the invention]

[0011] All patent, non-patent, and other publications cited herein are incorporated herein by reference in their entirety.

[0012] In this invention, the terms "nucleic acid" or "polynucleotide" mean DNA or RNA. DNA includes cDNA, genomic DNA, and synthetic DNA, and RNA includes total RNA, mRNA, rRNA, tRNA, non-coding RNA, and synthetic RNA.

[0013] In this invention, "gene" means a double-stranded DNA including human genomic DNA, as well as single-stranded DNA (positive strand) including cDNA, single-stranded DNA (complementary strand) having a sequence complementary to the positive strand, and fragments thereof, in which the sequence information of the bases constituting the DNA contains some kind of biological information. Furthermore, the term "gene" in question encompasses not only genes represented by a specific base sequence, but also nucleic acids encoding their homologs (i.e., orthologs), variants such as genetic polymorphisms, and derivatives.

[0014] In this invention, gene names follow the Official Symbols listed in NCBI ([www.ncbi.nlm.nih.gov / ]).

[0015] In this invention, the term "expression product" of a gene is a concept that encompasses both the transcript and the translation product of a gene. A "transcript" is RNA produced by transcription from a gene (DNA), and a "translation product" is a protein encoded by a gene that is synthesized through translation based on RNA.

[0016] In this invention, "Premenstrual Syndrome (PMS)" is a group of physical or mental symptoms that occur about a week before menstruation (luteal phase), first reported in 1931 by Frank (Archives of Neurology and Psychiatry, 1931, 26:1053). Medically, PMS is defined as "physical and mental symptoms that begin 3 to 10 days before the start of menstruation and decrease or disappear with the onset of menstruation." Typical symptoms of PMS include, as physical symptoms, lower abdominal pain, back pain, bloating, headache, stiff shoulders, cold hands and feet, increased appetite, diarrhea, constipation, swelling, breast pain, breast tenderness, acne, rough skin, fatigue, drowsiness, etc., and as mental symptoms, irritability, anger, aggression, depression, tearfulness, increased anxiety, etc.

[0017] In this specification, the severity of PMS refers to the degree of severity of the physical or mental symptoms of PMS, and is classified, for example, as asymptomatic (no symptoms), mild (mild degree), moderate (moderate degree), severe (severe degree). The severity of PMS can be determined, for example, based on the evaluation score of the Japanese version of the Daily Records of Severity of Problems (DRSP-J, Clinical Gynecology and Obstetrics (in Japanese). 2019;73(8):807-811) (hereinafter also referred to as the DRSP score). DRSP is a PMS symptom evaluation tool form used as a global standard, and DRSP-J has been proven to have sufficient reliability and validity for a wide range of subjects (International Journal of Womens Health 2024 Feb 26;16:299-308., International Journal of Womens Health 2021 Mar 29;13:361-367.). The higher the DRSP score, the higher the severity of PMS is evaluated. For example, evaluation is performed using the average value over 5 days of the total DRSP (items 1 to 21) score in the late luteal phase (5 days), and the higher the score, the more severe the PMS symptoms can be evaluated. In the present invention, the severity of PMS is preferably the severity evaluated based on the DRSP score.

[0018] In the present invention, the "detection" of the severity of PMS means to clarify the severity of PMS, and can also be paraphrased by terms such as examination, measurement, determination, or evaluation support. In this specification, the terms "detection", "examination", "measurement", "determination", or "evaluation" do not include the diagnosis of the severity of PMS by a doctor.

[0019] As shown in the examples described below, when gene expression analysis was performed on SSL collected from the entire face of PMS patients and asymptomatic individuals, 37 genes showing a positive correlation with the DRSP score of the severity of PMS, 165 genes showing a negative correlation with the score, and a total of 202 genes were identified. Among the 202 gene groups, 15 genes each with a larger absolute value of the correlation coefficient, specifically 15 genes from the genes showing a positive correlation with the DRSP score of the severity of PMS (Table 3 below) and 15 genes from the genes showing a negative correlation with the score (Table 4 below), a total of 30 genes were selected. Among the 30 gene groups, it was shown that it is possible to detect the severity of PMS using a discriminant (prediction model of the DRSP score) using 10 or more genes as characteristic genes. Therefore, at least 10 genes selected from the 30 gene groups shown in Table 3 and Table 4 below or their expression products can be detection markers for detecting the severity of PMS. In this case, the 15 genes showing a positive correlation with the DRSP score in Table 3 are positive markers whose expression levels increase stepwise with the severity of PMS, and the 15 genes showing a negative correlation with the DRSP score in Table 4 are negative markers whose expression levels decrease stepwise with the severity of PMS.

[0020]

Table 3

[0021]

Table 4

[0022] The 30 genes shown in Table 3 and Table 4 have not been reported to be related to PMS so far, and the genes selected from these gene groups or their expression products are novel detection markers for detecting the severity of PMS. In one embodiment, the detection marker of the present invention is a nucleic acid marker such as the DNA of the gene or its transcript, RNA. In another embodiment, the detection marker of the present invention is a protein marker which is the translation product of the gene. Preferably, the detection marker of the present invention is a nucleic acid marker, and more preferably, an RNA marker.

[0023] In the present invention, at least 10 genes or their expression products selected from the 30 gene groups shown in Tables 3 and 4 above are used as detection markers, and the severity of PMS can be detected based on their expression levels. From the viewpoint of improving accuracy, it is preferable to select 10 or more genes, preferably 15 or more, more preferably 20 or more, and even more preferably all 30 genes from the 30 gene groups shown in Tables 3 and 4 below. In this case, it is preferable that the selected genes include at least one gene from the 15 gene groups shown in Table 3 and at least one gene from the 15 gene groups shown in Table 4. Furthermore, in the detection of PMS severity according to the present invention, it is preferable to select at least 10 genes shown in Table 5 below. Furthermore, from the viewpoint of improving accuracy, it is preferable to select genes in which the absolute value of the correlation coefficient shown in the examples described later is preferably 0.38 or higher, and more preferably 0.42 or higher.

[0024] [Table 5]

[0025] The genes that can serve as detection markers for detecting the severity of PMS (hereinafter also referred to as "target genes") include genes that have substantially identical base sequences to the DNA that constitutes the target gene, insofar as they can serve as biomarkers for detecting the severity of PMS. Here, substantially identical base sequences mean, for example, that when searching using the homology calculation algorithm NCBI BLAST with the conditions expected value = 10; gap allowed; filtering = ON; match score = 1; mismatch score = -3, the base sequence of the target gene has 90% or more identity with the DNA that constitutes the target gene, preferably 95% or more, more preferably 98% or more, and even more preferably 99% or more.

[0026] The present invention provides a method for detecting the severity of PMS, which, in one embodiment, includes measuring the expression levels of at least 10 genes or their expression products selected from the 30 gene groups shown in Tables 3 and 4, in a biological sample taken from a subject.

[0027] The subjects in this invention are not particularly limited in terms of age, race, etc., and include subjects who menstruate and desire or need to have the severity of PMS detected. Preferably, the subjects are in the luteal phase of the menstrual cycle (follicular phase, ovulation phase, and luteal phase).

[0028] The biological samples used in the present invention may be any cells, tissues, and biomaterials in which the gene expression of the present invention is altered. Specifically, examples include organs, skin, blood, urine, saliva, sweat, stratum corneum, surface lipids (SSL), tissue exudates and other bodily fluids, serum and plasma prepared from blood, and others such as feces and hair. Preferably, skin or surface lipids (SSL), and more preferably, surface lipids (SSL). In the present invention, the skin site from which SSL is collected is not particularly limited and may include skin from any part of the body such as the head, face, neck, trunk, hands and feet, but preferably, the face, and more preferably, the entire face.

[0029] Here, "superficial lipids (SSL)" refers to the lipid-soluble fraction present on the surface of the skin, and is sometimes called sebum. Generally, SSL mainly consists of secretions from exocrine glands such as sebaceous glands in the skin, and exists on the skin surface in the form of a thin layer covering the skin surface. SSL contains RNA expressed in skin cells. (See Patent Document 1 above). Furthermore, in the present invention, unless otherwise specified, "skin" is a general term for the region including the stratum corneum, epidermis, dermis, hair follicles, and tissues such as sweat glands, sebaceous glands, and other glands.

[0030] Any means used for the collection or removal of SSL from the skin can be employed to collect SSL from the subject's skin. Preferably, an SSL absorbent material, an SSL adhesive material, or an instrument for scraping off SSL from the skin, as described later, can be used. The SSL absorbent material or SSL adhesive material is not particularly limited as long as it is a material that has an affinity for SSL, and examples include polypropylene and pulp. More detailed examples of procedures for collecting SSL from the skin include methods of absorbing SSL onto a sheet material such as oil-blotting paper or oil-blotting film, methods of adhering SSL to a glass plate or tape, and methods of scraping off and collecting SSL with a spatula, scraper, etc. To improve the adsorption of SSL, an SSL absorbent material containing a highly lipid-soluble solvent beforehand may be used. On the other hand, since the adsorption of SSL is inhibited if the SSL absorbent material contains a highly water-soluble solvent or water, it is preferable that the content of highly water-soluble solvents or water is low. It is preferable to use the SSL absorbent material in a dry state.

[0031] SSL collected from subjects may be stored for a certain period of time. To minimize the degradation of the contained RNA, it is preferable to store the collected SSL under low temperature conditions as quickly as possible after collection. The storage temperature conditions for SSL in this invention may be 0°C or lower, preferably -20±20°C to -80±20°C, more preferably -20±10°C to -80±10°C, even more preferably -20±20°C to -40±20°C, even more preferably -20±10°C to -40±10°C, even more preferably -20±10°C, and even more preferably -20±5°C. The storage period for SSL under these low temperature conditions is not particularly limited, but is preferably 12 months or less, for example, 6 hours to 12 months, more preferably 6 months or less, for example, 1 day to 6 months, and even more preferably 3 months or less, for example, 3 days to 3 months.

[0032] In the present invention, the objects to be measured for the expression level of a target gene or its expression product include RNA, the DNA encoding that RNA, the protein encoded by that RNA, molecules that interact with that protein, molecules that interact with that RNA, or molecules that interact with that DNA, with RNA being preferred and mRNA being more preferred. Here, molecules that interact with RNA, DNA, or proteins include DNA, RNA, proteins, polysaccharides, oligosaccharides, monosaccharides, lipids, fatty acids, and their phosphorylated, alkylated, and glycosidic compounds, as well as complexes of any of the above. Furthermore, the expression level comprehensively refers to the amount of expression or activity of the gene or expression product in question.

[0033] In the method of the present invention, in a preferred embodiment, SSL is used as the biological sample. In this case, the expression level of mRNA contained in SSL is analyzed. Specifically, the RNA is converted to cDNA by reverse transcription, and then the cDNA or its amplified product is measured. For RNA extraction from SSL, methods commonly used for RNA extraction or purification from biological samples can be employed, such as the phenol / chloroform method, the AGPC (acid guanidinium thiocyanate-phenol-chloroform extraction) method, or methods using columns such as TRIzol®, RNeasy®, or QIAzol®, or methods using special silica-coated magnetic particles, methods using Solid Phase Reversible Immobilization magnetic particles, or extraction using commercially available RNA extraction reagents such as ISOGEN.

[0034] For the reverse transcription, primers targeting a specific RNA to be analyzed may be used, but for more comprehensive nucleic acid preservation and analysis, random primers are preferable. A general reverse transcriptase or reverse transcription reagent kit can be used for the reverse transcription. Preferably, a reverse transcriptase or reverse transcription reagent kit with high accuracy and efficiency is used, such as M-MLV Reverse Transcriptase and its variants, or commercially available reverse transcriptase or reverse transcription reagent kits, such as the PrimeScript® Reverse Transcriptase series (Takara Bio Inc.) and the SuperScript® Reverse Transcriptase series (Thermo Scientific Inc.). SuperScript® III Reverse Transcriptase and SuperScript® VILO cDNA Synthesis kit (both from Thermo Scientific Inc.) are preferably used. In the reverse transcription extension reaction, it is preferable to adjust the temperature to preferably 42°C ± 1°C, more preferably 42°C ± 0.5°C, and even more preferably 42°C ± 0.25°C, while adjusting the reaction time to preferably 60 minutes or more, more preferably 80 to 120 minutes.

[0035] Methods for measuring expression levels can be selected from nucleic acid amplification methods such as PCR, real-time RT-PCR, multiplex PCR, SmartAmp, and LAMP, which use DNA that hybridizes to RNA, cDNA, or DNA as primers; hybridization methods (DNA chips, DNA microarrays, dot blot hybridization, slot blot hybridization, Northern blot hybridization, etc.) which use nucleic acids that hybridize to these as probes; methods for determining the base sequence (sequencing); or methods combining these.

[0036] In PCR, a primer pair targeting a specific DNA to be analyzed may be used to amplify only that specific DNA, or multiple primer pairs may be used to amplify multiple specific DNAs simultaneously. Preferably, the PCR is multiplex PCR. Multiplex PCR is a method of simultaneously amplifying multiple gene regions by using multiple primer pairs simultaneously in the PCR reaction system. Multiplex PCR can be performed using commercially available kits (for example, the Ion AmpliSeqTranscriptome Human Gene Expression Kit; Life Technologies Japan Co., Ltd., etc.). The temperature for the annealing and extension reactions in the PCR cannot be generalized as it depends on the primers used, but when using the above-mentioned multiplex PCR kit, it is preferably 62°C ± 1°C, more preferably 62°C ± 0.5°C, and even more preferably 62°C ± 0.25°C. Therefore, in the PCR, the annealing and extension reactions are preferably performed in one step. The duration of the annealing and extension reaction steps can be adjusted depending on the size of the DNA to be amplified, but is preferably 14 to 18 minutes. The conditions for the denaturation reaction in the PCR can be adjusted depending on the DNA to be amplified, but is preferably 95 to 99°C for 10 to 60 seconds. Reverse transcription and PCR at the above temperatures and times can be performed using a thermal cycler commonly used for PCR.

[0037] The purification of the reaction product obtained by the PCR is preferably carried out by size separation of the reaction product. Size separation allows the target PCR reaction product to be separated from primers and other impurities contained in the PCR reaction mixture. DNA size separation can be carried out, for example, by a size separation column, a size separation chip, or magnetic beads that can be used for size separation. Preferred examples of magnetic beads that can be used for size separation include Solid Phase Reversible Immobilization (SPRI) magnetic beads such as Ampure XP.

[0038] The purified PCR reaction product may be subjected to further processing necessary for subsequent quantitative analysis. For example, the purified PCR reaction product may be prepared into a suitable buffer solution for DNA sequencing, the PCR primer region in the PCR-amplified DNA may be cleaved, or adapter sequences may be further added to the amplified DNA. For instance, the purified PCR reaction product can be prepared into a buffer solution, the amplified DNA can be subjected to removal of PCR primer sequences and adapter ligation, and the resulting reaction product can be amplified as needed to prepare a library for quantitative analysis. These operations can be performed, for example, using the 5×VILO RT Reaction Mix included with the SuperScript® VILO cDNA Synthesis kit (Life Technologies Japan Co., Ltd.), the 5×Ion AmpliSeq HiFi Mix included with the Ion AmpliSeq Transcriptome Human Gene Expression Kit (Life Technologies Japan Co., Ltd.), and the Ion AmpliSeq Transcriptome Human Gene Expression Core Panel, according to the protocols included with each kit.

[0039] When measuring the expression level of a target gene or nucleic acid derived therefrom using Northern blot hybridization, for example, a probe DNA is first labeled with a radioisotope, a fluorescent substance, etc. Then, the resulting labeled DNA is hybridized with RNA derived from a biological sample transferred to a nylon membrane, etc., according to a conventional method. Subsequently, the double helix formed between the labeled DNA and RNA is measured by detecting the signal originating from the label.

[0040] When measuring the expression level of a target gene or nucleic acid derived therefrom using RT-PCR, for example, cDNA is first prepared from RNA derived from a biological sample according to a conventional method, and a pair of primers (a positive strand that binds to the cDNA (- strand), and a reverse strand that binds to the + strand) prepared to amplify the target gene of the present invention are hybridized with this cDNA as a template. Then, PCR is performed according to a conventional method, and the resulting amplified double-stranded DNA is detected. For the detection of the amplified double-stranded DNA, a method can be used to detect labeled double-stranded DNA produced by performing the above PCR using primers that have been previously labeled with an RI, a fluorescent substance, etc.

[0041] When measuring the expression level of a target gene or nucleic acid derived therefrom using a DNA microarray, for example, an array in which at least one nucleic acid (cDNA or DNA) derived from the target gene of the present invention is immobilized on a support is used, labeled cDNA or cRNA prepared from mRNA is bound to the microarray, and the expression level of mRNA can be measured by detecting the label on the microarray. The nucleic acid immobilized on the array can be any nucleic acid that hybridizes specifically (i.e., substantially only to the target nucleic acid) under stringent conditions. For example, it may be a nucleic acid having the entire sequence of the target gene of the present invention, or it may be a nucleic acid consisting of a partial sequence. Here, "partial sequence" refers to a nucleic acid consisting of at least 15 to 25 bases. Here, stringent conditions can typically be washing conditions of about "1×SSC, 0.1%SDS, 37°C", more stringent hybridization conditions can be about "0.5×SSC, 0.1%SDS, 42°C", and even more stringent hybridization conditions can be about "0.1×SSC, 0.1%SDS, 65°C". Hybridization conditions are described in J. Sambrook et al., Molecular Cloning: A Laboratory Manual, Third Edition, Cold Spring Harbor Laboratory Press (2001), etc.

[0042] When measuring the expression level of a target gene or nucleic acid derived therefrom by sequencing, for example, analysis can be performed using a next-generation sequencer (e.g., the Ion S5 / XL system, Life Technologies Japan Co., Ltd.). RNA expression can be quantified based on the number of reads generated by sequencing (read count).

[0043] The probes or primers used in the above measurements, namely primers for specifically recognizing and amplifying the target gene of the present invention or nucleic acids derived therefrom, or probes for specifically detecting the RNA or nucleic acids derived therefrom, can be designed based on the base sequence constituting the target gene. Here, "specifically recognizing" means, for example, in the Northern blotting method, that substantially only the target gene of the present invention or nucleic acids derived therefrom can be detected, or for example, in the RT-PCR method, that substantially only the nucleic acids in question are amplified, so that the detected substance or product can be determined to be the gene or nucleic acids derived therefrom. Specifically, the present invention can utilize DNA consisting of a base sequence constituting the target gene, or oligonucleotides containing a certain number of nucleotides complementary to its complementary strand. Here, "complementary strand" refers to the other strand of a double-stranded DNA consisting of A:T (U in the case of RNA) and G:C base pairs. Furthermore, "complementary" is not limited to cases where the sequence is perfectly complementary in the given number of consecutive nucleotide regions, but is preferable to have 80% or more, more preferably 85% or more, even more preferably 90% or more, even more preferably 95% or more, and even more preferably 98% or more of identity in the base sequence. The identity of the base sequence can be determined by algorithms such as BLAST. When used as a primer, such oligonucleotides only need to be able to perform specific annealing and chain extension, and typically have a chain length of, for example, 10 bases or more, preferably 15 bases or more, more preferably 20 bases or more, and for example 100 bases or less, preferably 50 bases or less, and more preferably 35 bases or less. When used as a probe, it is sufficient that specific hybridization can be performed, and oligonucleotides that have at least a part or all of the sequence of DNA (or its complementary strand) consisting of the base sequence constituting the target gene of the present invention are used, and for example, oligonucleotides with a chain length of 10 bases or more, preferably 15 bases or more, and for example 100 bases or less, preferably 50 bases or less, and more preferably 25 bases or less are used. Here, "oligonucleotides" can be DNA or RNA, and may be synthetic or naturally occurring. Furthermore, the probes used for hybridization are usually labeled.

[0044] Furthermore, when measuring the translation product (protein) of the target gene of the present invention, molecules that interact with said protein, molecules that interact with RNA, or molecules that interact with DNA, methods such as protein chip analysis, immunoassay (e.g., ELISA), mass spectrometry (e.g., LC-MS / MS, MALDI-TOF / MS), 1-hybrid methods (PNAS 100, 12271-12276 (2003)), and 2-hybrid methods (Biol. Reprod. 58, 302-311 (1998)) can be used and can be appropriately selected depending on the target. For example, when a protein is used as the target of measurement, the method is carried out by contacting a biological sample with an antibody that specifically recognizes the expression product of the present invention, specifically an antibody that recognizes a structural characteristic site (epitope) that can distinguish the expression product protein from other proteins, detecting the polypeptide or protein in the sample bound to the antibody, and measuring its level. For example, using the Western blotting method, the above antibody is used as the primary antibody, and then the primary antibody is labeled with a radioisotope, fluorescent substance, or enzyme, and the signal derived from these labeling substances is measured with a radiation detector, fluorescence detector, etc. Furthermore, the antibodies against the above-mentioned translation products may be polyclonal or monoclonal antibodies. These antibodies can be manufactured according to known methods. Specifically, polyclonal antibodies can be obtained by using proteins expressed and purified in E. coli or other bacteria according to conventional methods, or by synthesizing partial polypeptides of such proteins according to conventional methods, immunizing non-human animals such as rabbits, and then obtaining them from the serum of the immunized animals according to conventional methods. On the other hand, monoclonal antibodies can be obtained from hybridoma cells prepared by immunizing non-human animals such as mice with proteins expressed and purified in E. coli or other bacteria according to conventional methods, or with partial polypeptides of said proteins, and then fusing the resulting spleen cells with myeloma cells. Monoclonal antibodies may also be produced using phage display (Griffiths, AD; Duncan, AR, Current Opinion in Biotechnology, Volume 9, Number 1, February 1998, pp. 102-108(7)).

[0045] Thus, the expression level of the target gene of the present invention or its expression product in a biological sample taken from a subject is measured, and the severity of PMS in the subject is detected based on the expression level. Specifically, detection is performed by comparing the measured expression level of the target gene or its expression product of the present invention with a cutoff value (reference value). When analyzing the expression levels of multiple target genes by sequencing, it is preferable to use the following as indicators, as described above: the read count value, which is the expression level data; the RPM value, which is the read count value corrected for the difference in the total number of reads between samples; the value obtained by converting the RPM value to a base-2 logarithm (Log2RPM value) or the base-2 logarithm with an integer 1 added (Log2(RPM+1) value); or the count value corrected using DESeq2 (Normalized count value) or the base-2 logarithm with an integer 1 added (Log2(count+1) value). Alternatively, values ​​calculated using common RNA-seq quantitative methods such as fragments per kilobase of exon per million reads mapped (FPKM), reads per kilobase of exon per million reads mapped (RPKM), or transcripts per million (TPM) may be used. Furthermore, signal values ​​obtained by microarray methods and their corrected values ​​may also be used. Furthermore, when analyzing the expression level of only specific target genes using methods such as RT-PCR, it is preferable to either convert the expression level of the target gene to a relative expression level based on the expression level of housekeeping genes (relative quantification) for analysis, or to quantify the absolute copy number using a plasmid containing the region of the target gene (absolute quantification) for analysis. The copy number obtained by digital PCR may also be used.

[0046] Here, the "cutoff value" ("reference value") can be predetermined based on, for example, the relationship between the severity of PMS (e.g., DRSP score) and the expression level of the target gene or its expression product of the present invention. For example, a population can be divided into groups with different severity levels based on the DRSP score (e.g., asymptomatic, mild, moderate, and severe), and a value determined by referring to statistical values ​​such as the mean and standard deviation of the expression level of the target gene or its expression product in each group can be determined as the cutoff value (reference value) for determining whether or not a person belongs to each group. When using multiple genes as target genes, it is preferable to determine a cutoff value (reference value) for each gene or its expression product. In the present invention, it is preferable to determine cutoff values ​​(reference values) for at least 10 genes or their expression products. Groups may be formed based on race, age, etc.

[0047] For example, if the target gene or its expression product of the present invention is a positive marker, and the expression level of the marker is higher than the cutoff value (reference value), the subject may be detected as having a high severity of PMS; otherwise, the subject may be detected as having a low severity of PMS. For example, if the expression level of the target gene or its expression product in a subject is statistically significantly higher than the cutoff value (reference value), the subject may be detected as having a high severity of PMS; otherwise, the subject may be detected as having a low severity of PMS. Also, for example, if the expression level of the target gene or its expression product in a subject is preferably 110% or more, more preferably 150% or more, and even more preferably 200% or more relative to the cutoff value (reference value), the subject may be detected as having a high severity of PMS; otherwise, the subject may be detected as having a low severity of PMS.

[0048] On the other hand, if the target gene or its expression product of the present invention is a negative marker, if the expression level of the marker is lower than the cutoff value (reference value), the subject may be detected as having a high severity of PMS; otherwise, the subject may be detected as having a low severity of PMS. For example, if the expression level of the target gene or its expression product in a subject is statistically significantly lower than the cutoff value (reference value), the subject may be detected as having a high severity of PMS; otherwise, the subject may be detected as having a low severity of PMS. Also, for example, if the expression level of the target gene or its expression product in a subject is preferably 90% or less, more preferably 75% or less, and even more preferably 50% or less of the cutoff value (reference value), the subject may be detected as having a high severity of PMS; otherwise, the subject may be detected as having a low severity of PMS.

[0049] Alternatively, the severity of PMS in a subject can be detected based on whether a certain percentage of at least 10 target genes or their expression products, for example 50% or more, preferably 70% or more, more preferably 90% or more, and even more preferably 100%, meet the above-mentioned expression level criteria.

[0050] Furthermore, by utilizing the measured expression levels (expression profiles) of the target gene or its expression product according to the present invention, a discriminant formula (predictive model) can be constructed to differentiate the severity of PMS (e.g., DRSP score), and the severity of PMS can be detected using this discriminant formula. For example, by using machine learning with measured expression levels of the target gene or its expression product from symptomatic PMS groups (e.g., mild PMS groups, severe PMS groups) and asymptomatic groups (no symptoms) as explanatory variables and the severity of PMS (e.g., DRSP score) as the dependent variable, an optimal discriminant formula (predictive model) for detecting the severity of PMS (e.g., DRSP score) can be constructed. Then, the expression level of the target gene or its expression product of the present invention is similarly measured from a biological sample taken from a subject, the obtained measurement values ​​are input into the discriminant formula (predictive model), and the result obtained from the discriminant formula (for example, the predicted value of the DRSP score) can be detected as the severity of PMS in the subject.

[0051] Furthermore, by utilizing the measured expression levels of the target gene or its expression product of the present invention, a discriminant formula (predictive model) can be constructed to separate the PMS symptomatic group (e.g., mild PMS group, severe PMS group) with different PMS severity levels from the asymptomatic group (no symptoms), and the severity of PMS can be detected using this discriminant formula. Specifically, using the measured expression levels of the target gene or its expression product from the PMS symptomatic group with different PMS severity levels and the expression levels of the target gene or its expression product from the asymptomatic group as training samples, a discriminant formula (predictive model) can be constructed to separate the PMS symptomatic group (e.g., mild PMS group, severe PMS group) with different PMS severity levels from the asymptomatic group (no symptoms), and a cutoff value (reference value) can be determined to distinguish each group with different PMS severity levels based on this discriminant formula. Then, by similarly measuring the expression level of a target gene or its expression product from a biological sample taken from the subject, substituting the obtained measurement value into the discriminant formula, and comparing the result obtained from the discriminant formula with a cutoff value (reference value), the severity of PMS in the subject can be detected.

[0052] Algorithms used in constructing discriminant formulas can be publicly known algorithms, such as those used in machine learning. Examples of machine learning algorithms include linear regression, Lasso regression (Least Absolute Shrinkage and Selection Operator Regression), Ridge regression, random forest, ElasticNet, partial least squares regression, decision tree regression, linear kernel support vector machine (SVM linear), polynomial kernel support vector machine (SVM rbf), RBF kernel support vector machine (SVM polynomial), and neural networks. By inputting validation data into the constructed prediction model and calculating predicted values, the model that best matches the predicted values ​​to the observed values ​​can be selected. For example, the Pearson product-moment correlation coefficient, a statistical indicator that indicates the degree of similarity between two values, can be calculated between the observed and predicted values, and the model with the highest degree of similarity can be selected as the optimal prediction model.

[0053] When using the Random Forest algorithm to construct a discriminant, the out-of-bounds (OOB) error rate can be calculated as an indicator of the accuracy of the predictive model (Breiman L. Machine Learning (2001) 45;5-32).

[0054] In random forests, a method called bootstrapping is used to randomly select approximately two-thirds of the total sample size, allowing for overlaps, and create a classifier called a decision tree. Samples that are not selected are called Out of Bug (OOB). By using one decision tree to predict the target variable of the OOBs and comparing it with the correct label, the error rate can be calculated (OOB error rate in the decision tree). This process is repeated 500 times, and the average of the OOB error rates across the 500 decision trees can be used as the OOB error rate of the random forest model.

[0055] The number of decision trees used to construct the random forest model (n_estimators value) is 100 by default, but this can be changed to any number as needed. Furthermore, the number of variables used to create the sample discriminant in a single decision tree (max_features value) is the square root of the number of explanatory variables by default, but this can be changed to any value from one to the total number of explanatory variables as needed.

[0056] The Python library "scikit-learn" can be used to determine the max_features value. By specifying a random forest as the machine learning algorithm in the "scikit-learn" library, eight different max_features values ​​can be tried, and for example, the max_features value that maximizes accuracy can be selected as the optimal max_features value. Note that the number of trials for the max_features value can be changed to any number of trials as needed.

[0057] Alternatively, you can use the GridSearchCV function from the "scikit-learn" library to find the optimal values ​​for hyperparameters such as n_estimators, max_features, and max_depth (a parameter that controls how deep the decision tree can branch) from all possible combinations of specified parameters using cross-validation.

[0058] When using the Random Forest algorithm to construct a discriminant, the importance of the explanatory variables used in building the model can be quantified (variable importance). For example, the Mean Decrease Gini can be used as the variable importance value.

[0059] The method for determining the cutoff value (reference value) is not particularly restricted and can be determined according to known methods. For example, it can be obtained from an ROC (Receiver Operating Characteristic Curve) curve created using a discriminant. In an ROC curve, the vertical axis plots the probability of a positive result in a positive patient (sensitivity), and the horizontal axis plots the value obtained by subtracting the probability of a negative result in a negative patient (specificity) from 1 (false positive rate). Regarding the "true positive (sensitivity)" and "false positive (1-specificity)" shown in the ROC curve, the value (Youden index) at which "true positive (sensitivity)" - "false positive (1-specificity)" is maximized can be used as the cutoff value (reference value).

[0060] The feature genes used to construct the predictive model for detecting the severity of PMS according to the present invention include a group of 30 genes shown in Tables 3 and 4 above. In the present invention, at least 10 genes selected from the 30 gene groups shown in Tables 3 and 4 are used as feature genes to construct a predictive model for detecting the severity of PMS. More preferably, the 10 genes shown in Table 5 are used as feature genes to construct a predictive model for detecting the severity of PMS.

[0061] The present invention provides a test kit for detecting the severity of PMS, which contains a test reagent for measuring the expression level of the target gene or its expression product in a biological sample isolated from a subject. Specifically, this includes reagents for nucleic acid amplification and hybridization, such as oligonucleotides (e.g., primers for PCR) that specifically bind (hybridize) to the target gene or nucleic acid derived therefrom, or reagents for immunological measurement, such as antibodies that recognize the expression product (protein) of the target gene. The oligonucleotides, antibodies, etc., included in the kit can be obtained by known methods as described above. In addition to the antibodies and nucleic acids mentioned above, the test kit may also include labeling reagents, buffer solutions, chromogenic substrates, secondary antibodies, blocking agents, equipment necessary for the test, control reagents to be used as positive and negative controls, tools for collecting biological samples (for example, if the biological sample is SSL, an oil-removing film for collecting SSL), reagents for storing the collected biological samples, and storage containers. [Examples]

[0062] Example 1: Identification of PMS severity markers 1) Subjects The study included 48 adult women (ages 20-45, BMI between 18.5 and 25.0) who had previously been confirmed to have a stable menstrual cycle of approximately 25-30 days and were not currently receiving medical treatment for any disease, including PMS. The subject group consisted of 24 individuals who had previously completed the Japanese version of the Daily Records of Severity of Problems (DRSP), a severity scale for menstrual-related symptoms, over two menstrual cycles. Based on their DRSP scores, they were classified as having PMS (6 with moderate to severe symptoms, 18 with mild symptoms) and 24 individuals who were classified as having no PMS symptoms.

[0063] 2) SSL collection During the luteal phase of the menstrual cycle (luteal phase: any day between 21 and 23 days after the start of menstruation), sebum was collected from the entire face of each subject using one oil-absorbing film (polypropylene, 5.0 cm x 8.0 cm, 3M). The oil-absorbing films were transferred to plastic tubes and stored at -80°C until used for RNA extraction.

[0064] 3) RNA preparation and sequencing The oil-absorbing film described in 2) above was cut to an appropriate size, and RNA was extracted using QIAzol Lysis Reagent (Qiagen) according to the provided protocol. The extracted RNA was reverse transcribed at 42°C for 105 minutes, and the resulting cDNA was amplified by multiplex PCR. Multiplex PCR was performed at an annealing and extension temperature of 62°C. The obtained PCR products were purified using Ampure XP (Beckman Coulter, Inc.). The purified product solution was mixed with the 5×VILO RT Reaction Mix included with the SuperScript VILO cDNA Synthesis kit (Thermo Fisher Scientific), the 5×Ion Ampliseq HiFi Mix included with the Ion AmpliSeqTranscriptome Human Gene Expression Kit (Life Technologies Japan Co., Ltd.), and the Ion AmpliSeq Transcriptome Human Gene Expression Core Panel (Thermo Fisher Scientific). Then, according to the protocols included with each kit, the primer sequences were digested, adapter ligation and purification were performed, and amplification was carried out to prepare libraries. The prepared libraries were loaded onto an Ion 540 Chip and sequenced using the Ion S5 / XL system (Life Technologies Japan Co., Ltd.). By gene mapping each read sequence obtained from sequencing against the human genome reference sequence, hg19 AmpliSeq Transcriptome ERCC v1, the genes from which each read sequence originated were determined, and expression level data for each gene were obtained.

[0065] 4) Data Analysis The RNA expression data (read count values) derived from SSL of subjects obtained in step 3) above were corrected using the DESeq2 method (Love MI et al. Genome Biol. 2014). However, samples in which 4161 or more genes were not detected were excluded, and only 6729 genes for which expression data without missing values ​​was obtained for more than 90% of the subjects in the remaining samples were used for the following analysis. For the analysis, the count values ​​corrected using the DESeq2 method (normalized count values) were used, and count values ​​of 0 were treated as missing values. Based on the SSL-derived RNA expression levels (Normalized count values) obtained above, correlation coefficients and p-values ​​were calculated using Spearman's rank correlation analysis between samples. As a result, 202 genes were identified as having an absolute correlation coefficient of 0.35 or higher and a p-value of 0.05 or lower, and thus correlated with the DRSP score: 37 genes with a positive correlation and 165 genes with a negative correlation. Furthermore, the top 15 genes with the highest absolute correlation coefficients were selected from each of these positively and negatively correlated gene groups (Tables 6 and 7). These genes were all genes that had not been directly reported to be associated with PMS in previous studies. The nucleic acids derived from these genes in Tables 6 and 7, and their translation products, can each be used as markers of PMS severity. Moreover, the nucleic acids derived from these genes in Tables 6 and 7, and their translation products, are novel PMS severity markers.

[0066] [Table 6]

[0067] [Table 7]

[0068] Example 2: Construction of a PMS severity prediction model based on luteal phase RNA expression data. 1) Dataset splitting Of the RNA expression datasets obtained in Example 1 for subjects during the luteal phase, the datasets from 38 subjects (19 subjects classified as having PMS and 19 subjects classified as asymptomatic, representing 80% of the total number of subjects analyzed) were used as training data for the PMS severity prediction model, and the datasets from the remaining 20% ​​(5 subjects each, representing 10 subjects) were used as test data to evaluate the model's accuracy.

[0069] 2) Feature Selection From the 30 genes extracted in Example 1, shown in Tables 6 and 7, which have not been reported to be associated with PMS, three sets of features were selected to be used in constructing the predictive model according to the following criteria. From the genes shown in Tables 6 and 7, 10 to 30 genes with a larger absolute correlation coefficient with the DRSP score were selected. Specifically, the expression levels of 10 genes (Nos. 1-5 in Table 6 and Nos. 1-5 in Table 7), 20 genes (Nos. 1-10 in Table 6 and Nos. 1-10 in Table 7), or 30 genes (Nos. 1-15 in Table 6 and Nos. 1-15 in Table 7) were selected as features for predicting the DRSP score.

[0070] 3) Model Construction Model construction was performed using the carnet package in the statistical analysis environment R. A predictive model was constructed using the three sets of features selected in 2) above as explanatory variables and the luteal phase DRSP score as the dependent variable. For each feature, the predictive model was trained using 11 different algorithms: linear regression, Lasso regression, Lidge regression, Random forest, ElasticNet, partial least squares regression, decision tree regression, linear kernel SVM, polynomial kernel SVM, RBF kernel SVM, and neural network, with 10x cross-validation. For each algorithm, the expression level data of SSL-derived RNA from the test data was input into the trained model to calculate the predicted value of the luteal phase DRSP score. To check the degree of agreement between the predicted and actual values, the Pearson correlation coefficient was calculated between the predicted and actual values, and the model with the largest value was selected as the optimal predictive model.

[0071] 4) Results Table 8 shows the features, optimal algorithm, and correlation coefficients used to predict the DRSP score. Based on luteal phase expression data of genes related to PMS severity, or their principal components, the DRSP score could be predicted with a high correlation coefficient. Therefore, it was demonstrated that DRSP score prediction is possible using luteal phase expression data of SSL-derived RNA.

[0072] [Table 8]

Claims

1. A method for detecting the severity of premenstrual syndrome in a subject, comprising measuring the expression levels of at least 10 genes or their expression products selected from the 30 gene groups shown in Tables 1 and 2 below, in a biological sample taken from the subject. Table 1 Table 2

2. The method according to claim 1, comprising detecting the severity of premenstrual syndrome in the subject based on the expression level of the gene or its expression product.

3. The method according to claim 1, wherein the at least 10 genes are the genes shown in Table 3 below. Table 3

4. The method according to claim 1, wherein the subject is in the luteal phase of the menstrual cycle.

5. This includes detecting the severity of premenstrual syndrome in a subject using a predictive model based on the expression level of the gene or its expression product, The method according to claim 1, wherein the predictive model is constructed with measured values ​​of the expression levels of at least 10 genes or their expression products in symptomatic and asymptomatic groups of premenstrual syndrome as explanatory variables, and the severity of premenstrual syndrome as the dependent variable.

6. The method according to claim 1, wherein the expression level of the gene or its expression product is the expression level of mRNA.

7. The method according to claim 1, wherein the biological sample is lipids on the surface of the subject's skin.

8. The method according to claim 1, wherein the severity of premenstrual syndrome is a severity scale based on the Daily Records of Severity of Problems (DRSP) evaluation score.

9. A test kit for detecting the severity of premenstrual syndrome, used in the method according to any one of claims 1 to 8, comprising an oligonucleotide that specifically hybridizes with the gene or nucleic acid derived therefrom, or an antibody that recognizes the expression product of the gene.

10. A detection marker for detecting the severity of premenstrual syndrome, comprising at least 10 genes or their expression products selected from the 30 gene groups shown in Tables 4 and 5 below. Table 4 Table 5