Method for detecting the severity of premenstrual syndrome (PMS)
The method addresses the invasive and subjective nature of PMS detection by using specific gene expression markers in skin surface lipids to objectively assess PMS severity, providing earlier and non-invasive evaluation of symptom intensity.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- KAO CORP
- Filing Date
- 2021-04-23
- Publication Date
- 2026-04-14
AI Technical Summary
Existing methods for detecting the severity of premenstrual syndrome (PMS) are invasive, subjective, and lack objective markers for evaluating symptom severity, particularly before the onset of menstruation.
A method utilizing the expression levels of specific genes (SH3BGRL3, NDUFA11, PSME2, SNORA24, SSR4, ZNF91, SNORA70, HSPA6, ALDH2, BRK1, C20orf111, CSNK1A1, STK10, GHITM, RPL21, MED13, RPL32, UBE2R2, RPS3, PLEKHM2, DDX6, ALDH1A3, TSPYL1, NUB1, KRT6A, CCL22, ELOVL3, TMEM66, GDI2, MYL12B, CCNI, PSMC6, RABGEF1, ARF1, ACTN1, CHMP2A, GAK, STX3, SERP1, and GAS7) and their translation products, measured from skin surface lipids (SSL), to objectively assess PMS severity.
Enables non-invasive, objective detection of PMS severity by analyzing gene expression in SSL, allowing for earlier identification of symptom severity without burdening the subject, using novel markers that include previously unreported genes for PMS association.
Smart Images

Figure 0007845827000001 
Figure 0007845827000002 
Figure 0007845827000003
Abstract
Description
Technical Field
[0001] The present invention relates to a method for detecting the severity of premenstrual syndrome.
Background Art
[0002] Physical and mental discomfort in women before and / or during menstruation is called menstrual-related symptoms, and is represented by premenstrual syndrome (Premenstrual Syndrome; hereinafter referred to as PMS) and dysmenorrhea. Since it has been pointed out that PMS is also related to the work productivity of women, alleviating or reducing menstrual-related symptoms represented by PMS and dysmenorrhea and improving the QOL (Quality of Life) of women is not only health-related but also socially important.
[0003] Techniques have been developed to examine the current and even future physiological states in the human body by analyzing nucleic acids such as DNA and RNA in biological samples. Analysis using nucleic acids has the advantages that a comprehensive analysis method has been established and abundant information can be obtained by one analysis, and functional linking of the analysis results is easy based on many research reports on single nucleotide polymorphisms and RNA functions. Nucleic acids derived from living organisms can be extracted from tissues such as blood, body fluids, secretions, etc. Patent Document 1 describes using RNA contained in skin surface lipids (SSL) as a sample for biological analysis, and detecting marker genes of the epidermis, sweat glands, hair follicles, and sebaceous glands from SSL.
[0004] There are genes whose expression changes with the production of estrogen or progesterone, which are female hormones. For example, so far, AQP3 (aquaporin 3), HIF-1α (hypoxia inducible factor 1 subunit α), MFN2 (mitofusin 2), DEFB4A (defensin beta 4A, or hBD-2), hBD-3, etc. have been reported (Non-Patent Documents 1 to 3, Patent Document 2).
Prior Art Documents
[0005] [Patent Document 1] International Public Gazette No. 2018 / 008319 [Patent Document 2] Japanese Patent Publication No. 2015-228829 [Non-patent literature]
[0006] [Non-Patent Document 1] Hum Reprod, 2018, 33(11):2060-2073 [Non-Patent Document 2] Biol Reprod, 2018, 99(2):308-318 [Non-Patent Document 3] Mol Cell Endocrinol, 2013, 371:79-86 [Overview of the Initiative] [Problems that the invention aims to solve]
[0007] The present invention relates to a PMS severity marker that enables the detection of PMS severity, and a method for detecting PMS severity using the marker. [Means for solving the problem]
[0008] In other words, the present invention provides a method for detecting the severity of PMS in a subject, comprising measuring the expression of at least one selected from the group consisting of the following genes: SH3BGRL3, NDUFA11, PSME2, SNORA24, SSR4, ZNF91, SNORA70, HSPA6, ALDH2, BRK1, C20orf111, CSNK1A1, STK10, GHITM, RPL21, MED13, RPL32, UBE2R2, RPS3, PLEKHM2, DDX6, ALDH1A3, TSPYL1, NUB1, KRT6A, CCL22, ELOVL3, TMEM66, GDI2, MYL12B, CCNI, PSMC6, RABGEF1, ARF1, ACTN1, CHMP2A, GAK, STX3, SERP1, and GAS7, as well as the translation products of said genes. The present invention also provides a marker for the severity of PMS, comprising at least one selected from the group consisting of nucleic acids derived from the following genes: SH3BGRL3, NDUFA11, PSME2, SNORA24, SSR4, ZNF91, SNORA70, HSPA6, ALDH2, BRK1, C20orf111, CSNK1A1, STK10, GHITM, RPL21, MED13, RPL32, UBE2R2, RPS3, PLEKHM2, DDX6, ALDH1A3, TSPYL1, NUB1, KRT6A, CCL22, ELOVL3, TMEM66, GDI2, MYL12B, CCNI, PSMC6, RABGEF1, ARF1, ACTN1, CHMP2A, GAK, STX3, SERP1, and GAS7, as well as the translation products of said genes. [Effects of the Invention]
[0009] This invention provides a method for detecting the severity of PMS using a novel PMS severity marker. Since the PMS severity marker used in this invention can be collected from surface lipids (SSL) of the skin, it can be obtained simply and non-invasively. Therefore, according to this invention, it is possible to detect the severity of PMS without burdening the subject. [Modes for carrying out the invention]
[0010] All patent, non-patent, and other publications cited herein are incorporated herein by reference in their entirety.
[0011] Premenstrual syndrome (PMS) is a group of physical or mental symptoms that begin about a week before menstruation (during the luteal phase), first reported in 1931 by Frank (Archives of Neurology and Psychiatry, 1931, 26:1053). Medically, PMS is defined as "physical and mental symptoms that begin 3 to 10 days before the start of menstruation and decrease or disappear with the onset of menstruation." Typical symptoms of PMS include, as physical symptoms, lower abdominal pain, back pain, bloating, headache, stiff shoulders, cold hands and feet, increased appetite, diarrhea, constipation, swelling, breast pain, breast tenderness, acne, skin problems, fatigue, and drowsiness, while mental symptoms include irritability, anger, aggression, depression, tearfulness, and increased anxiety.
[0012] In this specification, severe PMS refers to a condition in which the physical or mental symptoms of PMS are severe, and mild PMS refers to a condition in which the physical or mental symptoms of PMS are mild. Traditionally, the severity of menstrual symptoms such as PMS has been evaluated based on scores from questionnaires such as the Menstrual Distress Questionnaire (MDQ). The higher the MDQ score during the luteal phase, the more severe the PMS is considered to be. For example, according to the MDQ described in Psychological Measurement Scales III (Science Co., Ltd., published in 2001), a luteal phase MDQ score of 66 or higher can be evaluated as severe PMS, and a score of 48 or lower can be evaluated as mild PMS.
[0013] The present invention enables the more simple and objective detection of the severity of PMS, which conventionally had to be evaluated based on the interview of subjects during the luteal phase as described above, based on the expression of a gene or its translation product. Further, the present invention enables the prior detection of the severity of PMS based on the expression of a gene or its translation product during the follicular phase prior to the luteal phase in which PMS occurs.
[0014] In one aspect, the present invention provides a PMS severity marker. As shown in the examples described later, the present inventors have found that the expression of 51 genes shown in Table 1 below has a correlation with the severity of PMS. Further, 40 genes shown in Table 2 among these 51 genes were genes for which no association with PMS had been reported conventionally. Therefore, the nucleic acid derived from the gene shown in Table 1 and the translation product of the gene can each be used as a marker for the severity of PMS. Further, the nucleic acid derived from the gene shown in Table 2 and the translation product of the gene are novel PMS severity markers.
[0015]
Table 1
[0016]
Table 2
[0017] The names of the genes disclosed in the present specification follow the Official Symbol described in NCBI ([www.ncbi.nlm.nih.gov / ]). In the present specification, "nucleic acid derived from a gene" includes mRNA transcribed from the gene, cDNA prepared from the mRNA, reaction products from the cDNA (for example, PCR products, cloned DNA, etc.), and the like.
[0018] In the present invention, it is possible to detect the severity of PMS in a subject based on the expression level of each gene shown in Table 1 or its translation product. In the present invention, the expression level of the gene shown in Table 1 or its translation product can be measured by quantifying the nucleic acid derived from the gene or the translation product of the gene. For example, the expression level of the gene or its translation product can be measured by quantifying mRNA transcribed from the gene, cDNA prepared from the mRNA, the reaction product from the cDNA, or the translation product (such as a protein) synthesized by translation of the mRNA.
[0019] The nucleic acid derived from the gene shown in Table 1 or the translation product of the gene can be prepared from a biological sample collected from a subject, such as cells, body fluids, secretions, etc., according to a conventional method. Preferably, the nucleic acid is mRNA prepared from the skin surface lipids (SSL) of the subject. More preferably, the translation product is a translation product prepared from the SSL of the subject.
[0020] In the present invention, the "detection" of the severity of premenstrual syndrome can also be expressed in terms such as prediction, examination, measurement, determination, or evaluation support. Note that the terms "detection", "prediction", "examination", "measurement", "determination", or "evaluation" of the severity of premenstrual syndrome in the present invention do not include the diagnosis of the severity of premenstrual syndrome by a doctor.
[0021] In the present specification, "skin surface lipids (SSL)" refers to the lipid-soluble fraction present on the surface of the skin, and is sometimes called sebum. Generally, SSL mainly contains secretions secreted from exocrine glands such as sebaceous glands in the skin and exists on the skin surface in the form of a thin layer covering the skin surface. SSL contains RNA expressed in skin cells (see Patent Document 1). Also in the present specification, "skin" is a general term for a region including the epidermis, dermis, hair follicles of the body surface, and tissues such as sweat glands, sebaceous glands, and other glands, unless otherwise particularly limited.
[0022] Any means used for the recovery or removal of SSL from the skin can be employed to collect SSL from the subject's skin. Preferably, SSL absorbent materials, SSL adhesive materials, or instruments for scraping SSL off the skin, as described later, can be used. The SSL absorbent material or SSL adhesive material is not particularly limited as long as it is a material that has an affinity for SSL, and examples include polypropylene and pulp. More detailed examples of procedures for collecting SSL from the skin include methods of absorbing SSL onto a sheet material such as oil-blotting paper or oil-blotting film, methods of adhering SSL to a glass plate or tape, and methods of scraping off and recovering SSL with a spatula, scraper, etc. To improve the adsorption of SSL, an SSL absorbent material containing a highly lipid-soluble solvent beforehand may be used. On the other hand, since the adsorption of SSL is inhibited if the SSL absorbent material contains a highly water-soluble solvent or water, it is preferable that the content of highly water-soluble solvents or water is low. It is preferable to use the SSL absorbent material in a dry state. The skin from which SSL is collected is not particularly limited and can be any part of the body, such as the head, face, neck, trunk, hands, or feet. Areas with high sebum secretion, such as the skin of the face, are preferred.
[0023] SSL collected from a subject may be stored for a certain period of time. To minimize the degradation of the contained RNA, it is preferable to store the collected SSL under low temperature conditions as quickly as possible after collection. The storage temperature conditions for the RNA-containing SSL in this invention may be 0°C or lower, preferably -20±20°C to -80±20°C, more preferably -20±10°C to -80±10°C, even more preferably -20±20°C to -40±20°C, even more preferably -20±10°C to -40±10°C, even more preferably -20±10°C, and even more preferably -20±5°C. The storage period for the RNA-containing SSL under these low temperature conditions is not particularly limited, but is preferably 12 months or less, for example, 6 hours to 12 months, more preferably 6 months or less, for example, 1 day to 6 months, and even more preferably 3 months or less, for example, 3 days to 3 months.
[0024] For RNA extraction from collected SSL, methods commonly used for RNA extraction or purification from biological samples can be used, such as the phenol / chloroform method, the AGPC (acid guanidinium thiocyanate-phenol-chloroform extraction) method, or methods using columns such as TRIzol®, RNeasy®, or QIAzol®, or methods using special magnetic particles coated with silica, or methods using Solid Phase Reversible Immobilization magnetic particles, or extraction using commercially available RNA extraction reagents such as ISOGEN. For the extraction of translation products from SSL, methods commonly used for protein extraction or purification from biological samples can be used, such as commercially available protein extraction reagents such as QIAzol Lysis Reagent (Qiagen).
[0025] The severity of PMS in a subject can be detected by examining the expression of one or more genes or their translation products shown in Table 1 using the nucleic acids or translation products prepared by the above procedure. Therefore, in another embodiment, the present invention provides a method for detecting the severity of PMS in a subject, comprising measuring the expression of at least one selected from the group consisting of genes or their translation products shown in Table 1.
[0026] The subject in the method for detecting the severity of PMS according to the present invention (hereinafter referred to as the "method of the present invention") may be, for example, a female human or non-human mammal who requires detection of the severity of PMS. Preferably, the subject is an animal with SSL on its skin, and more preferably a human.
[0027] In one embodiment, the method of the present invention includes measuring the expression of at least one gene selected from the group of genes shown in Table 1 (hereinafter also referred to as the target gene(s)) in a subject. In another embodiment, the method of the present invention includes measuring the expression of at least one gene selected from the group of translation products of the genes shown in Table 1 (hereinafter also referred to as the target product(s)) in a subject. Alternatively, both the expression of the target gene(s) and the expression of the target product(s) may be measured. Preferably, the expression of the target gene(s) is measured.
[0028] Among the genes shown in Table 1, those shown in Table 2 are genes that have not been previously reported to be associated with PMS. Therefore, in a preferred embodiment of the method of the present invention, the target gene(s) is at least one selected from the group consisting of the genes shown in Table 2, and the target product(s) is at least one selected from the group consisting of the translation products of the genes shown in Table 2. Furthermore, in the method of the present invention, a combination of at least one selected from the group consisting of the genes or translation products shown in Table 2 and at least one selected from the group consisting of other genes (i.e., CTNNB1, AQP3, ANXA1, ARHGDIB, DEFB4A, BIRC3, S100A7, SFN, HMGN2, HIF1A, and MFN2 listed in Table 1) or their translation products may be used as the target gene group or target product group, and their expression may be measured.
[0029] Preferably, the measurement of the expression of the target gene(s) or target product(s) in the method of the present invention is performed using RNA or a translation product contained in the subject's SSL. In one embodiment, the method of the present invention may further include collecting the subject's SSL. In one embodiment, the method of the present invention may further include extracting RNA or a translation product from the SSL collected from the subject.
[0030] The expression level of a target gene can be measured by quantifying the amount of RNA derived from the target gene contained in a sample derived from the subject, for example, RNA extracted from SSL. The amount of subject-derived RNA can be measured according to the procedure for RNA-based gene expression analysis commonly used in this field. Examples of RNA-based gene expression analysis methods include converting RNA to cDNA by reverse transcription, and then quantifying the cDNA or its amplification product using real-time PCR, multiplex PCR, microarrays, sequencing, chromatography, etc. The expression level of the translation product can be measured using protein quantification methods commonly used in this field, such as ELISA, immunostaining, fluorescence methods, electrophoresis, chromatography, and mass spectrometry. The measured expression level may be the absolute amount of expression of the target gene(s) or target product(s) in the biological sample, or it may be the relative expression level relative to other standard substances or the total expression level of the gene or translation product.
[0031] In a preferred embodiment, the method of the present invention involves converting the RNA derived from the test subject to cDNA by reverse transcription, then subjecting the cDNA to PCR, and purifying the resulting reaction product. While primers targeting specific RNAs to be analyzed may be used for the reverse transcription, random primers are preferable for more comprehensive nucleic acid preservation and analysis. General reverse transcriptases or reverse transcription reagent kits can be used for the reverse transcription. Preferably, highly accurate and efficient reverse transcriptases or reverse transcription reagent kits are used, such as M-MLV Reverse Transcriptase and its variants, or commercially available reverse transcriptases or reverse transcription reagent kits, for example, PrimeScript® Reverse Transcriptase series (Takara Bio Inc.), SuperScript® Reverse Transcriptase series (Thermo Scientific Inc.), etc. SuperScript® III Reverse Transcriptase and SuperScript® VILO cDNA Synthesis kit (both from Thermo Scientific Inc.) are preferred. In the PCR of the obtained cDNA, only the specific DNA to be analyzed may be amplified using a primer pair that targets that specific DNA, or multiple DNAs may be amplified using multiple primer pairs. Preferably, the PCR is multiplex PCR. Multiplex PCR is a method of simultaneously amplifying multiple gene regions by using multiple primer pairs simultaneously in the PCR reaction system. Multiplex PCR can be performed using commercially available kits (for example, the Ion AmpliSeqTranscriptome Human Gene Expression Kit; Life Technologies Japan Co., Ltd., etc.).
[0032] By adjusting the reaction conditions for reverse transcription and PCR, the yield of the PCR reaction product can be further improved, and consequently, the accuracy of the analysis using it can be further improved. Preferably, in the extension reaction of the reverse transcription, the temperature is adjusted to preferably 42°C ± 1°C, more preferably 42°C ± 0.5°C, and even more preferably 42°C ± 0.25°C, while the reaction time is adjusted to preferably 60 minutes or more, more preferably 80 to 120 minutes. Preferably, the temperature of the annealing and extension reactions in the PCR can be appropriately adjusted depending on the primer used, but for example, when using the above-mentioned multiplex PCR kit, it is preferably 62°C ± 1°C, more preferably 62°C ± 0.5°C, and even more preferably 62°C ± 0.25°C. Therefore, in the PCR, the annealing and extension reactions are preferably performed in one step. The time of the annealing and extension reaction steps can be adjusted depending on the size of the DNA to be amplified, but is preferably 14 to 18 minutes. The conditions for the denaturation reaction in PCR can be adjusted depending on the DNA to be amplified, but are preferably 95-99°C for 10-60 seconds. Reverse transcription and PCR at the above temperatures and times can be performed using a thermal cycler commonly used for PCR.
[0033] The purification of the reaction product obtained by the PCR is preferably carried out by size separation of the reaction product. Size separation allows the target PCR reaction product to be separated from primers and other impurities contained in the PCR reaction mixture. DNA size separation can be carried out, for example, by a size separation column, a size separation chip, or magnetic beads that can be used for size separation. Preferred examples of magnetic beads that can be used for size separation include Solid Phase Reversible Immobilization (SPRI) magnetic beads such as Ampure XP.
[0034] The purified PCR reaction product may be subjected to further processing necessary for subsequent quantitative analysis. For example, the purified PCR reaction product may be prepared into a suitable buffer solution for DNA sequencing, the PCR primer regions in the PCR-amplified DNA may be cleaved, or adapter sequences may be further added to the amplified DNA. For example, the purified PCR reaction product can be prepared into a buffer solution, the amplified DNA can be subjected to removal of PCR primer sequences and adapter ligation, and the resulting reaction product can be amplified as needed to prepare a library for quantitative analysis. These operations can be performed, for example, using the 5×VILO RT Reaction Mix included with the SuperScript® VILO cDNA Synthesis kit (Life Technologies Japan Co., Ltd.), the 5×Ion AmpliSeq HiFi Mix included with the Ion AmpliSeq Transcriptome Human Gene Expression Kit (Life Technologies Japan Co., Ltd.), and the Ion AmpliSeq Transcriptome Human Gene Expression Core Panel, according to the protocol included with each kit. For sequencing, a next-generation sequencer (e.g., Ion S5 / XL system, Life Technologies Japan Co., Ltd.) is preferably used. RNA expression can be quantified based on the number of reads generated by sequencing (read count).
[0035] Preferably, the method of the present invention detects the severity of PMS in a subject based on the expression level of the target gene(s) or target product(s). For example, the genes shown in Table 1 can be divided into two groups, as shown in Tables 3 and 4, based on the pattern of changes in their expression levels associated with the severity of PMS. The severity of PMS in a subject can be detected by utilizing these patterns of change.
[0036] The genes shown in Table 3 are those whose expression increases when PMS is severe and decreases when PMS is mild. On the other hand, the genes shown in Table 4 are those whose expression decreases when PMS is severe and increases when PMS is mild. By using the expression levels of one or more of the genes or their translation products shown in Tables 3 and 4 as a basis, the severity of a subject's PMS (e.g., whether it is severe or not) can be detected. Of the genes shown in Tables 3 and 4, the 22 and 18 genes underlined, respectively, are genes that have not been previously reported to be associated with PMS (genes shown in Table 2).
[0037] [Table 3]
[0038] [Table 4]
[0039] In one embodiment of the method of the present invention, the expression levels of individual target genes(s) or target products(s) are directly used as parameters for detecting the severity of PMS. For example, the severity of PMS in a subject can be detected by regularly (e.g., monthly) measuring the expression levels of the target gene or target product in a subject at a specific time in the menstrual cycle (e.g., the follicular phase or the luteal phase, hereinafter the same) and tracking the fluctuations in these expression levels. Alternatively, the severity of PMS in a subject can be detected by comparing the expression level of the target gene or target product measured from the subject at a specific time in the menstrual cycle with the baseline expression level of the target gene or target product for the same time in the menstrual cycle. The baseline expression level of the target gene or target product can be a predetermined value, for example, a statistical value (e.g., average expression level) of the expression level of the target gene or target product at a specific time in the menstrual cycle. The baseline expression level may be set for each subject. For example, it may be calculated from statistical values (e.g., average values) of the expression levels of the target gene or target product at specific times in the menstrual cycle measured multiple times from the same subject. Alternatively, the baseline expression level may be a statistical value (e.g., an average value) of the expression level of the target gene or target product at a specific time in the menstrual cycle measured from a population. When using multiple genes or translation products as the target gene group or target product group, it is preferable to determine the baseline expression level for each of the genes or translation products.
[0040] For example, if the target gene or target product is one of the genes or its translation product shown in Table 3, a higher expression level may indicate severe PMS in the subject, while a lower expression level may indicate that PMS is not severe (or mild). On the other hand, if the target gene or target product is one of the genes or its translation product shown in Table 4, a lower expression level may indicate severe PMS in the subject, while a higher expression level may indicate that PMS is not severe (or mild).
[0041] Alternatively, two or more genes or their translation products shown in Tables 3-4 may be combined and used as target genes or target products. When multiple genes or translation products are used, the severity of PMS in the subject can be detected based on whether a certain percentage, for example, 50% or more, preferably 70% or more, more preferably 90% or more, and even more preferably 100%, of the genes or translation products meet the expression level criteria described above. For example, if the expression level of 50% or more of the target gene group or target product group selected from Table 3 is higher than the standard expression level, the subject's PMS may be detected as severe. Alternatively, if the expression level of 50% or more of the target gene group or target product group selected from Table 4 is lower than the standard expression level, the subject's PMS may be detected as severe. Alternatively, if the expression level of 50% or more of the target gene group or target product group selected from Table 3 is higher than the standard expression level, and the expression level of 50% or more of the target gene group or target product group selected from Table 4 is lower than the standard expression level, the subject's PMS may be detected as severe. However, the combination of target gene groups or target product groups in the method of the present invention is not limited to the examples above.
[0042] In another embodiment, the severity of PMS in a subject is detected based on a predictive model constructed using expression data of a target gene(s) or target product(s). The expression data can include data on the expression levels of the target gene(s) or target product(s) at specific stages of the menstrual cycle (e.g., follicular phase or luteal phase). For example, an optimal predictive model for detecting PMS severity can be constructed using machine learning, with expression data for one or more target genes(s) or target products(s) as explanatory variables and PMS severity (e.g., severe or mild) as the dependent variable. For example, multiple genes or their translation products showing a large difference in expression between the severe and mild PMS groups can be selected from those shown in Table 1 or Table 2, and their expression data can be used as explanatory variables.
[0043] Alternatively, when using expression data of multiple genes or translation products to construct a predictive model, the data may be compressed by dimensionality reduction as needed before constructing the predictive model. For example, multiple genes can be selected from the genes shown in Table 1 or Table 2, and these genes or their translation products can be extracted as a group of target genes or target products. Then, principal component analysis can be performed on the expression levels of the extracted group of target genes or target products. Using one or more principal components calculated by the principal component analysis as explanatory variables and the severity of PMS (e.g., severe or mild) as the dependent variable, an optimal predictive model for detecting the severity of PMS can be constructed using machine learning.
[0044] For constructing the predictive model, it is preferable to select one or more target genes or target products from among the genes or their translation products shown in Table 1 that show a greater difference in expression between the severe and mild groups in PMS severity. For example, one or more genes or their translation products selected from the 40 genes shown in Table 2, which have not been previously reported to be associated with PMS, are preferred target genes or target products to be used in constructing the predictive model. The same applies to target genes or target products used in principal component analysis.
[0045] In one embodiment, approximately 5 to 20 genes, preferably 5 to 15 genes, are selected from the genes shown in Table 2 that show a larger difference in expression between the severe and mild PMS groups during the luteal phase (for example, genes No. 1 to 20 in Table 2). More specifically, for example, 5 genes No. 1 to 5, 10 genes No. 1 to 10, or 15 genes No. 1 to 15 are selected from Table 2, preferably 10 genes No. 1 to 10 or 15 genes No. 1 to 15. A predictive model can be constructed using data on the expression levels of the selected genes during the luteal phase as explanatory variables. In another embodiment, approximately 5 to 20 genes, preferably 5 to 15 genes, are selected from the genes shown in Table 2 that show a larger difference in expression between the severe and mild PMS groups during the luteal phase (for example, genes No. 1 to 20 in Table 2). More specifically, for example, five genes No. 1 to 5, ten genes No. 1 to 10, or fifteen genes No. 1 to 15 are selected from Table 2, preferably five genes No. 1 to 5. A predictive model can be constructed using data on the expression levels of the selected genes during the follicular phase as explanatory variables.
[0046] When data compression is performed using principal component analysis, the dimensions and number of principal components used to construct the predictive model are not particularly limited, but it is preferable to use one or more principal components as explanatory variables that include components with a large contribution rate, for example, a first principal component, and whose total contribution rate is preferably 20% or more, more preferably 50% or more. In one embodiment, principal component analysis is performed on the expression data of target gene groups during the luteal phase, and principal components are selected such that they include a first principal component and have a total contribution rate of 20% or more, preferably 50% or more, for example, preferably 22-55% or more, more preferably 50-70%, and even more preferably 55-65%, and a predictive model can be constructed using the selected principal components as explanatory variables. In another embodiment, principal component analysis is performed on the expression data of target gene groups during the follicular phase, and principal components are selected such that they include a first principal component and have a total contribution rate of 80% or more, preferably 85% or more, more preferably 90% or more, and a predictive model can be constructed using the selected principal components as explanatory variables.
[0047] Algorithms used in building predictive models can be publicly known, such as those used in machine learning. Examples of machine learning algorithms include Random Forest, Neural Network, Support Vector Machine (SVM (linear)) with a linear kernel, and Support Vector Machine (SVM (rbf)) with an rbf kernel. By inputting validation data into the constructed predictive model and calculating predicted values, the model whose predicted values best match the actual values—for example, the model with a high accuracy rate, which indicates the degree of agreement between predicted and actual values—can be selected as the optimal predictive model.
[0048] By inputting the expression data (or principal components calculated from them) of the target gene(s) or target product(s) in the subject into the constructed predictive model, the severity of PMS in the subject can be detected.
[0049] Exemplary embodiments of the present invention are further disclosed herein, including the following substances, manufacturing methods, uses, and methods. However, the present invention is not limited to these embodiments.
[0050] [1] The genes listed below: SH3BGRL3, NDUFA11, CTNNB1, AQP3, PSME2, SNORA24, SSR4, ZNF91, ANXA1, SNORA70, HSPA6, ALDH2, ARHGDIB, DEFB4A, BRK1, C20orf111, CSNK1A1, BIRC3, S100A7, STK10, GHITM, SFN, RPL21, MED13, RPL32, UBE2R2, RPS3, PLEKHM2, DDX6 A method for detecting the severity of premenstrual syndrome in a subject, comprising measuring the expression of at least one selected from the group consisting of ALDH1A3, TSPYL1, NUB1, KRT6A, CCL22, ELOVL3, TMEM66, GDI2, MYL12B, CCNI, PSMC6, HMGN2, HIF1A, RABGEF1, ARF1, ACTN1, MFN2, CHMP2A, GAK, STX3, SERP1, and GAS7, as well as the translation products of said genes. [2] Preferably, the method according to [1], comprising measuring the expression of at least one selected from the group consisting of the following genes: SH3BGRL3, NDUFA11, PSME2, SNORA24, SSR4, ZNF91, SNORA70, HSPA6, ALDH2, BRK1, C20orf111, CSNK1A1, STK10, GHITM, RPL21, MED13, RPL32, UBE2R2, RPS3, PLEKHM2, DDX6, ALDH1A3, TSPYL1, NUB1, KRT6A, CCL22, ELOVL3, TMEM66, GDI2, MYL12B, CCNI, PSMC6, RABGEF1, ARF1, ACTN1, CHMP2A, GAK, STX3, SERP1, and GAS7, and the translation products of said genes. [3] Preferably, the method according to [2] further comprises measuring the expression of at least one selected from the group consisting of the following genes: CTNNB1, AQP3, ANXA1, ARHGDIB, DEFB4A, BIRC3, S100A7, SFN, HMGN2, HIF1A, and MFN2, and the translation products of said genes. [4] Preferably, the method includes detecting the severity of premenstrual syndrome based on the expression level of the gene or translation product, More preferably, the severity of premenstrual syndrome is detected based on the expression level of the gene, The method described in any one of items [1] to [3]. [5] Preferably, the method according to [4], comprising detecting the severity of premenstrual syndrome based on the expression level of the gene or translation product during the luteal or follicular phase. [6] Preferably, the method includes detecting the severity of premenstrual syndrome using a predictive model based on the expression level of the gene or translation product, The predictive model is constructed with the expression levels of the gene or its translation product in severe and mild cases of premenstrual syndrome (PMS), or at least one principal component calculated by principal component analysis of the expression levels, as explanatory variables, and the severity of PMS as the dependent variable. The method described in any one of items [1] to [5]. [7] Preferably, the method according to any one of [1] to [6], wherein the expression level of the gene is measured by quantifying the mRNA of the gene contained in the lipids on the surface of the skin of the subject. [8] Nucleic acids derived from the following genes: SH3BGRL3, NDUFA11, CTNNB1, AQP3, PSME2, SNORA24, SSR4, ZNF91, ANXA1, SNORA70, HSPA6, ALDH2, ARHGDIB, DEFB4A, BRK1, C20orf111, CSNK1A1, BIRC3, S100A7, STK10, GHITM, SFN, RPL21, MED13, RPL32, UBE2R2, RPS3, PLE Use of at least one selected from the group consisting of KHM2, DDX6, ALDH1A3, TSPYL1, NUB1, KRT6A, CCL22, ELOVL3, TMEM66, GDI2, MYL12B, CCNI, PSMC6, HMGN2, HIF1A, RABGEF1, ARF1, ACTN1, MFN2, CHMP2A, GAK, STX3, SERP1, and GAS7, as well as the translation products of said genes, as a severity marker for premenstrual syndrome. [9] Preferably, the use according to [8] is used, wherein at least one selected from the group consisting of nucleic acids derived from the following genes: SH3BGRL3, NDUFA11, PSME2, SNORA24, SSR4, ZNF91, SNORA70, HSPA6, ALDH2, BRK1, C20orf111, CSNK1A1, STK10, GHITM, RPL21, MED13, RPL32, UBE2R2, RPS3, PLEKHM2, DDX6, ALDH1A3, TSPYL1, NUB1, KRT6A, CCL22, ELOVL3, TMEM66, GDI2, MYL12B, CCNI, PSMC6, RABGEF1, ARF1, ACTN1, CHMP2A, GAK, STX3, SERP1, and GAS7, and the translation products of said genes.
[10] Preferably, the use according to [9] further includes at least one selected from the group consisting of nucleic acids derived from the following genes: CTNNB1, AQP3, ANXA1, ARHGDIB, DEFB4A, BIRC3, S100A7, SFN, HMGN2, HIF1A, and MFN2, and the translation products of said genes.
[11] Preferably, the use according to any one of [8] to
[10] , wherein the nucleic acid derived from the gene is the mRNA of the gene contained in the lipids on the skin surface.
[12] A marker of the severity of premenstrual syndrome comprising at least one selected from the group consisting of nucleic acids derived from the following genes: SH3BGRL3, NDUFA11, PSME2, SNORA24, SSR4, ZNF91, SNORA70, HSPA6, ALDH2, BRK1, C20orf111, CSNK1A1, STK10, GHITM, RPL21, MED13, RPL32, UBE2R2, RPS3, PLEKHM2, DDX6, ALDH1A3, TSPYL1, NUB1, KRT6A, CCL22, ELOVL3, TMEM66, GDI2, MYL12B, CCNI, PSMC6, RABGEF1, ARF1, ACTN1, CHMP2A, GAK, STX3, SERP1, and GAS7, and the translation products of said genes.
[13] Preferably, the marker according to
[12] further comprises at least one selected from the group consisting of nucleic acids derived from the following genes: CTNNB1, AQP3, ANXA1, ARHGDIB, DEFB4A, BIRC3, S100A7, SFN, HMGN2, HIF1A, and MFN2, and the translation products of said genes.
[14] Preferably, the marker according to
[12] or
[13] , wherein the nucleic acid derived from the gene is the mRNA of the gene contained in the lipids on the skin surface. [Examples]
[0051] The present invention will be described in more detail below based on examples, but the present invention is not limited thereto.
[0052] Example 1: Identification of PMS severity markers 1) SSL collection Thirty-eight healthy women (20-45 years old, BMI 18.5 or higher and less than 25.0) were included as subjects. The healthy subjects had previously been confirmed to have stable menstrual cycles of approximately 25-30 days. Sebum was collected from each subject's entire face using an oil-absorbing film (5.0cm x 8.0cm, 3M) at each stage of the menstrual cycle (menstrual phase: any day between 2 and 4 days after the start of menstruation; follicular phase: any day between 10 and 12 days after the start of menstruation; luteal phase: any day between 21 and 23 days after the start of menstruation). The oil-absorbing films were transferred to glass vials and stored at -80°C for approximately one month until use for RNA extraction. Furthermore, after collecting sebum samples, the severity of menstrual symptoms (dysmenorrhea, premenstrual syndrome, etc.) at each stage of the menstrual cycle (menstrual phase, follicular phase, and luteal phase) was evaluated using the Menstrual Distress Questionnaire (MDQ), described in Psychological Measurement Scales III (Science Co., Ltd., published in 2001), which is a questionnaire for evaluating menstrual-related symptoms.
[0053] 2) RNA preparation and sequencing The oil-absorbing film described in 1) above was cut to an appropriate size, and RNA was extracted using QIAzol Lysis Reagent (Qiagen) according to the provided protocol. The extracted RNA was reverse transcribed at 42°C for 105 minutes, and the resulting cDNA was amplified by multiplex PCR. Multiplex PCR was performed at an annealing and extension temperature of 62°C. The obtained PCR products were purified using Ampure XP (Beckman Coulter, Inc.). The solution of the obtained purified product was mixed with 5×VILO RT Reaction Mix included with the SuperScript VILO cDNA Synthesis kit, 5×Ion Ampliseq HiFi Mix included with the Ion AmpliSeq Transcriptome Human Gene Expression Kit (Life Technologies Japan Co., Ltd.), and the Ion AmpliSeq Transcriptome Human Gene Expression Core Panel. Then, primer sequence digestion, adapter ligation and purification, and amplification were performed according to the protocols provided with each kit to prepare libraries. The prepared library was loaded onto an Ion 540 Chip and sequenced using the Ion S5 / XL system (Life Technologies Japan Co., Ltd.). The genes from which each read sequence originated were determined by gene mapping of each read sequence to the human genome reference sequence, hg19 AmpliSeq Transcriptome ERCC v1.
[0054] 3) Grouping of subjects based on MDQ score In the Menstrual Questionnaire (MDQ), a higher score indicates a greater severity of menstrual symptoms. Therefore, the 38 subjects were divided into two groups based on their MDQ scores obtained during the luteal phase: 15 subjects with a score of 66 or higher (mean score 87.4) were classified as the severe PMS group, and 15 subjects with a score of 48 or lower (mean score 40.3) were classified as the mild PMS group. The 8 subjects in the intermediate group were excluded from the following data analysis.
[0055] 4) Data Analysis The read counts obtained from sequencing SSL-derived RNA from subjects during the follicular and luteal phases, as described in 2) above, were used as data for the expression level of each RNA, and reads with a read count of less than 10 were treated as missing values. To correct for differences in total read counts between samples, each read count was converted to an RPM (Reads per million mapped reads) value. 1884 genes for which expression level data without missing values was obtained for more than 80% of subjects were selected. For these 1884 genes, a Student's t test was performed between groups (severe PMS group and mild PMS group) based on their expression level data (RPM values) during the luteal phase when PMS occurs, and genes with p<0.05 were extracted. Missing values were imputed using Singular Value Decomposition (SVD) input. 51 genes with alternating expression levels between groups were obtained (Table 5), and of these, 40 genes shown in Table 6 were genes that had not been previously reported to be associated with PMS. The nucleic acids derived from the genes listed in Table 5, and their translation products, can each be used as markers of PMS severity. Furthermore, the nucleic acids derived from the genes shown in Table 6, among those listed in Table 5, and their translation products are novel PMS severity markers.
[0056] [Table 5]
[0057] [Table 6]
[0058] Example 2: Construction of a PMS severity prediction model based on luteal phase RNA expression data. 1) Dataset splitting From the luteal phase SSL-derived RNA expression dataset obtained in Example 1, data from 10 subjects in each group (severe PMS group and mild PMS group), representing 67% of the total number of subjects analyzed, were used as training data for the PMS severity prediction model, and data from 5 subjects in each group, representing the remaining 33%, were used as test data to evaluate the model's accuracy. In order to approximate the RPM values, which follow a negative binomial distribution, to a normal distribution, the RNA expression data (RPM values of read counts) obtained in Example 1 were converted to base-2 logarithmic RPM values (Log2RPM values).
[0059] 2) Feature Selection From the 40 genes listed in Table 6, which were extracted in Example 1 and not reported to be associated with PMS, features to be used in constructing the predictive model were selected using the following two methods. (Method 1) From the genes shown in Table 6, 5 to 15 genes with small p-value differences in expression between the severe and mild PMS groups were selected as features for detecting PMS severity. Specifically, the expression levels (Log2RPM values) of the 5 genes No. 1 to 5 in Table 6, the 10 genes No. 1 to 10 in Table 6, or the 15 genes No. 1 to 15 in Table 6 were selected. (Method 2) Principal component analysis was performed on the gene expression data shown in Table 6 using the tidyverse package in the statistical analysis environment R, and variable compression was carried out. The data with a new dimension obtained by principal component analysis (1st to 10th principal components, Table 7) were selected as features for PMS severity detection.
[0060] [Table 7]
[0061] 3) Model Construction Model construction was performed using the carnet package in the statistical analysis environment R. Using the features selected in methods 1 and 2 above as explanatory variables and PMS severity (severe or mild) as the dependent variable, a predictive model was constructed. For each feature, four algorithms—Random Forest, Neural Network, Linear Kernel Support Vector Machine (SVM (linear)), and rbf Kernel Support Vector Machine (SVM (rbf))—were used to train the predictive model with 10x cross-validation. For each algorithm, the RNA expression data (Log2RPM value) or principal component data of the test data were input into the trained model to calculate the predicted value of PMS severity. The accuracy, which indicates the degree of agreement between the predicted value and the measured value, was calculated, and the model with the highest accuracy was selected as the optimal predictive model.
[0062] 4) Results Table 8 shows the features used to detect PMS severity, the optimal algorithm that yielded the highest accuracy, and its accuracy. High accuracy in detecting PMS severity was possible based on luteal phase expression data of genes related to PMS severity, or their principal components. Therefore, it was demonstrated that PMS severity can be detected using luteal phase expression data of SSL-derived RNA. Furthermore, in the case of prediction models based on principal components, the accuracy of the prediction model using principal components 1-5 was higher than that of the prediction model using principal component 1, while being equivalent to that of the prediction model using principal components 1-6. These results indicate that a more accurate PMS severity prediction model can be constructed by using one or more principal components as explanatory variables such that the total contribution rate is 20% or more, for example, 22-55% or higher, preferably 50-70%, more preferably 55-65%. Additionally, in the case of prediction models directly constructed from gene expression levels, the accuracy of the prediction model using 10 genes was higher than that of the prediction model using 5 genes, while being equivalent to that of the prediction model using 15 genes. It was shown that a more accurate PMS severity prediction model can be constructed by using expression data from approximately 10 to 15 genes, or more, that show significant differences in expression depending on PMS severity.
[0063] [Table 8]
[0064] Example 3: Construction of a PMS severity prediction model based on follicular phase RNA expression data. 1) Dataset splitting and feature selection Training and test data were prepared using the same procedure as in Example 2, 1), except that a dataset of follicular phase SSL-derived RNA expression levels from subjects for the genes shown in Table 6 was used. Then, using the same procedure as in Example 2, 2), the Log2RPM values of the RNA expression data and the 1st to 15th principal components (Table 9) obtained by principal component analysis of the expression data were selected as features for PMS severity detection.
[0065] [Table 9]
[0066] 2) Model construction Using the features selected in 1) as explanatory variables, the prediction model was trained using the same procedure as in 3) of Example 2, and the model with the highest accuracy was selected as the optimal prediction model.
[0067] 3) Results Table 10 shows the features used to detect PMS severity, the optimal algorithm that yielded the highest accuracy, and its accuracy. Based on the expression levels of genes related to PMS severity during the follicular phase, or their principal components, PMS severity could be detected with high accuracy. Therefore, it was shown that it is possible to detect the severity of PMS associated with the next menstrual cycle using the expression level data of SSL-derived RNA during the follicular phase. Furthermore, in the case of a prediction model based on principal components, the accuracy of the prediction model using the 1st to 15th principal components was the highest, and it was shown that a more accurate PMS severity prediction model can be constructed by using one or more principal components as explanatory variables such that the total contribution rate is 80% or more, preferably 85% or more, and more preferably 90% or more. On the other hand, in the case of a prediction model directly constructed from gene expression levels, it was shown that a more accurate PMS severity prediction model can be constructed by using the expression data of about 5 to 20 genes, preferably 5 to 15 genes, that have a high correlation with PMS severity.
[0068] [Table 10]
Claims
1. A method for supporting the assessment of the severity of premenstrual syndrome in a subject, The mRNA expression levels of all genes selected from the group consisting of SH3BGRL3, NDUFA11, PSME2, SNORA24, and SSR4 in the test subjects will be measured. To predict the severity of premenstrual syndrome in the subject using a predictive model based on the expression level of the mRNA, Includes, The expression level of the mRNA is measured by quantifying the mRNA contained in the surface lipids of the subject's skin during the luteal or follicular phase. This predictive model is a machine learning model pre-constructed using training data that uses the expression level of mRNA contained in skin surface lipids during the luteal or follicular phase as the explanatory variable and the severity of premenstrual syndrome (PMS) (mild or severe) as the dependent variable. method.
2. A method for supporting the assessment of the severity of premenstrual syndrome in a subject, The mRNA expression levels of all genes selected from the group consisting of SH3BGRL3, NDUFA11, PSME2, SNORA24, SSR4, ZNF91, SNORA70, HSPA6, ALDH2, and BRK1 in the test subjects will be measured. To predict the severity of premenstrual syndrome in the subject using a predictive model based on the expression level of the mRNA, Includes, The expression level of the mRNA is measured by quantifying the mRNA contained in the surface lipids of the subject's skin during the luteal or follicular phase. This predictive model is a machine learning model pre-constructed using training data that uses the expression level of mRNA contained in skin surface lipids during the luteal or follicular phase as the explanatory variable and the severity of premenstrual syndrome (PMS) (mild or severe) as the dependent variable. method.
3. A method for supporting the assessment of the severity of premenstrual syndrome in a subject, The mRNA expression levels of all genes selected from the group consisting of SH3BGRL3, NDUFA11, PSME2, SNORA24, SSR4, ZNF91, SNORA70, HSPA6, ALDH2, BRK1, C20orf111, CSNK1A1, STK10, GHITM, and RPL21 in the test subjects will be measured. To predict the severity of premenstrual syndrome in the subject using a predictive model based on the expression level of the mRNA, Includes, The expression level of the mRNA is measured by quantifying the mRNA contained in the surface lipids of the subject's skin during the luteal or follicular phase. This predictive model is a machine learning model pre-constructed using training data that uses the expression level of mRNA contained in skin surface lipids during the luteal or follicular phase as the explanatory variable and the severity of premenstrual syndrome (PMS) (mild or severe) as the dependent variable. method.
4. A method for supporting the assessment of the severity of premenstrual syndrome in a subject, The mRNA expression levels of all genes selected from the group consisting of SH3BGRL3, NDUFA11, PSME2, SNORA24, SSR4, ZNF91, SNORA70, HSPA6, ALDH2, BRK1, C20orf111, CSNK1A1, STK10, GHITM, RPL21, MED13, RPL32, UBE2R2, RPS3, PLEKHM2, DDX6, ALDH1A3, TSPYL1, NUB1, KRT6A, CCL22, ELOVL3, TMEM66, GDI2, MYL12B, CCNI, PSMC6, RABGEF1, ARF1, ACTN1, CHMP2A, GAK, STX3, SERP1, and GAS7 will be measured in the subjects. To predict the severity of premenstrual syndrome in the subject using a predictive model based on the expression level of the mRNA, Includes, The expression level of the mRNA is measured by quantifying the mRNA contained in the surface lipids of the subject's skin during the luteal or follicular phase. This predictive model is a machine learning model pre-constructed using training data that uses the expression level of mRNA contained in skin surface lipids during the luteal or follicular phase as the explanatory variable and the severity of premenstrual syndrome (PMS) (mild or severe) as the dependent variable. method.
Citation Information
Patent Citations
Selection method for antimicrobial agent
JP2015228829A
Method of preparing nucleic acid sample
JP2018000206A
Method for preparing nucleic acid sample
WO2018008319A1