Method for classifying skin types
By analyzing RNA expression in skin surface lipids through cluster analysis, this method effectively classifies skin types, addressing the limitations of current methods and enabling personalized cosmetic product selection.
Patent Information
- Application Number
- JP2024184188
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-20
- Filing Date
- 2024-10-18
- Publication Date
- 2025-05-02
AI Technical Summary
Current methods for classifying skin types rely on subjective self-diagnosis and temporary, measurable skin conditions, which are not accurate or sustainable for personalized cosmetic product selection.
A method involving the collection of skin surface lipids (SSL) from healthy individuals, followed by cluster analysis based on the expression status of RNA in SSL, to classify skin types into distinct clusters.
This method enables accurate classification of skin types, facilitating the selection of optimal cosmetic products tailored to individual skin types and contributing to customized skincare solutions.
Smart Images

Figure 2025071065000015 
Figure 2025071065000016 
Figure 2025071065000001
Abstract
Description
[Technical field]
[0001] The present invention relates to a method for classifying skin type using biological information from the skin. [Background technology]
[0002] In recent years, various cosmetics have been provided according to the skin characteristics and skin concerns of the cosmetic user. For example, there are a variety of cosmetics, such as cosmetics for different skin types, such as for dry skin, combination skin, and oily skin, cosmetics for solving skin concerns, such as age spots, sagging skin, and wrinkles, as well as cosmetics for blocking ultraviolet rays.
[0003] Conventionally, in order to provide a cosmetic product suited to the individual skin characteristics of a user, counseling has been provided based on the results of a self-diagnosis questionnaire by the user and the measurement results of skin information such as texture, skin barrier function, moisture content, elasticity, and viscosity. However, the user's self-diagnosis is a subjective judgment and does not accurately reflect the user's skin characteristics. Moreover, the objective information measured by the device is merely a temporary numerical representation of the measurable skin condition, which may change from day to day.
[0004] In recent years, technology has been developed to examine the current and future physiological state of the human body by analyzing nucleic acids such as DNA and RNA in biological samples, and a method has been reported for evaluating human skin characteristics based on genetic factors using SNP analysis (Patent Document 1).
[0005] On the other hand, it has been found that RNA contained in skin surface lipids (SSL) can be used as a sample for biological analysis, and it has been reported that marker genes for the epidermis, sweat glands, hair follicles, and sebaceous glands can be detected from SSL (Patent Document 2). [Prior art documents] [Patent documents]
[0006] [Patent Document 1] International Publication No. 2021 / 029332 [Patent Document 2] International Publication No. 2018 / 008319 Summary of the Invention [Problem to be solved by the invention]
[0007] The present invention relates to providing a method for classifying the skin type of a subject who is a cosmetic user, using biological information from the skin. [Means for solving the problem]
[0008] The present inventors collected SSL from several healthy female subjects, performed cluster analysis based on the expression state of RNA contained in the SSL, and classified the SSL into two clusters. They found that each cluster had a characteristic gene expression. Furthermore, they found that the skin type of the subjects who are cosmetic users can be classified based on the gene expression.
[0009] That is, the present invention relates to the following. A method for classifying the skin of a subject, the classification being based on the expression states of genes contained in lipids on the skin surface of a plurality of healthy subjects and being generated from two clusters observed; Acquiring biological information from the skin of a subject; and and determining from the biological information which of the skin types the subject is classified into. Effect of the Invention
[0010] According to the present invention, the skin of a cosmetic user can be classified, which contributes to the selection of optimal cosmetics for each skin type or the provision of customized cosmetics. [Brief description of the drawings]
[0011] [Figure 1] Hierarchical clustering analysis using SSL-RNA. [Diagram 2]Hierarchical clustering analysis of keratinization and immune response using SSL-RNA. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0012] All patents, non-patent publications, and other publications cited herein are hereby incorporated by reference in their entirety.
[0013] In the present invention, "DNA" includes any of cDNA, genomic DNA, and synthetic DNA, and "RNA" includes any of total RNA, mRNA, rRNA, tRNA, non-coding RNA, and synthetic RNA.
[0014] In the present invention, the term "gene" refers to double-stranded DNA including human genomic DNA, single-stranded DNA (positive strand) including cDNA, single-stranded DNA (complementary strand) having a sequence complementary to the positive strand, and fragments thereof, and refers to DNA that contains some biological information in the sequence information of the bases that make up the DNA. Furthermore, the "gene" in question includes not only "genes" represented by a specific base sequence, but also nucleic acids that code for their homologues (i.e., homologs or orthologs), mutants such as gene polymorphisms, and derivatives.
[0015] The method for classifying skin of the present invention is a method for classifying a subject's skin into a skin type generated from two clusters recognized by cluster analysis based on the expression states of genes contained in skin surface lipids of a plurality of healthy subjects, Acquiring biological information from the skin of a subject; and and determining, from the biological information, into which of the skin types the subject is classified.
[0016] 1. Skin type generation In the method of classifying skin of the present invention, the expression states of genes contained in lipids on the skin surface of multiple healthy subjects are obtained in advance, and based on this, cluster analysis is performed to classify the skin into two clusters, and a skin type is generated based on the classified clusters. Here, the subject "healthy person" is a person with good skin condition, preferably a person who has not been diagnosed with a skin disease by a doctor. In addition, although age and sex are not important, the expression state of genes contained in the lipids on the skin surface of multiple healthy people for each age, generation, and sex may be obtained according to the age and sex of the subject to be subjected to skin classification, and a skin type may be generated for each. For example, if the subject is a woman, the skin type is generated from the expression state of genes contained in the lipids on the skin surface of multiple healthy female people, if the subject is a man ... male people, if the subject is both male and female, the skin type is generated from the expression state of genes contained in the lipids on the skin surface of multiple healthy female people and healthy male people. Here, "plurality" refers to an integer of 2 or more, preferably 10 or more, more preferably 100 or more, even more preferably 150 or more, and still more preferably 250 or more.
[0017] In the present invention, the biological sample used in gene expression analysis is skin surface lipids (SSL). "Skin surface lipids (SSL)" refers to a fat-soluble fraction present on the surface of the skin, and is also called sebum. In general, SSL mainly contains secretions secreted from exocrine glands such as sebaceous glands in the skin, and exists on the skin surface in the form of a thin layer covering the skin surface. SSL contains RNA expressed in skin cells (see Patent Document 2). In this specification, "skin" is a general term for an area including tissues such as the epidermis, dermis, hair follicles, sweat glands, sebaceous glands, and other glands on the body surface, unless otherwise specified.
[0018] To collect SSL from the skin, any means used for recovering or removing SSL from the skin can be used. Preferably, the SSL absorbent material, SSL adhesive material, or a tool for scraping SSL from the skin, described later, can be used. The SSL absorbent material or SSL adhesive material is not particularly limited as long as it has an affinity for SSL, and examples thereof include polypropylene, pulp, and the like. More detailed examples of procedures for collecting SSL from the skin include a method of absorbing SSL into a sheet-like material such as oil blotting paper or oil blotting film, a method of adhering SSL to a glass plate, tape, or the like, and a method of scraping SSL off and collecting it with a spatula, scraper, or the like. In order to improve the adsorption of SSL, an SSL absorbent material that has been previously impregnated with a highly lipid-soluble solvent may be used. On the other hand, if the SSL absorbent material contains a highly water-soluble solvent or moisture, the adsorption of SSL is inhibited, so it is preferable that the content of highly water-soluble solvent or moisture is low. It is preferable to use the SSL absorbent material in a dry state. The site of skin from which the SSL is collected is not particularly limited and may be any site of the body such as the head, face, neck, trunk, limbs, etc., although sites with high sebum secretion, such as the skin of the face, are preferred.
[0019] The collected RNA-containing SSL may be stored for a certain period of time. In order to minimize the degradation of the RNA contained therein, the collected SSL is preferably stored under low temperature conditions as soon as possible after collection. The temperature condition for storing the RNA-containing SSL in the present invention may be 0° C. or lower, preferably −20±20° C. to −80±20° C., more preferably −20±10° C. to −80±10° C., even more preferably −20±20° C. to −40±20° C., even more preferably −20±10° C. to −40±10° C., even more preferably −20±10° C., and even more preferably −20±5° C. The period for storing the RNA-containing SSL under the low temperature conditions is not particularly limited, but is preferably 12 months or less, for example, 6 hours or more and 12 months or less, more preferably 6 months or less, for example, 1 day or more and 6 months or less, even more preferably 3 months or less, for example, 3 days or more and 3 months or less.
[0020] In the present invention, the expression state of a gene means an indicator showing the degree or tendency of gene expression. The expression state of a gene contained in an SSL is obtained by measuring the expression level (expression amount) of RNA, specifically, by converting RNA into cDNA by reverse transcription, and then measuring the cDNA or its amplification product. RNA can be extracted from SSL by methods commonly used for extracting or purifying RNA from biological samples, such as the phenol / chloroform method, the acid guanidinium thiocyanate-phenol-chloroform extraction (AGPC) method, or methods using columns such as TRIzol (registered trademark), RNeasy (registered trademark), or QIAzol (registered trademark), methods using special silica-coated magnetic particles, methods using Solid Phase Reversible Immobilization magnetic particles, or extraction using a commercially available RNA extraction reagent such as ISOGEN.
[0021] For the reverse transcription, a primer targeting a specific RNA to be analyzed may be used, but for more comprehensive nucleic acid storage and analysis, it is preferable to use a random primer. For the reverse transcription, a general reverse transcriptase or reverse transcription reagent kit can be used. Preferably, a highly accurate and efficient reverse transcriptase or reverse transcription reagent kit is used, and examples thereof include M-MLV Reverse Transcriptase and its variants, or commercially available reverse transcriptase or reverse transcription reagent kits, such as PrimeScript (registered trademark) Reverse Transcriptase series (Takara Bio Inc.), SuperScript (registered trademark) Reverse Transcriptase series (Thermo Scientific Co., Ltd.), etc. SuperScript (registered trademark) III Reverse Transcriptase, SuperScript (registered trademark) VILO cDNA Synthesis kit (both Thermo Scientific Co., Ltd.), etc. are preferably used. The temperature of the extension reaction in the reverse transcription is preferably adjusted to 42°C±1°C, more preferably 42°C±0.5°C, and even more preferably 42°C±0.25°C, while the reaction time is preferably adjusted to 60 minutes or more, more preferably 80 to 120 minutes.
[0022] The expression level of RNA can be measured, for example, by quantifying the cDNA or its amplification product by real-time PCR, multiplex PCR, microarray, sequencing, chromatography, etc. In one embodiment of the method of the present invention, a library containing DNA derived from expressed genes is prepared using multiplex PCR, the prepared library is sequenced by a next-generation sequencer, and calculation is performed based on the number of reads generated (read count).
[0023] Multiplex PCR is a method for amplifying multiple gene regions simultaneously by using multiple primer pairs simultaneously in a PCR reaction system. Multiplex PCR can be performed using a commercially available kit (e.g., Ion AmpliSeq Transcriptome Human Gene Expression Kit; Life Technologies Japan, Inc., etc.) under the following conditions: 99°C, 2 minutes → (99°C, 15 seconds → 62°C, 16 minutes) × 20 cycles → 4°C, hold.
[0024] The purification of the reaction product obtained by the PCR is preferably carried out by size separation of the reaction product. By size separation, the target PCR reaction product can be separated from primers and other impurities contained in the PCR reaction solution. Size separation of DNA can be carried out, for example, by a size separation column, a size separation chip, magnetic beads usable for size separation, etc. A preferred example of magnetic beads usable for size separation includes Solid Phase Reversible Immobilization (SPRI) magnetic beads such as Ampure XP.
[0025] For the purified PCR reaction product, the purified PCR reaction product is prepared in an appropriate buffer solution for DNA sequencing, the PCR primer region contained in the PCR-amplified DNA is cut, and an adapter sequence is further added to the amplified DNA. For example, the purified PCR reaction product is prepared in a buffer solution, the PCR primer sequence is removed from the amplified DNA and adapter ligation is performed, and the resulting reaction product is amplified as necessary to prepare a library for sequencing. These operations can be performed, for example, using the 5×VILO RT Reaction Mix included in the SuperScript (registered trademark) VILO cDNA Synthesis kit (Life Technologies Japan, Inc.), and the 5×Ion AmpliSeq HiFi Mix and Ion AmpliSeq Transcriptome Human Gene Expression Kit (Life Technologies Japan, Inc.) and the Ion AmpliSeq Transcriptome Human Gene Expression Core Panel included in the Ion AmpliSeq Transcriptome Human Gene Expression Kit according to the protocol included with each kit.
[0026] Sequencing can be performed using a next-generation sequencer (e.g., Ion S5 / XL system, Life Technologies Japan Co., Ltd.), and the genes derived from each read sequence are determined by genetic mapping of each read sequence obtained by sequencing to the hg19 AmpliSeq Transcriptome ERCC v1, which is the reference sequence of the human genome. RNA expression levels are calculated based on the number of reads (read count) generated by sequencing.
[0027] The read count value, which is expression level data obtained by sequencing, is appropriately corrected. Examples of corrected values for the read count value include an RPM value obtained by correcting the read count value for differences in the total number of reads between samples, a value obtained by converting the RPM value to a logarithmic value with base 2 (Log2RPM value) or a logarithmic value with base 2 with an integer 1 added (Log2(RPM+1) value), or a count value corrected using DESeq2 (Love MI et al. Genome Biol. 2014) (Normalized count value) or a logarithmic value with base 2 with an integer 1 added (Log2(count+1) value), with the Normalized count value corrected using DESeq2 being preferred.
[0028] Based on the obtained information on the expression state of genes contained in the lipids on the skin surface, cluster analysis is carried out.The method of cluster analysis is not particularly limited as long as a plurality of objects can be classified into two or more clusters, and may be hierarchical clustering or non-hierarchical clustering.Cluster analysis can be carried out by software capable of statistical analysis. For example, as shown in the examples described below, 281 SSL samples taken from the entire face of healthy female subjects aged 20 to 60 were targeted, and the correlation coefficient ρ between the samples was calculated using Spearman's rank correlation analysis based on the expression levels (normalized count values) of RNA derived from the SSL. Hierarchical clustering analysis was then performed with the distance between the samples set to 1-ρ, and the samples were classified into two clusters, Cluster 1 and Cluster 2, as shown in Figure 1.
[0029] In the present invention, it is possible to generate skin types associated with the two clusters based on gene expression information that characterizes the clusters.
[0030] When generating a skin type based on gene expression information, as shown in the examples described below, by extracting RNA in which the SSL-derived RNA expression ratio between the two clusters is 1.2 times or more and the corrected p-value (FDR) by the likelihood ratio test is less than 0.05, 614 genes with relatively higher expression in Cluster 1 compared to Cluster 2 and 855 genes with relatively higher expression in Cluster 2 compared to Cluster 1 were identified. Gene Ontology (GO) enrichment analysis confirmed that the group of genes with relatively higher expression in Cluster 1 included many genes related to epidermal keratinization, and the group of genes with relatively higher expression in Cluster 2 included many genes related to immune response. That is, genes annotated with GO terms such as skin development (GO:0043588) and epidermis development (GO:0008544) were included in the keratinization-related genes with relatively high expression in Cluster 1, while genes annotated with GO terms such as immune system process (GO:0002376) and response to cytokines (GO:0034097) were included in the immune response-related genes with relatively high expression in Cluster 2. GO terms (Gene Ontology) are common vocabulary for explaining the biological concept of a gene (biological process of a gene, components of a cell, and molecular function), and GO terms are annotated (linked) to genes whose functions have been clarified. As a result, a GO term that describes the function is assigned to each gene, and the gene to which the term is assigned can be confirmed from the GO term. As shown in the examples below, cluster analysis was performed based on information on the expression state of genes annotated with the keratinization-related GO terms skin development (GO:0043588) and epidermis development (GO:0008544), and genes annotated with the immune response-related GO terms immune system process (GO:0002376) and response to cytokine (GO:0034097). The clusters (Cluster 1' and Cluster 2', Figure 2) were almost equivalent to Cluster 1 and Cluster 2 (Table 1). Therefore, Cluster 1 can be said to be a group with high expression of keratinization-related genes, and Cluster 2 can be said to be a group with high expression of immune response-related genes. From these two clusters, two skin types, namely, a skin type with high expression of keratinization-related genes and a skin type with high expression of immune response-related genes, which are found in the skin of healthy women, are generated.
[0031] Furthermore, by analyzing various biological information from the skin of healthy subjects belonging to the high keratinization-related gene expression group of Cluster 1 or Cluster 1' and the high immune response-related gene expression group of Cluster 2 or Cluster 2', it is possible to link the respective characteristic biological information to the two generated skin types. As biological information from the skin, any information that can be obtained from the skin, regardless of whether it is physiological information or physical information, can be used.
[0032] 2. Classification of subjects' skin types In the present invention, biological information is obtained from the skin of a subject, and it is determined from the biological information whether the subject is classified into one of the two skin types generated from the cluster analysis. Here, the "subject" may be a cosmetic product user or a potential user considering using the cosmetic product, and there are no limitations on gender or age. The term "biological information from the skin" is not particularly limited as long as it is information obtainable from the skin, regardless of whether it is physiological information or physical information. Specifically, examples include information on gene expression status in the SSL used in the above-mentioned cluster analysis, information on protein expression status and metabolites in the SSL, information on gene expression status, protein expression status and metabolites in biological samples that can be taken from the face (skin biopsy, stratum corneum, sweat, tears, saliva, etc.), information on gene expression status, protein expression status and metabolites of the flora of normal skin bacteria or bacteria-derived gene expression status, protein expression status and metabolites, and measurements of skin property values measured by equipment. Preferable examples of equipment measurements include transepidermal water loss, stratum corneum moisture content, sebum content, skin glycation level, skin viscoelasticity, skin color, hemoglobin content, melanin content, skin blood flow, skin temperature, biophotons, etc. When information on the gene expression state in SSL is used as biological information from the skin, collection of SSL and measurement of gene expression levels in SSL can be performed in the same manner as described above.
[0033] Methods for determining the skin type of a subject based on the measured gene expression state include, for example, 1) a method of determining from the distance between the center of gravity of the two pre-generated clusters (Cluster 1 and Cluster 2, or Cluster 1' and Cluster 2') and the subject's SSL-derived RNA expression level, and 2) a method of determining based on a comparison with the expression level of specific genes related to keratinization or immune response. The method of determining the distance between the center of gravity of the two clusters and the expression level of SSL-derived RNA of the subject can be carried out by calculating the distance (similarity) between the two clusters and the expression level of SSL-derived RNA obtained from the subject. For example, for two clusters (Cluster 1 and Cluster 2, or Cluster 1' and Cluster 2') that are the population, the center of gravity is calculated from a multidimensional vector with the SSL-derived RNA expression level as a variable, and the distance between the multidimensional vector based on the SSL-derived RNA expression level obtained from the subject and the center of gravity of the two clusters is calculated, and the cluster with the closest distance is determined as the type of the subject. As for the distance, it is also possible to calculate the Spearman's rank correlation coefficient ρ between the center of gravity and the subject's data as in the cluster analysis, and use 1-ρ. In addition, at this time, as the variable that determines the center of gravity or the RNA expression level obtained from the subject, it is also possible to narrow it down to genes that are relatively highly expressed in Cluster 1 compared to Cluster 2, genes that are relatively highly expressed in Cluster 2 compared to Cluster 1, and genes that are annotated with GO terms related to keratinization or immune response.
[0034] In a method for determining skin type based on a comparison with the expression levels of specific genes related to keratinization or immune response, one or more gene species with no or less overlap in expression levels between two clusters (Cluster 1 and Cluster 2, or Cluster 1' and Cluster 2') that are the population are selected, a threshold for the expression level is set based on statistical values such as the average value and standard deviation, and the type can be determined from the SSL-derived RNA expression level of the relevant gene species obtained from the subject. Here, the specific gene related to keratinization or immune response is preferably selected from genes annotated with the keratinization-related or immune response-related GO term, and more preferably selected from the 25 genes shown in Table 7 below from the viewpoint of accuracy in discriminating between the two clusters.
[0035] In one embodiment, the expression level of one or more genes selected from the 25 genes is obtained from a subject, and the expression level is compared with a reference value to determine the skin type of the subject. Here, the reference value can be a threshold value in the Youden index of the ROC curve.
[0036] In another embodiment, the expression levels of one or more genes selected from specific genes related to keratinization or immune response are obtained from a subject, and the expression levels are substituted into a discriminant (discriminant model) for detecting skin type, thereby determining the skin type of the subject. Specifically, a discriminant equation can be constructed by machine learning using an arbitrary group as a teacher sample, obtaining the expression level of a specific gene related to keratinization or immune response from each person in the teacher sample as an explanatory variable and using each person's skin type as a target variable, and then measuring the expression level of the specific gene in the subject and substituting the expression level into the discriminant equation to discriminate the skin type of the subject. Here, the skin type of each person in the teacher sample can be determined by carrying out the above-mentioned cluster analysis based on information on the expression state of genes contained in the lipids on the skin surface of each person.
[0037] The specific genes related to keratinization or immune response used to construct the discriminant equation can be selected from genes annotated with the keratinization-related or immune response-related GO terms, but it is preferable to select one or more from the 25 genes shown in Table 7 below.
[0038] As an algorithm for constructing the discriminant, known algorithms such as algorithms used in machine learning can be used. Examples of machine learning algorithms include random forest, decision tree, gradient boosting, XGBossti (eXtreme Gradient Boosting), LightGBM, support vector machine (SVM linear) with a linear kernel, support vector machine (SVM rbf) with an rbf kernel, logistic regression, regularized logistic regression, generalized linear model, regularized linear discriminant analysis, k-nearest neighbors, neural net, multi-layer perceptron, convotional neural network, recurrent neural network, etc. The prediction model constructed is input with data for verification to calculate a prediction value, and the model whose prediction value matches the actual measurement value best, for example, the model with the highest accuracy, can be selected as the optimal prediction model. In addition, the recall, precision, and the F value, which is the harmonic mean of these, can be calculated from the prediction value and the actual measurement value, and the model with the highest F value can be selected as the optimal prediction model.
[0039] When using information other than information on gene expression state in SSL as biometric information from skin, it is preferable to use biometric information linked to the two clusters (Cluster 1 and Cluster 2, or Cluster 1' and Cluster 2') that are generated. For the two clusters (Cluster 1 and Cluster 2, or Cluster 1' and Cluster 2') that are the population, the biometric information is acquired in advance, and a criterion for classifying into one of the two clusters, for example a threshold, is set, or a discriminant for classifying into one of the two clusters by machine learning is constructed. Then, based on the biometric information acquired from the subject, it is possible to determine which of the two clusters the skin type generated corresponds to.
[0040] The method for classifying the skin of a subject of the present invention can be performed using a calculation device (computer). Therefore, in another aspect, the present invention provides a calculation device for executing the above-mentioned method, a program for causing the calculation device to execute the above-mentioned method, and an information recording medium readable by the calculation device and having the program recorded thereon.
[0041] The computing device of the present invention has a means for inputting biometric information from the skin obtained from a subject, and includes a step of executing the above-mentioned method of determining which of the above-mentioned skin types the subject is classified into from the biometric information, in accordance with a program for executing the method of classifying the subject's skin of the present invention.
[0042] Examples of information recording media readable by a computing device and recording a program for executing the method for classifying the skin of a subject of the present invention include magnetic disks, optical disks, magneto-optical disks, flash memories, etc. In the present invention, "readable by a computing device" also includes the case where the program is distributed via a telecommunications line, etc. EXAMPLES
[0043] The present invention will be described in more detail below based on examples, but the present invention is not limited to these examples. Example 1: Skin type classification analysis of SSL-RNA 1) SSL collection Sebum samples were collected from 281 healthy female subjects aged 20 to 60 (24 samples from those in their 20s, 126 samples from those in their 30s, 29 samples from those in their 40s, and 102 samples from those over 50). Sebum samples were collected from the entire face of each subject using a sheet of oil blotting film (polypropylene, 5.0 cm x 8.0 cm, 3M).
[0044] 2) RNA preparation and sequencing The oil blotting film from which sebum was collected in 1) above was cut to an appropriate size, and RNA was extracted using QIAzol Lysis Reagent (Qiagen) according to the attached protocol. Based on the extracted RNA, reverse transcription was performed at 42°C for 90 minutes using SuperScript VILO cDNA Synthesis kit (Life Technologies Japan, Inc.) to synthesize cDNA. The random primers included in the kit were used as primers for the reverse transcription reaction. From the obtained cDNA, a library containing DNA derived from the 20802 gene was prepared by multiplex PCR. Multiplex PCR was performed using Ion AmpliSeq Transcriptome Human Gene Expression Kit (Life Technologies Japan, Inc.) under the following conditions: [99°C, 2 minutes → (99°C, 15 seconds → 62°C, 16 minutes) × 20 cycles → 4°C, hold]. The obtained PCR product was purified with Ampure XP (Beckman Coulter, Inc.), followed by buffer reconstitution, digestion of the primer sequence, adapter ligation and purification, and amplification to prepare a library. The prepared library was loaded onto an Ion 540 chip and sequenced using an Ion S5 / XL system (Life Technologies Japan, Inc.). The genes from which each read sequence was derived were determined by genetic mapping of each read sequence obtained by sequencing against the hg19 AmpliSeq Transcriptome ERCC v1, a reference sequence for the human genome.
[0045] 3) Data Analysis The RNA expression data (read count values) derived from the SSL of the subjects obtained in 2) above were corrected using a method called DESeq2 (Love MI et al. Genome Biol. 2014). However, samples in which 4161 or more genes were not detected were excluded, and only 2605 genes for which expression data that was not a missing value was obtained in 90% or more of the sample subjects out of all the sample data after exclusion were used in the following analysis. Normalized count values corrected using a method called DESeq2 were used for the analysis. Based on the SSL-derived RNA expression levels (normalized count values) obtained above, the correlation coefficient ρ between samples was calculated by Spearman's rank correlation analysis. At this time, the variance of the count values was corrected using a method called Valiance Stabilizing Transformation. Based on the correlation coefficient ρ, a hierarchical clustering analysis was performed in which the distance between samples was set to 1-ρ and the Ward method was used to link clusters. Based on the dendrogram obtained from the analysis, the samples were divided into two clusters (Cluster 1 and Cluster 2) (Figure 1).
[0046] By extracting RNA with an SSL-derived RNA expression ratio of 1.2 times or more between the two clusters and a p-value corrected value (FDR) of the likelihood ratio test of less than 0.05, 614 genes with relatively high expression in Cluster 1 and 855 genes with relatively high expression in Cluster 2 were identified. We used the public database PANTHER to search for biological process (BP) by gene ontology (GO) enrichment analysis. As a result, 215 GO terms were obtained from the genes with relatively high expression in Cluster 1, and 1287 GO terms were obtained from the genes with relatively high expression in Cluster 2. Among these were terms related to keratinization and immune response, which are major functions of the epidermal stratum corneum. In Cluster 1, the terms with the lowest FDR were "skin development (GO:0043588, FDR:5.23E-18)" and "epidermis development (GO:0008544, FDR:8.66E-17)," and three of the top five terms were related to keratinization. In Cluster 2, the terms with the lowest FDR were "immune system process (GO:0002376, FDR:1.84E-66)" and "response to cytokine (GO:0034097, FDR:7.22E-47)," and four of the top five terms were related to immune response.
[0047] Example 2 Skin type classification analysis using SSL-RNA in relation to keratinization and immune response 1) Usage Data As in Example 1, based on the subject's SSL-derived RNA expression levels (normalized count values), only 2,605 genes for which non-missing expression data was obtained in 90% or more of all samples were used in the following analysis.
[0048] 2) Selection of genes related to keratinization and immune response From the 2,605 genes, 76 genes annotated with the GO term "skin development (GO:0043588)" and 83 genes annotated with "epidermis development (GO:0008544)" were extracted as genes related to keratinization, and 453 genes annotated with the GO term "immune system process (GO:0002376)" and 227 genes annotated with "response to cytokine (GO:0034097)" were extracted as genes related to immune response.
[0049] 3) Data Analysis Based on the expression levels (normalized count values) of a total of 606 genes annotated with specific GO terms related to keratinization and immune response selected in 2) from the SSL-derived RNA, the correlation coefficient ρ between samples was calculated by Spearman's rank correlation analysis. At this time, the variance of the count values was corrected using a method called Valiance Stabilizing Transformation. Based on the correlation coefficient ρ, a hierarchical clustering analysis was performed in which the distance between samples was set to 1-ρ and the Ward method was used to link clusters. Based on the dendrogram obtained from the analysis, the samples were divided into two clusters (Cluster 1' and Cluster 2') (Figure 2).
[0050] 4) Comparison of cluster analysis results The two clusters observed in Example 1 were compared with the two clusters observed from the above data analysis. Of the 107 samples belonging to Cluster 1, 106 samples were included in Cluster 1', and of the 174 samples belonging to Cluster 2, 135 samples were included in Cluster 2', with a concordance rate of 85% or more (Table 1).
[0051] [Table 1]
[0052] Example 3: Estimation of skin type using cluster centroids 1) Usage Data The data on the expression levels of SSL-derived RNA obtained in Example 1 (read count values) were converted to RPM values corrected for the difference in the total number of reads between samples. However, only 2605 types of genes for which non-missing expression level data was obtained in 90% or more of all samples were used in the following analysis. To calculate the center of gravity of each cluster, RPM values converted to logarithmic values with base 2 (Log2RPM values) were used to approximate the RPM values that follow a negative binomial distribution to a normal distribution.
[0053] 2) Dividing the sample The samples were divided into training samples (227 samples) and test samples (54 samples). The skin type of each sample was determined to be Cluster 1 or Cluster 2 recognized in the hierarchical clustering analysis of Example 1. At this time, the samples were divided so that there was no bias in the proportion of Cluster 1 and Cluster 2 belonging to the training samples and the test samples, respectively (Table 2).
[0054] [Table 2]
[0055] 3) Calculation of cluster centroids The expression data (Log2RPM value) of the gene to be analyzed of the SSL-derived RNA of the training sample (training data) was used as a variable, and the centroids (Centroid1, Centroid2) of each cluster (Cluster1, Cluster2) were calculated. In calculating the centroids, the expression data of each sample was expressed as a vector with the expression amount of each gene as a component, and the expression data was positioned as one data point in a 2605-dimensional space, which is the number of genes to be analyzed, and the average of the position vectors of the data points of all samples was used as the centroid. For example, if the data points belonging to Cluster1 are x_1, x_2,..., x_87, the centroid1 of the cluster can be expressed as (1 / 87)*(x_1+x_2+...+x_87). The centroids calculated in this way were used to determine the skin type of the test sample.
[0056] 4) Skin type determination based on center of gravity The expression level data (Log2RPM value) of the gene to be analyzed in the SSL-derived RNA of the test sample (test data) was used as a variable to calculate the Euclidean distance between the test sample and Centroid1 and Centroid2. The distance between each test sample and Centroid1 was compared with the distance between each test sample and Centroid2, and the cluster with the closest center of gravity was determined to be the skin type (Cluster1 and Cluster2) of each test sample.
[0057] 5) Verification of judgment accuracy When the skin type of the test sample was compared between that determined in Example 1 and that determined based on the center of gravity of the training data, the accuracy rate (match rate) was 95% or higher (Table 3).
[0058] [Table 3]
[0059] Example 4: Selection of skin type determining marker genes 1) Data used and sample division The data used were the expression level data (read count values) of SSL-derived RNA obtained in Example 1. As in Example 3, the samples were divided into training samples (227 samples) and test samples (54 samples) (Table 2). The skin type of each sample was determined to be Cluster 1 or Cluster 2 as determined by the hierarchical clustering analysis in Example 1.
[0060] 2) Analysis of training data The RNA expression data (read count values) derived from the SSL of the training samples were corrected using DESeq2 to create the training data. Only 2563 genes for which non-missing expression data was obtained in more than 90% of the sample subjects were used in the following analysis. Normalized count values corrected using a method called DESeq2 were used for the analysis. By extracting RNA in which the SSL-derived RNA expression ratio of the training data was 1.2 times or more and the corrected p-value (FDR) by the likelihood ratio test was less than 0.05 between the two skin types of the training sample, 591 genes with relatively high expression in Cluster 1 and 823 genes with relatively high expression in Cluster 2 were identified.
[0061] 3) Selection of differentially expressed genes related to keratinization and immune response 2) Among the 591 genes that showed relatively high expression in Cluster 1, 54 genes were extracted that were annotated with either the GO terms "skin development (GO:0043588)" or "epidermis development (GO:0008544)" as genes related to keratinization. In addition, among the 823 genes that showed relatively high expression in Cluster 2, 313 genes were extracted that were annotated with either the GO terms "immune system process (GO:0002376)" or "response to cytokine (GO:0034097)" as genes related to immune response.
[0062] 4) Evaluation of the discrimination accuracy of individual genes in the training sample For each of the 367 genes extracted in 3), an ROC curve was drawn using the expression data (Log2RPM value) as a variable, and genes more suitable for discriminating between Cluster 1 and Cluster 2 of the training sample were extracted, and the discrimination accuracy was evaluated using the genes. When discriminating whether a certain sample is negative or positive using the magnitude of a certain continuous value as an index, the ROC curve plots the probability of a positive result in the positive group (sensitivity or true positive rate) on the vertical axis and the value obtained by subtracting the probability of a negative result in the negative group (specificity) from 1 (false positive rate) on the horizontal axis for the negative and positive groups. The larger the area under the curve (AUC) of the ROC curve, the higher the discrimination accuracy of the two groups is judged to be. In this study, the AUC of the 367 genes was calculated and summary statistics were obtained (Table 4). Of the 367 genes, 93 genes with AUCs above the third quartile (0.808) were extracted. Next, the gene expression level in the Youden index of the ROC curve for each of the 93 genes was set as a threshold value, and the skin type of the training sample was classified. The accuracy rate was calculated as an index of classification accuracy, and summary statistics were obtained (Table 5).
[0063] [Table 4]
[0064] [Table 5]
[0065] 5) Evaluation of the discrimination accuracy of individual genes in test samples Of the 93 genes evaluated for discrimination accuracy in 4), 46 genes with better discrimination accuracy and a correct answer rate equal to or higher than the median (0.795) in Table 5 were extracted. For these 46 genes, the threshold set in the skin type discrimination of the training sample in 4) was used, and the threshold was compared with the expression amount data (Log2RPM value) of the test sample to discriminate the skin type of the test sample. The correct answer rate was calculated as an index of discrimination accuracy, and summary statistics were obtained (Table 6). Of the 46 genes used for discrimination, 25 genes with better discrimination accuracy in the test sample and a correct answer rate equal to or higher than the median (0.796) in Table 6 are optimal skin type determination marker genes that can discriminate Cluster 1 or Cluster 2 with higher accuracy. Table 7 shows the gene names of the 25 genes, the AUC of the ROC curve, the threshold used to discriminate skin type, the discrimination method, the accuracy rate in training samples (training accuracy rate), and the accuracy rate in test samples (test accuracy rate).
[0066] [Table 6]
[0067] [Table 7]
[0068] Example 5: Construction of a skin type prediction model using skin type determination marker genes 1) Data used and sample division The corrected value (Log2RPM value) of the data on the expression level of SSL-derived RNA used in Example 3 was used. The samples were divided into training samples (227 samples) and test samples (54 samples) in the same manner as in Example 3. The skin type of each sample was determined to be the skin type (Cluster 1 or Cluster 2) recognized by the hierarchical clustering analysis in Example 1.
[0069] 2) Feature selection Several genes were randomly extracted from the 25 genes used as skin type determination markers in 5) of Example 4, and the extracted gene set was used as a feature. Specifically, the feature amount obtained by extracting 2 genes from 25 genes was named Feature 1, the feature amount obtained by extracting 3 genes from 25 genes was named Feature 2, the feature amount obtained by extracting 4 genes from 25 genes was named Feature 3, the feature amount obtained by extracting 5 genes from 25 genes was named Feature 4, the feature amount obtained by extracting 10 genes from 25 genes was named Feature 5, the feature amount obtained by extracting 15 genes from 25 genes was named Feature 6, the feature amount obtained by extracting 20 genes from 25 genes was named Feature 7, and the feature amount obtained by using all 25 genes as features was named Feature 8. Extraction of features related to Features 1 to 8 was attempted 10 times for each Feature. The genes extracted as features for each feature are shown in Tables 8-1 to 8-7.
[0070] 3) Model construction A binary classification model was constructed using the expression data (Log2RPM value) of the features selected from the SSL-derived RNA of the training sample as explanatory variables and skin type (Cluster 1 or Cluster 2) as the objective variable. Logistic regression was used as the algorithm. Eight models were constructed using Features 1 to 8 as features. As mentioned above, Features 1 to 7 exist for the number of random extraction trials (10 times), so a total of 71 models were constructed.
[0071] 4) Verification of the discriminant model The expression data (Log2RPM value) of the features of the training samples was input to each of 10 models (Models 1-10) constructed using Features 1-7 and a model constructed using Feature 8, and the skin type of the training sample was determined. The accuracy rate (match rate) of the skin type of the training sample determined by the model relative to the skin type of the training sample recognized in Example 1 is shown as the training accuracy rate in Tables 8-1-7. The training accuracy rate of the model constructed using Feature 8 (25 genes) was 1.00. A high accuracy rate was achieved using either model, confirming that the discrimination model constructed in 3) can function as a skin type discrimination model.
[0072] 5) Identifying the skin type of the test sample The expression data (Log2RPM value) of the features of the test samples were input to each of 10 models (Models 1-10) constructed using Features 1-7 and a model constructed using Feature 8, and the skin type of the test sample was determined. The accuracy rate (match rate) of the skin type of the test sample determined by the model relative to the skin type of the test sample recognized in Example 1 is shown as the test accuracy rate in Tables 8-1-7. The test accuracy rate for the model constructed using Feature 8 (25 genes) was 0.981. As with the training samples, both models showed high accuracy rates when using the test samples, demonstrating that the 25 genes shown in Table 7 are useful as markers for discriminating skin types.
[0073] [Table 8-1]
[0074] [Table 8-2]
[0075] [Table 8-3]
[0076]
Table 8-4
[0077]
Table 8-5
[0078]
Table 8-6
[0079]
Table 8-7
Claims
1. A method for classifying the skin of a subject, the classification being into skin types generated from two clusters obtained by performing cluster analysis based on expression states of genes contained in lipids on the skin surface of a plurality of healthy subjects; Acquiring biological information from the skin of a subject; and and determining from the biological information which of the skin types the subject is classified into.
2. The method according to claim 1, further comprising a preparatory step of dividing the skin of a plurality of healthy subjects into two clusters by cluster analysis based on the expression state of genes contained in the skin surface lipids, and generating skin types from each cluster.
3. The method according to claim 1 or 2, wherein the genes include genes annotated with at least one GO term selected from skin development (GO:0043588), epidermis development (GO:0008544), immune system process (GO:0002376, and response to cytokine (GO:0034097).
4. The method according to any one of claims 1 to 3, wherein the biological information from the skin is the expression state of genes contained in lipids on the surface of the skin of the subject.
5. The method according to claim 4, wherein the expression state of the gene is an expression state of a gene annotated with at least one GO term selected from skin development (GO:0043588), epidermis development (GO:0008544), immune system process (GO:0002376, and response to cytokine (GO:0034097).
6. The method according to claim 4, wherein the expression state of the gene is the expression state of at least one gene selected from the following 25 genes: ASPRV1, KRT17, KRT80, KRT79, DNASE1L2, SPRR1A, DSP, CDSN, CST6, LCP1, KRT6B, LCE1C, CALML5, CD83, JUP, TYROBP, SPINT1, CNFN, TMSB4X, NFKB1, B2M, MSN, CXCL16, CD58 and KRT72.
Citation Information
Patent Citations
Method for preparing nucleic acid sample
WO2018008319A1
Genetic testing method for implementing skin care counseling
WO2021029332A1