Data processing method for RNA information
By setting thresholds for sample and gene selection and applying normalization techniques, the method effectively addresses the challenges of missing values and variations in RNA data from secretions, ensuring accurate and reproducible analysis.
Patent Information
- Application Number
- JP2021082826
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-05-14
- Filing Date
- 2021-05-14
- Publication Date
- 2025-07-24
- Estimated Expiration
- 2041-05-14
AI Technical Summary
RNA information collected from secretions such as sebum and saliva, particularly from skin surface lipids (SSL), often has many missing values and large variations, leading to inaccuracies and reduced reproducibility in statistical analysis.
A data processing method that involves setting thresholds for sample and gene selection criteria, specifically using TD and SD values to exclude samples and genes with low expression levels, followed by normalization using methods like DESeq2 to achieve effective normalization and reproducible statistical analysis.
Enables accurate and reproducible comparison of RNA expression profiles across multiple samples by addressing the issues of missing values and variations, enhancing the reliability of RNA analysis.
Smart Images

Figure 0007712792000007 
Figure 0007712792000001 
Figure 0007712792000002
Abstract
Description
Technical Field
[0001] The present invention relates to a data processing method for RNA information in human-derived secretions.
Background Art
[0002] In recent years, techniques have been developed to examine the current and even future physiological states in the human body by analyzing nucleic acids such as DNA and RNA in biological samples. Analyses using nucleic acids have the advantages that comprehensive analysis methods have been established and abundant information can be obtained in a single analysis, and functional associations of analysis results are easily achievable based on many research reports on single nucleotide polymorphisms, RNA functions, and the like. Nucleic acids derived from living organisms can be extracted from body fluids such as blood, secretions, tissues, etc. Recently, it has been reported that RNA contained in skin surface lipids (SSL) can be used as a sample for biological analysis, and marker genes of the epidermis, sweat glands, hair follicles, and sebaceous glands can be detected from SSL (Patent Document 1).
[0003] RNA sequencing (RNA-Seq) analysis, which directly quantifies RNA sequences expressed in cells, is an analysis method currently attracting attention because it enables the detection of low-expression genes that were difficult to quantify with microarrays using signal intensity ratios and can obtain a highly accurate expression profile. In gene expression analysis, the concentration and / or relative or absolute amount of a specific RNA in a sample is determined, and the specific RNA is quantified. In this case, a method with high accuracy and reproducibility is desired. However, in biological samples collected from different individuals, there may be a bias in the expression level profile depending on the biological sample and the analysis process, so the quantity of a specific RNA cannot always be directly compared. Therefore, in order to compare the quantity of a specific RNA well in biological samples derived from two or more different individuals, normalization of the quantity of RNA between samples is carried out.
[0004] In RNA-seq analysis, the number of sequence reads mapped to the genome is used to quantify the gene expression level. Therefore, for normalization, methods such as RPM; Reads Per Million reads mapped (Non-Patent Document 1) and RLE; Relative Log Expression (Non-Patent Document 2), which are correction methods using the total number of reads, are used. Normalization by RLE is implemented in an analysis method for performing a series of gene expression level analyses called DESeq2.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Non-Patent Documents
[0006]
Non-Patent Document 1
Non-Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0007] However, the information of RNA collected from secretions such as sebum and saliva, particularly RNA collected from SSL, has many missing values and large variations. Therefore, when performing the same data processing as other RNA information, even if subsequent statistical processing such as machine learning is performed, problems may occur in terms of accuracy and reproducibility. The present invention relates to providing a data processing method for RNA information for performing effective normalization processing when analyzing RNA information obtained from a biological sample which is a secretion derived from a subject.
Means for Solving the Problems
[0008] As a result of examining the data used for normalizing expression values for use in various statistical methods, using the expression state of RNA contained in SSL as sequence information, the inventors have found that effective normalization processing can be achieved by extracting RNA information by setting the threshold serving as the selection criterion for the sample to be analyzed and the threshold serving as the selection criterion for the gene to be analyzed within a specific range.
[0009] The present invention relates to the following (1) to (3). 1) A data processing method for analyzing RNA expression information obtained from biological samples, which are secretions collected from a plurality of subjects, the method comprising the following steps a) to d). a) Among the RNAs to be detected, RNAs with an expression level of zero or that can be regarded as zero are determined to be undetectable, the number of detectable RNAs is counted, and for each sample, a ratio 1 (TD value) of the number of detectable RNAs to the total number of RNAs to be detected is obtained. b) Excluding samples among the samples with a ratio 1 less than the threshold set within the range of 5 to 29%, and selecting samples to be analyzed. c) Based on the RNA expression information of the selected samples to be analyzed, for each RNA to be detected, a ratio 2 (SD value) of the number of samples with an expression level greater than the expression level that can be regarded as zero or zero to the total number of all samples to be analyzed is obtained. d) Excluding RNAs among the RNAs to be detected with a ratio 2 less than the threshold set within the range of 81 to 99%, and extracting the expression information of the remaining RNAs as the analysis target. 2) A method for correcting RNA expression values, which normalizes with respect to the total number of RNA expression information extracted by the method of (1). 3) A program for executing the data processing method of (1) or the correction method of (2), an information recording medium on which the program is recorded, a computing device for executing the program, and an RNA analysis data set obtained by the data processing method or the correction method.
Advantages of the Invention
[0010] According to the present invention, when comparing RNA expression profiles derived from a plurality of samples in a biological sample with many missing values and variations in RNA expression information, effective normalization processing becomes possible, and highly accurate and reproducible statistical analysis based on RNA information becomes possible.
Brief Description of the Drawings
[0011]
Figure 1
Embodiments for Carrying Out the Invention
[0012] In the method of the present invention, the "RNA" to be analyzed may be any RNA derived from a living body, and may be total RNA, mRNA, rRNA, tRNA, non-coding RNA, etc., but preferably mRNA.
[0013] The biological sample used in the method of the present invention is a secretion derived from a subject, and specifically includes samples containing sebum, saliva, nasal discharge, tears, sweat, urine, semen, vaginal fluid, amniotic fluid, milk, feces, etc. Among these, the method of the present invention is effective for applying to skin surface lipids (SSL) with many RNA information deficiencies and variations. "Skin surface lipids (SSL)" refers to the fat-soluble fraction present on the surface of the skin, and is sometimes called sebum. Generally, SSL mainly contains secretions secreted from exocrine glands such as sebaceous glands in the skin and exists on the skin surface in the form of a thin layer covering the skin surface. SSL contains RNA expressed in skin cells. Here, "skin" is a general term for a region including the stratum corneum, epidermis, dermis, hair follicles, and tissues such as sweat glands, sebaceous glands, and other glands, unless otherwise particularly limited.
[0014] For the collection of SSL from the skin of a subject, any means used for the recovery or removal of SSL from the skin can be adopted. Preferably, an SSL absorbent material, an SSL adhesive material, or an instrument for scraping off SSL from the skin can be used. The SSL absorbent material or SSL adhesive material is not particularly limited as long as it has an affinity for SSL, and examples include polypropylene and pulp. More detailed examples of the procedure for collecting SSL from the skin include a method of absorbing SSL into a sheet-like material such as blotting paper or blotting film, a method of adhering SSL to a glass plate, tape, etc., a method of scraping off and collecting SSL with a spatula, scraper, etc. To improve the adsorption of SSL, an SSL absorbent material pre-containing a highly fat-soluble solvent may be used. On the other hand, since the adsorption of SSL is inhibited when the SSL absorbent material contains a highly water-soluble solvent or moisture, it is preferable that the content of the highly water-soluble solvent or moisture is low. The SSL absorbent material is preferably used in a dry state. The site of the skin from which SSL is collected is not particularly limited, and examples include the skin of any part of the body such as the head, face, neck, trunk, hands, and feet. A site with a large amount of sebum secretion, such as the skin of the head or face, is preferable, and the skin of the face is more preferable.
[0015] The RNA-containing SSL collected from the subject may be stored for a certain period. The collected SSL is preferably stored under low-temperature conditions as soon as possible after collection in order to minimize the degradation of the contained RNA. The temperature condition for storing the RNA-containing SSL may be below 0°C, preferably -20 ± 20°C to -80 ± 20°C, more preferably -20 ± 10°C to -80 ± 10°C, still more preferably -20 ± 20°C to -40 ± 20°C, still more preferably -20 ± 10°C to -40 ± 10°C, still more preferably -20 ± 10°C, and still more preferably -20 ± 5°C. The storage period of the RNA-containing SSL under the low-temperature condition is not particularly limited, but preferably 12 months or less, for example, 6 hours or more and 12 months or less, more preferably 6 months or less, for example, 1 day or more and 6 months or less, still more preferably 3 months or less, for example, 3 days or more and 3 months or less.
[0016] In the method of the present invention, the method for obtaining the expression information of RNA is not particularly limited. For example, it can be obtained by converting RNA contained in a sample into cDNA by reverse transcription and then measuring the cDNA or its amplification product. As means for measuring the expression level, DNA chips, DNA microarrays, RNA-Seq, etc. can be mentioned, and RNA-Seq is preferably used. When using microarray analysis, the expression level of RNA is quantified by the signal intensity ratio, and in RNA-seq analysis, it is quantified by the number of sequence reads mapped to the genome (read count value).
[0017] The method of the present invention includes a step of obtaining information on the expression level of RNA, and includes a step of obtaining the above-mentioned quantified number of sequence reads (read count value) as the expression level of RNA. After that step, the data on the expression level of the RNA is stored in a server or a recording medium of a computer, input into the computer, and based on the input data, the data processing of the present invention can be executed by a program installed in the computer.
[0018] In the data processing method of the RNA information of the present invention, by setting a threshold value as a selection criterion for a sample to be analyzed and a threshold value as a selection criterion for a gene to be analyzed, the expression information of the RNA to be analyzed is extracted and normalization is performed. As shown in the examples described later, regarding the RNA expression level data (read count value by RNA-Seq) in a sample derived from a subject, the following examinations were conducted on the selection criteria for a sample (subject) to be analyzed and the selection criteria for a gene to be analyzed. As a selection index for a sample (j) to be analyzed, TD j value obtained by the following formula for each sample is used. The TD value is Targets Detected and corresponds to the gene detection rate (%).
[0019]
Equation
[0020] Here, the total number of genes to be detected refers to the total number of genes that can be theoretically detected in RNA expression analysis, and it can be appropriately determined based on the RNA expression analysis method used. In the case of the sequencing method (AmpliSeq) in the examples described later, it is determined based on the number of primer pairs in multiplex PCR. Also, the number of detectable genes can be calculated by subtracting the number of undetectable genes from the total number of genes to be detected. Here, the number of undetectable genes means the number of genes with zero or negligible expression.
[0021] On the other hand, for the selection of the gene (i) to be analyzed, the SD i value obtained by the following formula for each gene is used. The SD value is Samples Detected, and for each gene in the RNA expression level data of the samples to be analyzed after selection using the TD value, it is the ratio (detection sample rate) of the samples in which RNA expression from the gene could be detected. Here, that RNA expression could be detected means that the expression was detected exceeding zero or a negligible amount.
[0022]
Number
[0023] And samples (subjects) with TD j values less than 0%, 20%, and 30% are excluded, and the remaining samples (subjects) are selected as samples (subjects) to be analyzed, and then SD iGenes with values less than 70%, 80%, 90%, and 100% were excluded, and the remaining genes were selected as genes for data analysis. For the RNA expression level data extracted for these genes, normalization was performed using DESeq2 (Love MI et al. Genome Biol. 2014), and the degree of approximation to a normal distribution was verified. As a result, it was shown that by excluding samples with TD values less than 0% or less than 20% or less than 30%, and excluding genes with SD values less than 80% or less than 90% or less than 100%, it is possible to approximate more closely to a normal distribution in the normalization by DESeq2. However, in this case, it was shown that the number of samples for analysis can ensure about 80% of analyzable samples when excluding samples with TD values less than 20%, while it decreases to about 60% when excluding samples with TD values less than 30%. Also, the number of genes for analysis was slightly less than 20% of analyzable genes when excluding genes with SD values less than 90%, but it was shown to decrease to a few percent when excluding genes with SD values less than 100%.
[0024] Therefore, in the present invention, RNAs with an expression level of zero or that can be regarded as zero are determined to be undetectable, the number of detectable RNAs is counted, and for each sample, the ratio 1 (TD value) of the detectable RNAs to the total number of RNAs to be detected is obtained (step a). Samples with a ratio 1 less than the threshold value set within the range of 5 to 29% are excluded, and after selecting the samples for analysis (step b), for each RNA to be detected in the selected samples, the ratio 2 (SD value) of the number of samples with an RNA expression level higher than the expression level that can be regarded as zero or zero to the total number of all samples for analysis is obtained (step c). By excluding RNAs with a ratio 2 less than the threshold value set within the range of 81 to 99% and extracting the expression information of the remaining RNAs as the analysis target (step d), it can be said that effective normalization is possible in subsequent normalization processing.
[0025] In Project A, the RNAs with an expression level of zero or that can be regarded as zero can be appropriately determined by the measuring means. For example, in RNA-seq analysis, RNAs with a read count value of less than 20, preferably less than 15, more preferably less than 10 are exemplified.
[0026] In the selection of the analysis target sample in Project B, the threshold of Ratio 1 of the detectable RNAs to the total number of RNAs to be detected is set to 5% or more from the viewpoint of effective normalization, preferably 10% or more, more preferably 15% or more, still more preferably 18% or more. On the other hand, the threshold of Ratio 1 is set to 29% or less from the point of ensuring the number of analysis target samples in the analysis after normalization, preferably 27% or less, more preferably 25% or less, still more preferably 23% or less. Also, the threshold of Ratio 1 is appropriately set within the range of 5 to 29%, preferably within the range of 10 to 27%, more preferably within the range of 15% to 25%, still more preferably within the range of 18 to 23%. It is particularly preferable that the threshold of Ratio 1 is 20%.
[0027] In Project C, for each RNA to be detected, the ratio 2 (SD value) of the number of samples with an expression level greater than the expression level of zero or that can be regarded as zero to the total number of all analysis target samples is calculated. Here, the expression level that can be regarded as zero means, for example, in RNA-seq analysis, that the read count value is less than 5, preferably less than 3, more preferably less than 1. In the present invention, it is preferable to use, as the ratio 2 (SD value), the ratio of the number of samples with an expression level greater than zero (in RNA-seq analysis, the number of samples with a read count value greater than 0) to the total number of all analysis target samples.
[0028] In addition, in the selection of the RNA to be analyzed in step d, the threshold of ratio 2, which is the ratio of the number of samples with an RNA expression level greater than zero or an expression level that can be regarded as zero to the total number of samples, is set to 81% or more from the perspective of effective normalization, preferably 84% or more, more preferably 87% or more. On the other hand, the threshold of ratio 2 is set to 99% or less from the point of ensuring the number of genes to be analyzed in the analysis after normalization, preferably 96% or less, more preferably 93% or less. Also, the threshold of ratio 2 is appropriately set within the range of 81 - 99%, preferably within the range of 84 - 96%, more preferably within the range of 87 - 93%. It is particularly preferable that the threshold of ratio 2 is 90%.
[0029] When the threshold of ratio 1 in step b is low, it is desirable to increase the threshold of ratio 2 in step d for efficient normalization. When the threshold of ratio 2 in step d is low, it is desirable to increase the threshold of ratio 1 in step b for efficient normalization.
[0030] Thus, by performing normalization on the total number of expression information of the extracted RNA to be analyzed, effective correction of RNA expression values approximated to a normal distribution becomes possible. The normalization method used in this case is not particularly limited. For example, in addition to the RPM method and RLE method described above, the FPKM (fragments per kilobase of exon per million reads mapped) method, RPKM (reads per kilobase of exon per million reads mapped), TPM (transcripts per million) method, TMM (Trimmed mean of M values) method, etc. can be adopted, and the RLE method is preferably used. The RLE method is implemented in an analysis method for performing a series of gene expression level analyses called DESeq2.
[0031] A data processing method and a correction method for analyzing the above RNA expression information can be performed using a computer (computing device). That is, the present invention can provide a computing device for executing the above method, a program for causing the computer to execute the above method, and a computer-readable information recording medium on which the program is recorded. Further, the present invention can provide a data set for RNA analysis obtained by the above data processing method. Also, the present invention can input information such as ratio 1, ratio 2, or a threshold value used in the above data processing to perform data processing, or can select appropriate ratio 1, ratio 2, and threshold values by calculation.
[0032] The computing device of the present invention has means for inputting RNA expression information obtained from a sample collected from a subject, and according to a program for executing the data processing method and correction method of the present invention, includes one or more steps selected from the following steps: a step of selecting an analysis target sample, a step of selecting an analysis target gene, a step of extracting RNA expression information of the analysis target gene, and a step of normalizing the RNA expression information.
[0033] Examples of the computer-readable information recording medium on which the program for executing the data processing method and correction method of the present invention is recorded include a magnetic disk, an optical disk, a magneto-optical disk, and a flash memory. In the present invention, "computer-readable" shall include cases where it is distributed via an electric communication line or the like.
[0034] Aspects and preferred embodiments of the present invention are shown below. <1> A data processing method for analyzing RNA expression information obtained from secretions collected from a plurality of subjects as biological samples, the method comprising the following steps a) to d). a) Among the RNAs to be detected, RNAs with an expression level of zero or that can be regarded as zero are determined to be undetectable, the number of detectable RNAs is counted, and for each sample, a ratio 1 (TD value) of the number of detectable RNAs to the total number of RNAs to be detected is obtained. b) Excluding samples in which ratio 1 is less than the threshold value set within the range of 5 to 29%, and selecting samples to be analyzed c) Based on the RNA expression information of the selected samples to be analyzed, for each RNA to be detected, calculating ratio 2 (SD value), which is the ratio of the number of samples with an expression level greater than zero or an expression level that can be regarded as zero to the total number of samples to be analyzed d) Excluding RNAs in which ratio 2 is less than the threshold value set within the range of 81 to 99% among the RNAs to be detected, and extracting the expression information of the remaining RNAs as the objects of analysis <2>The method according to <1>, wherein the secretion is skin surface lipid <3>The method according to <1> or <2>, wherein the information on the expression level of the RNA in step a) is the read count value by RNA-Seq <4>The method according to any one of <1> to <3>, wherein the RNA with an expression level of zero or an expression level that can be regarded as zero in step a) is an RNA with a read count value by RNA-seq of less than 20, preferably less than 15, more preferably less than 10 <5>In step b), setting the threshold value of ratio 1 to preferably 10% or more, more preferably 15% or more, still more preferably 18% or more, and preferably 27% or less, more preferably 25% or less, still more preferably 23% or less, or setting it within the range of preferably 10 to 27%, more preferably 15% to 25%, still more preferably 18 to 23%, the method according to any one of <1> to <4> <6>The method according to any one of <1> to <4>, wherein in step b), the threshold value of ratio 1 is set to 20% <7>The method according to any one of <1> to <6>, wherein the expression level regarded as zero in step c) is a read count value in RNA-seq of less than 5, preferably less than 3, more preferably less than 1 <8>The method according to any one of <1> to <6>, wherein the sample with an expression level greater than zero or an expression level that can be regarded as zero in step c) is a sample with a read count value greater than 0 in RNA-seq <9>In step d), set the threshold of ratio 2 to preferably 84% or more, more preferably 87% or more, and preferably 96% or less, more preferably 93% or less, or preferably within the range of 84 to 96%, more preferably within the range of 87 to 93%, by any one of the methods <1> to <8>. <10>In step d), set the threshold of ratio 2 to 90%, by any one of the methods <1> to <8>. <11>A method for correcting RNA expression values, which normalizes with respect to the total number of RNA expression information extracted by any one of the methods <1> to <10>. <12>The method of <11>, which performs normalization by the RLE method. <13>A program for executing a data processing method or a correction method for analyzing any one of the RNA expression information of <1> to <12>. <14>An information recording medium characterized by recording the program of <13>. <15>A computing device including one or more steps selected from the step of selecting a sample to be analyzed, the step of selecting a gene to be analyzed, the step of extracting RNA expression information of the gene to be analyzed, and the step of calculating normalization of the RNA information of the gene to be analyzed, which are executed by the program of <13>. <16>An RNA analysis data set obtained by a data processing method or a correction method for analyzing any one of the RNA expression information of <1> to <12>.
Example
[0035] Hereinafter, the present invention will be described in more detail based on examples, but the present invention is not limited thereto. Example 1 Normalization of RNA expression data extracted from SSL 1) SSL collection After collecting sebum from the entire face of 42 healthy subjects (20 - 59-year-old women) using an oil-absorbing film, the oil-absorbing film was transferred to a vial and stored at -80°C for about 1 month until used for RNA extraction.
[0036] 2) RNA preparation and sequencing The defatted film in 1) above was cut into an appropriate size, and RNA was extracted using QIAzol Lysis Reagent (Qiagen) according to the attached protocol. Based on the extracted RNA, reverse transcription was performed at 42°C for 90 minutes using the SuperScript VILO cDNA Synthesis kit (Life Technologies Japan Co., Ltd.) to synthesize cDNA. The random primer attached to the kit was used as the primer for the reverse transcription reaction. From the obtained cDNA, a library containing DNA derived from 20,802 genes was prepared by multiplex PCR. Multiplex PCR was performed using the Ion AmpliSeq Transcriptome Human Gene Expression Kit (Life Technologies Japan Co., Ltd.) under the conditions of [99°C, 2 minutes → (99°C, 15 seconds → 62°C, 16 minutes) × 20 cycles → 4°C, Hold]. The obtained PCR product was purified with Ampure XP (Beckman Coulter, Inc.), and then buffer reconstitution, digestion of the primer sequence, adapter ligation and purification, and amplification were performed to prepare a library. The prepared library was loaded onto an Ion 540 Chip and sequenced using an Ion S5 / XL system (Life Technologies Japan Co., Ltd.).
[0037] 3) Data analysis In the RNA expression level data (read count values) derived from the subjects measured in 2) above, the selection criteria for the subjects to be analyzed and the selection criteria for the genes to be analyzed were examined. As the selection criteria for the subjects to be analyzed, the value of Targets Detected (TD) calculated in Torrent Suite (Life Technologies Japan Co., Ltd.) was used, and the TD calculated for each subject jThe thresholds were set at 0, 20, and 30%, and subjects with values below the thresholds were excluded from the analysis targets, while the remaining subjects were selected as the subjects for data analysis. As the extraction criteria for the genes to be analyzed, for each gene in the RNA expression level data after the selection of the subjects for data analysis using TD, the percentage of subjects (Samples Detected, SD) with read count values exceeding 0 was used, and the SD calculated for each gene to be detected i The thresholds were set at 70, 80, 90, and 100%, and genes with values below the thresholds were excluded from the analysis targets, while the remaining genes were selected as the genes for data analysis. After selecting the subjects for data analysis and subsequently extracting the expression information of the selected genes for data analysis, the logarithm to the base 2 value (Log2(normalized count + 1) value) obtained by adding the integer 1 to the normalized read count value (normalized count value) using the method called DESeq2 was calculated. Figure 1 shows the box plot of the Log2(normalized count + 1) values for each subject. Here, the TD j in subject j (j is an integer from 1 to n, where n is the number of subjects) and the SD i in gene i (i is an integer from 1 to m, where m is the number of genes to be detected) were calculated as follows.
[0038]
Equation
[0039] 4) Setting of the optimal selection criteria Regarding the Log2(normalized count+1) values calculated in 3) above, as a result of calculating the variance of the median, the variance of the median decreased to 0.1 or less as the thresholds of the TD value or the SD value increased (Table 1, bold). Also, a multiplicative decrease in the variance of the median accompanying the increase in the thresholds of the TD value and the SD value was confirmed. Therefore, it was shown that by selecting the subjects and genes to be analyzed using the TD value and the SD value, it is possible to align the medians of each subject after normalization by DESeq2. However, when subjects with a TD value of less than 20% were excluded, the number of analyzable subjects decreased to approximately 83%, while when subjects with a TD value of less than 30% were excluded, the number of analyzable subjects decreased to approximately 64% (Table 2). Since it is necessary to ensure the number of subjects to be analyzed in the analysis after normalization, it was shown that it is preferable to set 20% of the TD value as the threshold in the selection of subjects to be analyzed (Table 2, bold). Also, when genes with an SD value of less than 90% were excluded, the number of analyzable genes was approximately 16%, while when genes with an SD value of less than 100% were excluded, the number of analyzable genes decreased to 2% or 6% (Table 3). Since it is necessary to ensure the number of genes to be analyzed in the analysis after normalization, it was shown that it is preferable to set 90% of the SD value as the threshold in the selection of genes to be analyzed (Table 3, bold).
[0040]
Table 1
[0041]
Table 2
[0042]
Table 3
Claims
1. A data processing method executed by a computing device for analyzing RNA expression information obtained from biological samples, which are secretions collected from multiple subjects, the method comprising the following steps a) to d). a) Among the RNAs to be detected, RNAs with an expression level of zero or that can be regarded as zero are determined to be undetectable, the number of detectable RNAs is counted, and for each sample, a ratio 1 (TD value) of the number of detectable RNAs to the total number of RNAs to be detected is obtained. b) Among the samples, samples with a ratio 1 less than a threshold value set within the range of 5-29% are excluded, and samples to be analyzed are selected. c) Based on the RNA expression information of the selected samples to be analyzed, for each RNA to be detected, a ratio 2 (SD value) of the number of samples with an expression level greater than zero or that can be regarded as zero to the total number of all samples to be analyzed is obtained. d) Among the RNAs to be detected, RNAs with a ratio 2 less than a threshold value set within the range of 81-99% are excluded, and the expression information of the remaining RNAs is extracted as the analysis target.
2. The method according to claim 1, wherein the secretion is skin surface lipid.
3. The method according to claim 1 or 2, wherein the information on the expression level of the RNA in step a) is the read count value by RNA-Seq.
4. The method according to any one of claims 1 to 3, wherein the RNA with an expression level of zero or that can be regarded as zero in step a) is an RNA with a read count value less than 10 by RNA-seq.
5. The method according to any one of claims 1 to 4, wherein in step b), the threshold value of ratio 1 is set to 20%.
6. The method according to any one of claims 1 to 5, wherein the sample with an expression level greater than zero or that can be regarded as zero in step c) is a sample with a read count value greater than 0 in RNA-seq.
7. The method according to any one of claims 1 to 6, wherein in step d), the threshold value of ratio 2 is set to 90%.
8. An RNA expression value correction method executed by a computing device for normalizing with respect to the total number of RNA expression information extracted by the method according to any one of claims 1 to 7.
9. A program for executing the data processing method for analyzing RNA expression information according to any one of claims 1 to 7.
10. An information recording medium characterized by recording the program according to claim 9.
11. A computing device including a step of selecting a sample to be analyzed, a step of selecting a gene to be analyzed, a step of extracting RNA expression information of the gene to be analyzed, and a step of calculating normalization of RNA information of the gene to be analyzed, which are executed by the program according to claim 9.
Citation Information
Patent Citations
Method for preparing nucleic acid sample
WO2018008319A1
Method for determining the susceptibility of a patient suffering from proliferative disease to treatment using an agent which targets a component of the PD1 / PD-l1 pathway
WO2018229487A1
Method for preparing nucleic acid derived from skin cell of subject
WO2020091044A1