Methods for detecting primary immunodeficiency

Gene expression analysis and metagenomic profiling using RNA sequencing and linear mixed models offer a timely and cost-effective method for diagnosing primary immunodeficiency disorders, addressing the inefficiencies of current diagnostic methods by providing comprehensive immune system insights.

JP7894317B2Inactive Publication Date: 2026-07-23IMMUNOSIS P T Y LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
IMMUNOSIS P T Y LTD
Filing Date
2021-02-05
Publication Date
2026-07-23
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Current diagnostic methods for primary immunodeficiency disorders (PIDs) are cumbersome, costly, and time-consuming, often taking an average of five years from symptom onset to diagnosis, and lack functional information about the immune system, making timely and accurate diagnosis challenging.

Method used

A method involving gene expression analysis through RNA sequencing and linear mixed models, combined with metagenomic profiling, to determine PID susceptibility by analyzing transcriptome profiles and detecting functional deficiencies in the immune system.

Benefits of technology

Provides efficient and accurate PID diagnosis, reducing diagnostic time, lowering healthcare costs, and improving treatment decisions by incorporating functional immune system information in a single step.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007894317000018
    Figure 0007894317000018
  • Figure 0007894317000019
    Figure 0007894317000019
  • Figure 0007894317000020
    Figure 0007894317000020
Patent Text Reader

Abstract

The present invention relates to a method for determining whether a subject has a primary immune deficiency (PID) or is predisposed to developing a PID, the method comprising fitting a transcriptome profile of the subject using a linear mixed model to a PID prediction equation created by fitting a transcriptome relationship matrix generated from a set of reference transcriptome profiles of reference subjects with and without a PID to the linear mixed model, wherein the result of the prediction equation indicates whether the subject has a PID or is predisposed to developing a PID. The present invention relates to a method for creating a primary immune deficiency (PID) prediction equation for determining whether a subject has a PID or is predisposed to developing a PID, the method comprising fitting a transcriptome relationship matrix generated from a set of reference transcriptome profiles of reference subjects with and without a PID to a linear mixed model to create the PID prediction equation.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Field of Invention The present invention relates to a method for determining whether a subject has primary immunodeficiency (PID) or is susceptible to developing PID.

[0002] Cross-reference of prior applications This application claims priority from Australian Patent Application Publication No. 2020900337, the entire contents of which are incorporated by reference. [Background technology]

[0003] Background of the Invention Primary immunodeficiency disorders (PIDs) are a group of diseases caused by congenital defects in the immune system, with as many as 200 different causative mutations known. PIDs are characterized by recurrent severe infections that can be life-threatening. Effective treatments for PIDs are available, including hematopoietic stem cell transplantation, gene therapy, enzyme replacement therapy, and intravenous immunoglobulins. Early diagnosis is critically important for reducing disease-related morbidity, treatment costs, and improving patient outcomes. Although the clinical phenotypes and molecular basis of an increasing number of immunological defects in PIDs have become clearer, timely and accurate diagnosis remains necessary in clinical practice.

[0004] Because the clinical symptoms of PID are diverse and current diagnostic procedures are complex, it takes an average of five years from symptom onset to diagnosis. Current diagnostic procedures involve countless specialized, costly, and cumbersome functional tests, including lymphocyte proliferation and cytotoxicity assays, flow cytometry, serum immunoglobulin level measurements, complete blood counts, neutrophil function tests, and complement assays.

[0005] Several DNA sequencing techniques are being explored to aid in the diagnosis of PID. Targeted Sanger or other gene exon sequencing or gene typing are used to establish PID classifications and devise optimal treatment strategies. In selecting candidate genes to test, the individual clinical and immunological characteristics of each patient often serve as a guide. However, although it is generally a monogenic disease, more than 200 different causative mutations have been reported, and there are likely several hundred more, making it not always clear which gene (or specific mutation) should be evaluated. Furthermore, mutations in different genes can manifest as similar phenotypes (locus heterogeneity), while mutations in different parts of the same gene can manifest as distinct phenotypes (allelic heterogeneity).

[0006] Next-generation sequencing (NGS), including whole-genome sequencing (WGS) or whole-exome sequencing (WES), enables the simultaneous amplification and sequencing of millions of DNA fragments from a single subject within days. However, identifying causative mutations can be challenging due to the large number of nucleotide variants to measure, and the fact that novel variants detected by NGS often relate to poorly characterized genes or are difficult to interpret because their biological effects on protein function are unpredictable.

[0007] One recognized limitation of DNA sequencing is that it does not provide functional information about the performance of the immune system, which is key information needed for PID diagnosis.

[0008] Gene expression analysis can provide insights into the functional effects of PID mutations [1]. Salem et al 2014 described the PID mutation IRF8 K108ERNA sequencing of blood cells obtained from patients with PID revealed decreased expression of target genes regulated by IRF8 and a deficiency of cell-type specific transcripts that indicate cytopenia [1]. While gene expression analysis is a useful research tool, it is not used, or considered to be used, as a direct diagnostic method for PID on its own. Current diagnostic methods rely on cell-based functional information about the composition and performance of the immune system, combined with knowledge of causative mutations, if possible. Functional insights from gene expression analysis allow for the identification of a set of genes whose expression indicates PID and which, by expression, can distinguish PID from immunocompetent individuals, including those with other immune system disorders.

[0009] Importantly, a comprehensive analysis of a patient's gene expression levels as a composite phenotypic analysis has never been used, or even considered, as a direct diagnostic method for PID. RNA sequencing has the advantage of being able to provide a measure of immune cell composition and activity, and therefore has potential diagnostic capabilities.

[0010] As described above, the defining characteristic of PID is recurrent infections due to the immune system's inability to control microbial colonization and invasion. While the identification of specific pathogens is useful and may provide information for treatment, monitoring the composition of the commensal microbial community can also provide useful information for managing PID. Microbial communities are increasingly shown to have functional interactions with the immune system[2], including those in the skin of PID patients[3], and these patients appear to exhibit some fundamental differences. [Overview of the project] [Problems that the invention aims to solve]

[0011] There is a need for efficient and accurate PID diagnostic methods that can be deployed at a low cost. This would have a significant impact on public health by influencing treatment decisions, improving patient survival and quality of life, increasing the speed and timeliness of diagnosis, and consequently significantly reducing patient healthcare costs and the demand for expensive pathology testing services.

[0012] Any reference to prior art in this specification does not constitute an endorsement or implied that such prior art is part of the common technical knowledge in any jurisdiction, or that such prior art can be reasonably expected to be understood, considered relevant, and / or combined with other prior art portions. [Means for solving the problem]

[0013] Summary of the Invention The inventors provide a method for determining whether a subject has PID or is susceptible to developing PID. This method includes gene expression, i.e., transcriptome, and optionally RNA sequencing (RNAseq) of gene sequence mutations, and further includes detecting functional deficiencies of the immune system reflected in the transcriptome using a linear mixed model with RNA expression levels (not sequences or SNPs) as input, and optionally detecting specific PID sequence mutations. In addition, the inventors provide a method for determining whether a subject has PID or is susceptible to developing PID by using metagenomic profiling as a measure of the commensal microbial community structure in combination with RNA sequencing mixed model analysis.

[0014] Accordingly, in one embodiment, the present invention relates to a method for determining whether a subject has primary immunodeficiency (PID) or is susceptible to developing PID, Applying the transcriptome profile of the subject to the PID prediction equation created by applying the transcriptome relationship matrix generated from the reference transcriptome profile sets of reference subjects with and without PID to a linear mixed model A method is provided that where the result of the prediction equation indicates whether the subject has PID or is susceptible to PID.

[0015] In another aspect, the present invention is a method for determining whether a subject has primary immunodeficiency (PID) or is prone to developing PID, comprising: - Generating a transcriptome profile from a sample; and Applying the transcriptome profile of the subject to the PID prediction equation created by applying the transcriptome relationship matrix generated from the reference transcriptome profile sets of reference subjects with and without PID to a linear mixed model A method is provided that where the result of the prediction equation indicates whether the subject has PID or is susceptible to PID.

[0016] In a further aspect, the present invention is a method for determining whether a subject has primary immunodeficiency (PID) or is prone to developing PID, comprising: - Obtaining a sample from the subject; - Generating a transcriptome profile from the sample; and Applying the transcriptome profile of the subject to the PID prediction equation created by applying the transcriptome relationship matrix generated from the reference transcriptome profile sets of reference subjects with and without PID to a linear mixed model A method is provided that Here, the result of the prediction equation indicates whether the subject has PID or is likely to suffer from PID.

[0017] In one aspect, the present invention is a method for creating a primary immunodeficiency disorder (PID) prediction equation for determining whether a subject has PID or is likely to develop PID, comprising: - fitting a transcriptome relationship matrix generated from a set of reference transcriptome profiles of reference subjects with and without PID to a linear mixed model to create a PID prediction equation and providing a method including this.

[0018] In another aspect, the present invention is a method for creating a primary immunodeficiency disorder (PID) prediction equation for determining whether a subject has PID or is likely to develop PID, comprising: - generating a reference transcriptome profile from a reference subject; - generating a set of reference transcriptome profiles; and - fitting a transcriptome relationship matrix generated from a set of reference transcriptome profiles of reference subjects with and without PID to a linear mixed model to create a PID prediction equation and providing a method including this.

[0019] In another aspect, the present invention is a method for creating a primary immunodeficiency disorder (PID) prediction equation for determining whether a subject has PID or is likely to develop PID, comprising: - obtaining one or more samples from one or more subjects with and without PID; - generating a reference transcriptome profile from each subject; - generating a set of reference transcriptome profiles; and - Create a PID prediction equation by fitting the transcriptome relation matrix generated from reference transcriptome profile sets with and without PIDs to a linear mixed model. This provides a method that includes [something].

[0020] In any embodiment of the above method, the method further includes measuring or determining the transcriptome profile of a subject for whom PID or susceptibility to PID will be determined.

[0021] In any embodiment, the reference transcriptome profile set and / or the subject transcriptome profile on which PID or susceptibility to PID will be determined includes at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, or all 500 of the genes listed in Table 1, Table 2, or Tables 1 and 2.

[0022] In a preferred embodiment of the present invention, the linear mixed model is Best Linear Unbiased Prediction (BLUP), BayesR, or a machine learning method. In a further embodiment of the present invention, the machine learning method is one of ElasticNet, Ridge Regression, Lasso Regression, Random Forest, Gradient Boosting Machine, Support Vector Machine, Multilayer Perceptron (MLP), or Convolutional Neural Network (CNN).

[0023] In one embodiment of the present invention, the PID prediction equation provides an absolute prediction score. In one embodiment, an absolute prediction score greater than 0.2, greater than 0.4, greater than 0.6, or about 0.2, about 0.4, or about 0.6.

[0024] In one embodiment of the present invention, the PID prediction equation provides a relative prediction score, which is calculated by subtracting a healthy control score (determined from a known population of healthy controls) from a patient score (determined from a sample of the subject to be diagnosed). In one embodiment, the relative prediction score is greater than 0, greater than 0.1, greater than 0.2, or about 0, about 0.1, or about 0.2.

[0025] In one embodiment of the present invention, when detecting known PID gene mutations, an absolute prediction score close to 1.0, preferably 1.0, can be specified.

[0026] In any embodiment of the present invention, the PID prediction equation further provides a readout of PID gene mutations.

[0027] In any embodiment of the above method, the reference set further includes RNA sequence mutation profiles.

[0028] In any embodiment of the above method, the method further includes measuring or determining the RNA sequence mutation profile of a subject on which PID or susceptibility to PID will be determined.

[0029] In any embodiment of the above method, the transcriptome profile is used to provide further information about the defective pathways in subjects with PID. For example, a report may be generated stating that the patient has deficiencies in the Fc receptor signaling pathway, complement pathway, or interferon signaling pathway. This provides clinicians with information that can help prescribe treatment options.

[0030] In a preferred embodiment of the present invention, the mutation profile is a) RNA sequences of PID genes containing known mutations associated with, involved in, or causing PID; b) Novel mutations affecting the structure or function of proteins encoded by known gene mutations associated with, involved in, or causing PID, optionally including frameshift mutations, amino acid-algorithmic missense mutations, or nonsense stop codons; c) A dominant mutation in one allele that is associated with, involved in, or causes PID; d) Two different mutations located on two different alleles of the same gene that are associated with, involved in, or cause PID; e) Known mutations in RNA that are inferred or imputed by association with co-occurrence markers for mutations associated with, involved in, or causing PID; f) Absence of expression of a gene normally expressed in a non-PID subject that indicates a dysregulation or destabilizing mutation; g) A defective exon structure that indicates a splicing defect; h) One or more, optionally selected, additional mutations of 1 to 3 that are associated with, involved in, or cause PID; or i) Sequences of two or more other genes associated with, involved in, or causing PID severity, or sequences into which two or more other genes are imputed. Includes.

[0031] In any embodiment of the above method, the reference set further includes DNA sequence mutation profiles.

[0032] In any embodiment of the above method, the method further includes measuring or determining the DNA sequence mutation profile of a subject for which PID or susceptibility to PID will be determined. Preferably, a linear mixed model is used to fit the transcriptome profile and DNA sequence mutation profile of the subject to a PID prediction equation.

[0033] In any embodiment of the above method, the reference set further includes a metagenomic profile, and a linear mixed model is used to fit the transcriptome profile and metagenomic profile of interest to the PID prediction equation.

[0034] In a preferred embodiment of the above method, the metagenomic profile is obtained from an oral swab, nasal swab, pharyngeal swab, saliva, feces, or skin.

[0035] In a more preferred embodiment, the subject is a human.

[0036] In a further embodiment of the present invention, a method for determining whether a subject has or is susceptible to primary immunodeficiency (PID), comprising fitting the subject's metagenomic profile to a PID prediction equation created by fitting a metagenomic relation matrix generated from a reference set of metagenomic profiles of reference subjects with and without PID to a linear mixed model, wherein the result of the prediction equation indicates whether the subject has or is susceptible to PID.

[0037] It will be understood that transcriptome profiles or sequence mutation profiles are obtained from sputum, blood, amniotic fluid, plasma, semen, bone marrow, tissue, urine, ascites, or pleural fluid, and are optionally obtained by fine-needle biopsy.

[0038] Furthermore, it will be understood that transcriptome profiles or sequence (DNA and / or RNA) mutation profiles can be generated in vitro or ex vivo.

[0039] Furthermore, it will be understood that transcriptome profiles or sequence (DNA and / or RNA) mutation profiles can be generated in vitro, exovivo, or in silico.

[0040] In some embodiments of the above method, the method is not performed on the body of a human or animal.

[0041] In some embodiments of the above method, the method excludes any set of direct data collection performed on the human or animal body.

[0042] In a preferred embodiment of the above method, the blood includes peripheral blood mononuclear cells.

[0043] In any aspect or embodiment, the transcriptome, sequence (DNA and / or RNA) mutation profile, and metagenomic profile are determined from samples previously obtained from the subject.

[0044] In another embodiment, the present invention provides a method for determining whether a subject has or is susceptible to primary immunodeficiency (PID), comprising fitting the transcriptome profile of the subject to a PID prediction equation created by fitting a transcriptome relation matrix generated from a set of reference transcriptome profiles of reference subjects having and not having PID to a linear mixed model, wherein the result of the prediction equation indicates whether the subject has or is susceptible to PID.

[0045] In another embodiment, the present invention relates to a method for treating primary immunodeficiency (PID) in a subject who has or is prone to developing primary immunodeficiency (PID), - Determining whether a subject has PID or is susceptible to PID by performing or having performed the methods described herein; and - If the subject has PID, or is susceptible to PID, then administer a PID-specific therapy to the subject. This provides a method that includes [something].

[0046] In another embodiment, the present invention relates to a method for treating primary immunodeficiency (PID) in a subject who has or is prone to developing primary immunodeficiency (PID), - The method involves fitting a transcriptome relation matrix generated from a set of reference transcriptome profiles of reference subjects with and without PID to a linear mixed model to create a PID prediction equation, and then fitting the subject's transcriptome profile to this equation using a linear mixed model to determine whether the subject has PID, or whether it is susceptible to PID, and the result of the prediction equation indicates whether the subject has PID or is susceptible to PID. In cases where the subject has primary immunodeficiency (PID) or is prone to developing PID, the subject should be administered a PID-specific therapy. This provides a method that includes [something].

[0047] In another embodiment, the present invention provides the use of a primary immunodeficiency (PID)-specific therapy in the manufacture of a therapeutic agent for primary immunodeficiency (PID) in subjects having or being prone to developing primary immunodeficiency (PID), wherein the subjects are diagnosed by the method described herein.

[0048] In another embodiment, the present invention relates to a method for determining the effectiveness of PID therapy in a subject, - Provide a first sample obtained from the subject before receiving PID therapy; - Provide a second sample obtained from the subject during or after receiving PID therapy; - The method involves fitting a transcriptome relation matrix generated from a set of reference transcriptome profiles of reference subjects with and without PID to a linear mixed model to create a PID prediction equation, and then fitting the transcriptome profiles of the first and second samples of the subject to this linear mixed model, thereby indicating whether the subject has PID or is susceptible to PID based on the result of the prediction equation. Provides a method that includes, Here, the changes in transcriptome profiles from the first and second samples indicate the effectiveness of PID therapy in the subjects.

[0049] In one embodiment of the above method, PID therapy is administration of intravenous immunoglobulin (IVIG). In a further embodiment, intravenous immunoglobulin (IVIG) is administered at a dose of 200-800 mg / kg. In a further embodiment, the dose of intravenous immunoglobulin (IVIG) is administered every 3-4 weeks.

[0050] In another embodiment of the above method, PID therapy is the administration of subcutaneous immunoglobulin (SCIG). In a further embodiment, the subcutaneous (SCIG) is administered daily, weekly, or bi-weekly (every two weeks) in a dose calculated for each patient according to the manufacturer's instructions, taking into account their immunoglobulin trough concentration and previous IVIG dose.

[0051] In any embodiment of the above method, the primary immunodeficiency may be selected from any one of the following types: antibody production deficiency, combined immunodeficiency, phagocytic dysfunction, immunodysregulation, or complement deficiency. Preferably, the primary immunodeficiency is antibody production deficiency.

[0052] In any embodiment of the above method, primary immunodeficiency is X-linked agammaglobulinemia, unclassifiable immunodeficiency, selective immunoglobulin deficiency, Wiscott-Aldrich syndrome, severe combined immunodeficiency (SCID), DiGeorge syndrome, ataxia-telangectasia, chronic granulomatous disease, transient hypogammaglobulinemia in infancy, agammaglobulinemia, complement deficiency, selective IgA deficiency, IL-12 receptor deficiency, IL-12p40 deficiency, IFN-γ receptor deficiency, STAT1 deficiency, γc deficiency, Jak3 deficiency, RAG The following may be selected from the group consisting of 1 / 2 deficiency, ADA deficiency, X-linked hyper-IgM syndrome, MHC class II deficiency, Chediak-Higashi syndrome, deficiency of the initial components of the classical pathway (C1, C2, C4), deficiency of the initial components of the alternative pathway (factor D, factor P), deficiency of the membrane-invasive components (C5-C9), adenosine deaminase deficiency, autoimmune polyglandular endocrine disorder type 1 (APECED), Bloom syndrome, chondrohair follicle dysplasia, chronic granulomatous disease, familial atypical mycobacterial disease, hyperimmuneglobulin D syndrome, lymphoproliferative disorders, X-linked Nijmogen breakage syndrome, properzine deficiency, purine nucleoside phosphorylase deficiency, X-linked severe combined immunodeficiency, or any other primary immunodeficiency disorders described herein.

[0053] In one embodiment, the present invention provides a computer-implemented method for processing genomic information, wherein the genomic information includes a target transcriptome profile. - Accessing the reference transcriptome profile set for each reference subject, where each reference subject either has or does not have a primary immunodeficiency disorder (PID); - Generating transcriptome relation matrices from a reference transcriptome profile set; - Applying the transcriptome relation matrix to a linear mixed model to generate PID prediction equations; and - Fitting the target transcriptome profile to the PID prediction equation. This provides a method that includes [something].

[0054] In another embodiment, the present invention relates to a computer-implemented method for generating a predictive equation for primary immunodeficiency (PID), - Accessing the reference transcriptome profile set for each reference subject, where each reference subject either has or does not have a primary immunodeficiency disorder (PID); - Generating transcriptome relation matrices from a reference transcriptome profile set; and - Applying the transcriptome relation matrix to a linear mixed model to generate PID prediction equations. This provides a method that includes [something].

[0055] Any embodiment of the above method further includes measuring or determining the transcriptome profile of a subject for whom PID or susceptibility to PID will be determined.

[0056] In a preferred embodiment of the present invention, the linear mixed model is a machine learning technique including Best Linear Unbiased Prediction (BLUP), BayesR, Random Forest, or as defined herein.

[0057] In any embodiment of the above method, the reference set further includes RNA sequence mutation profiles.

[0058] In any embodiment of the above method, the method further includes measuring or determining the RNA sequence mutation profile of a subject on which PID or susceptibility to PID will be determined.

[0059] In any embodiment of the above method, the reference set further includes DNA sequence mutation profiles.

[0060] In any embodiment of the above method, the method further includes measuring or determining the DNA sequence mutation profile of a subject for which PID or susceptibility to PID will be determined. Preferably, a linear mixed model is used to fit the transcriptome profile and DNA sequence mutation profile of the subject to a PID prediction equation.

[0061] In any embodiment of the above method, the reference set further includes a metagenomic profile, and a linear mixed model is used to fit the transcriptome profile and metagenomic profile of interest to the PID prediction equation.

[0062] In a further embodiment of the present invention, a method for determining whether a subject has or is susceptible to primary immunodeficiency (PID), comprising fitting the subject's metagenomic profile to a PID prediction equation created by fitting a metagenomic relation matrix generated from a set of reference metagenomic profiles of reference subjects with and without PID to a linear mixed model, wherein the result of the prediction equation indicates whether the subject has or is susceptible to PID.

[0063] In another embodiment, the present invention provides a non-temporary computer-readable medium for storing instructions, wherein when an instruction is executed by a processor, - Accessing the reference transcriptome profile set for each reference subject, where each reference subject either has or does not have a primary immunodeficiency disorder (PID); - Generating transcriptome relation matrices from a reference transcriptome profile set; - Applying the transcriptome relation matrix to a linear mixed model to generate PID prediction equations; - To receive the target transcriptome profile; and - Fitting the target transcriptome profile to the PID prediction equation. It provides a non-temporary, computer-readable medium for storing instructions that cause the processor to perform certain actions.

[0064] In another embodiment, the present invention provides a non-temporary computer-readable medium for storing instructions, wherein when an instruction is executed by a processor, - Accessing the reference transcriptome profile set for each reference subject, where each reference subject either has or does not have a primary immunodeficiency disorder (PID); - Generating transcriptome relation matrices from a reference transcriptome profile set; and - Applying the transcriptome relation matrix to a linear mixed model to generate PID prediction equations. It provides a non-temporary, computer-readable medium for storing instructions that cause the processor to perform certain actions.

[0065] In a preferred embodiment of the present invention, the linear mixed model is a machine learning technique including Best Linear Unbiased Prediction (BLUP), BayesR, or as defined herein. In a further embodiment of the present invention, the machine learning technique is one of ElasticNet, Ridge Regression, Lasso Regression, Random Forest, Gradient Boosting Machine, Support Vector Machine, Multilayer Perceptron (MLP), or Convolutional Neural Network (CNN).

[0066] In any embodiment of a non-temporary computer-readable medium storing the above instructions, the reference set further includes RNA sequence mutation profiles.

[0067] In any embodiment of a non-temporary computer-readable medium storing the above instructions, the reference set further includes DNA sequence mutation profiles.

[0068] In any embodiment of a non-temporary computer-readable medium storing the above instructions, the reference set further includes metagenomic profiles, and a linear mixed model is used to fit the transcriptome profile and metagenomic profile of interest to the PID prediction equation.

[0069] When used herein, unless specifically required by context, the term "comprise" and its variations, such as "comprising," "comprises," and "comprised," are not intended to exclude any further additional elements, components, completes, or steps.

[0070] Further aspects of the present invention and further embodiments of the aspects described in the preceding paragraphs will become apparent from the following description, provided by reference to the accompanying drawings, for example. [Brief explanation of the drawing]

[0071] Brief explanation of the drawing [Figure 1] A schematic diagram illustrating the procedure for extracting RNA from blood. [Figure 2] A schematic diagram illustrating the procedure for generating an RNA sequence library. [Figure 3] This gene expression difference analysis compares 19,521 genes expressed in the blood of PID patients and corresponding normal controls. [Figure 4] Application of a predictive model using a leave-one-out prediction method. [Figure 5] Receiver Operating Characteristic (ROC) curve. [Figure 6] Four distinct genes that exhibit different levels of regulation in PID, i.e., are either upwardly or downwardly regulated. [Figure 7] An analysis demonstrating examples of significant differences in specific bacterial populations between PID patients and age- and sex-matched controls generated by the examples in this disclosure. [Figure 8]Expression of known PID genes in 15 patients (mean ± standard deviation). The PID genes shown are those of patients included in this study who had known mutations. [Figure 9] Mutation detection by whole blood RNA sequencing. RNA sequencing detection of dominant missense CXCR4 gene mutations in alleles present in PID patients. [Figure 10] A block diagram of a configurable computer processing system that implements various features of this disclosure. [Modes for carrying out the invention]

[0072] Detailed explanation There is a need to determine, detect, or diagnose a target PID in a timely and accurate manner. The present invention provides a method for predicting, determining, detecting, or diagnosing a target PID using RNA-seq, and optionally metagenomics and linear mixed models.

[0073] As used herein, “primary immunodeficiency” includes, but is not limited to, combined immunodeficiency disorders such as combined immunodeficiency; combined immunodeficiency with associated or symptomatic findings such as congenital thrombocytopenia; antibody-deficient predominance such as unclassifiable immunodeficiency; complement deficiencies such as C1q deficiency; congenital deficiencies of the number, function, or both of phagocytic cells such as severe congenital neutropenia; innate immunity deficiencies such as anhidrotic ectodermal dysplasia with immunodeficiency and autoinflammatory disorders such as familial Mediterranean fever; and immunomodulatory disorders such as familial hemophagocytic lymphohistiocytosis syndrome.

[0074] RNA sequencing offers at least three advantages over DNA analysis.

[0075] a) Mutation detection. An advantage of mutation detection in RNA compared to genomic DNA is that RNA sequences only contain expressed genes. Since these sequences do not contain the majority (98%) of the genomic sequence that is not expressed, the total amount of sequence generation required to identify mutations is reduced. This results in a significant reduction in nucleic acid complexity (and an increase in information density that increases throughput and efficiency), especially when methods that deplete highly expressed globin transcripts are applied before sequencing. Expressed gene sequences in blood are also accumulated with respect to expressed immunogenes, including their coding sequences. As a result, less total sequence information is needed to determine the mutation status. Since RNA of expressed and spliced ​​genes is thus accumulated from the genome, fewer sequences need to be obtained, and consequently, sequencing costs are reduced. Sequence information obtained from RNA is highly relevant and concentrated (and consequently the level of irrelevant sequence information is reduced), improving the reliability and efficiency of bioinformatics processing. Recently reported genome sequencing methods for PID [4] can be used to confirm or supplement RNA sequence information.

[0076] b) Advantages of RNA sequencing for measuring the integrity of PID gene transcripts. RNA sequencing is advantageous over DNA sequencing in that it can be used to identify RNA structural variants, such as splicing variants and intron expression at incorrect locations. RNA sequencing can also identify PID genes that are abnormally low in expression in the blood, for example, as a result of difficulty in identifying transcript defects, destabilizing mutations, or regulatory region mutations that interfere with gene expression. Sequences appearing in RNA can include coding RNA and non-coding RNA. Short-read NGS technology is well suited for this, however, long-read sequencing technologies such as Pacific-Biosciences (PacBio) SMRT and Oxford Nanopore are preferred and advantageous for measuring the presence and integrity of transcripts.

[0077] c) Advantages of RNA sequencing for measuring the composition and activity of immune cells. RNA sequencing is superior to DNA sequencing in terms of mutation detection as a component of PID determination, detection, or diagnosis. In addition, it provides functional information (not included in DNA sequences) because it includes gene activity, in this case a comprehensive measure of gene activity in immune cells in the blood. Holistic analysis of gene expression levels can be useful in identifying immunodeficiency because the expression of many genes measured in the blood or blood-derived cells can identify gene expression defects or abnormalities that occur in association with changes in immune cell populations and immune cell function. The inability of PID patients to overcome infections is a direct result of such changes in immune cell populations and immune cell function in the blood, and these changes are likely to be clearly visible in RNA transcript profiles.

[0078] Because various cell types are involved in PID disorders, and numerous immunogenes subsequently or secondarily influence them, whole-transcriptome methods such as Best Linear Unbiased Prediction (BLUP) or BayesR[5], modified to use read count or converted read count rather than SNP information (direct correction), are required. BLUP or BayesR require evaluation of distinguishing features in PID patients across the entire spectrum. Immunofunctional information provided in a single step by RNA sequencing (when captured by appropriate analysis) offers advantages in terms of cost, time, and resolution compared to a combination of immunological status assays typically required for the determination, detection, or diagnosis of PID, such as lymphocyte proliferation and cytotoxicity assays, flow cytometry, serum immunoglobulin level measurement, complete blood count, neutrophil function tests, and complement assays.

[0079] RNA-seq is useful for investigation purposes and is used in disease research, but several challenges prevent its use in clinical settings for determination, detection, or diagnosis, or for routine disease assessment [1]. The challenges in making whole transcriptome RNA expression information usable stem from the complexity of the information (data equivalent to thousands of genes) required to monitor expression levels for the determination, detection, or diagnosis of diseases such as PIDs, and the lack of knowledge about the relevant components of the information (specific genes and pathways, etc.). Furthermore, there is no suitable statistical analysis method for identifying and utilizing predicted mRNA biomarkers. Even if mRNA biomarkers can be identified from RNA-seq data, the lack of standardized RNA sequencing and defined statistical analysis limits its potential for clinical application.

[0080] More advanced techniques exist for DNA sequencing, providing more established guidelines and standards for mutation detection that complement clinical information on the immune system. While RNA sequencing is useful for detecting mutations in expressed gene sequences, functional information that can be provided by transcriptome sampling can also be utilized. BLUP or BayesR linear mixed model techniques provide analysis of transcript abundance information in RNASeq data, enabling its direct use as a diagnostic method. A limitation of using RNA sequencing alone without RNA expression BLUP or BayesR analysis is that while mutations in expressed gene sequences can be detected, the capture and use of functional information that can be provided by RNA sequence profiles / transcriptome data regarding the immune system is not fully realized.

[0081] The BLUP or BayesR model provides techniques that enable the incorporation of a vast number of effects, including minor effects (resulting from immune system deficiencies), into diagnostic analysis and evaluation. Because these techniques can capture a wide range of functional effects at the RNA level, the need for immunological clinical testing may be eliminated. Diagnostic methods (without using BayesR or BLUP) would typically attempt to identify key genes (in addition to PID genes) as functional markers that could potentially be used as alternatives to immunological clinical testing. For example, measuring transcripts of specific markers such as CD4, CD14, CD3, CD56, and CD19 could assess cellular compositional changes in PID. Similarly, other specific pathways or gene networks known to be affected in PID could also be used, either as individual tests, in combination with other tests, or by deriving a distinct set of genetic information from RNA-seq data. BLUP and BayesR offer a solution because they can be applied directly using whole RNA-seq information, thus allowing for the inclusion of numerous affected genes in the analysis, and enabling the measurement of numerous small effects expected to occur as a result of PID mutations.

[0082] The BLUP and BayesR methods proposed by the inventors are advantageous over other more targeted diagnostic marker methods because they directly utilize the maximum information from gene expression profiles as diagnostic signatures (using all genes expressed in the blood for analysis), in contrast to using one or a limited number of informative markers and / or known markers (even if they have been discovered and are available for use in PID diagnostic applications) as separate gene expression assays, or deriving specific information from RNA-seq data. In addition, the BLUP or BayesR methods are direct and efficient, requiring only a single computational step without human intervention or the need to combine analytical methods. The transcriptome BLUP or BayesR methods are also best suited to enabling the identification of overlapping immunodeficiency gene expression patterns across a range of patients, reflecting disease from diverse causative mutations. A more limited set of diagnostic gene markers (even if they are available) may not be able to identify the diversity of PID disease across a range of conditions. In addition, the BLUP / BayesR method, when trained with appropriate affected and unaffected patient reference profiles, can be effectively performed even without specific knowledge of all aspects of the functional changes measured for diagnosis, and thus can capture informative results about mutations that have not yet been elucidated, which can be used to aid in diagnosis.

[0083] The inventors have overcome the challenges of determining, detecting, or diagnosing PID by providing sequencing and whole transcriptome BLUP / Bayes® methodologies. This eliminates the need for functional testing required for PID diagnosis by providing a method for simultaneously assaying genomic information and immune cell function in a single step using molecular means. Paths toward improving functional testing largely involve expanding the range of cell types examined using antibody markers and FACS, as well as investigating dysfunctional cells examined under activated conditions.

[0084] RNA-seq is used not as a diagnostic tool, but as a research tool to identify genes and pathways related to immune function. In this case, researchers would likely begin by selecting specific genes from various analyses as candidates for monitoring and diagnosing immune function. For example, similar to RNA-seq applications in other diseases, PID target samples and normal target samples may be compared using various methods, and transcripts with differing expression levels would be identified as differences between the PID and normal target samples. Gene ontology enrichment analysis would likely be performed using tools such as the DAVID website (https: / / david.ncifcrf.gov / ). Differential gene expression profiles may also be subjected to gene set enrichment analysis (GSEA) using MSigDB publicly available immunogene signatures. Researchers would likely perform RNA-seq for research purposes, typically running it on a subset of blood cells to search RNA-seq data for known genes and pathways, or known cellular markers. The BLUP method for RNA sequencing from whole blood can capture information from known and poorly understood gene networks, capturing both direct and indirect effects. However, it has never been envisioned as a direct diagnostic method, nor as a substitute for a range of cell-based assays. There is no suggestion anywhere to directly use whole blood transcriptome BLUP as a diagnostic method to replace cell and immunoassays, including those for PID.

[0085] BLUP is used to classify samples into subsets, aiding in research and enhancing the genetic diagnosis (SNP mutations) of polygenic diseases. In some cases, combining diverse types of clinical information using BLUP can provide more accurate prognosis. The application of BLUP to disease classification has been applied to neuroblastoma [6].

[0086] To aid in diagnosis, other clinical information, including microbial colonization information, can be used in combination with information from the RNA-based methods described above. Infection record keeping and management, including, in some cases, microbial diagnostic techniques for pathogenic organisms, are important components of PID diagnosis.

[0087] Metagenomic sequencing extends the analysis of microbial composition beyond pathogens by providing information that includes a comprehensive measure of microbial community activity. Because the presence of numerous organisms in mucous membranes or hair follicles can identify deficiencies, abnormalities, or combinations of specific organisms in the community structure that occur in conjunction with changes in immune cell populations and immune cell function, a holistic analysis of microbial interface maintenance may be useful in identifying immunodeficiency.

[0088] As used herein, “RNAseq” or “transcriptome” refers to a gene that is expressed and then sequenced, whose sequence reads are aligned with the exon sequence of its genome or a reference transcriptome database. A “transcriptome profile” is a vector of sequence read counts and therefore the overall composition that characterizes the genes expressed in the sample.

[0089] The transcriptome relation matrix may be generated from the transcriptome profile as described in the examples, generated as part of the method of the present invention, or may already exist.

[0090] In one embodiment of the present invention, the linear mixed model is BLUP or BayesR. As used herein, the “linear mixed model” is also called a “multilayer model” or “hierarchical model” and refers to a class of regression models that take into account both the variability explained by the independent variable of interest and the variability not explained by the independent variable of interest, i.e., random effects. Examples of linear mixed models include, but are not limited to, BayesR and Best Linear Unbiased Prediction (BLUP). Those skilled in the art will be aware of other suitable linear mixed models.

[0091] In one embodiment, the PID prediction equation is one of those described herein, including in the examples.

[0092] The generated predictive scores (either relative or absolute) can be used to classify subjects as having a high risk of having PID (e.g., a higher score indicates a higher risk) or a low risk (e.g., a lower score indicates a lower risk). For example, when using absolute predictive scores, a score greater than 0.2 provides a diagnostic assay that detects PID with 93% sensitivity and 47% specificity. A score greater than 0.4 provides a diagnostic assay that detects PID with 73% sensitivity and 73% specificity. A score greater than 0.6 provides a diagnostic assay that detects PID with 53% sensitivity and 100% specificity. In contrast, for example, when using relative predictive scores (where the relative predictive score for each patient relative to the control group is determined by subtracting the healthy control score from the patient score), scores greater than 0, greater than 0.1, and greater than 0.2 provide diagnostic assays that detect PID with 93%, 80%, and 73% sensitivity, respectively.

[0093] In one embodiment of the present invention, the reference set further includes RNA sequence mutation profiles. In a further embodiment of the present invention, the reference set further includes RNA sequence mutation profiles, and a linear mixed model is used to fit the target transcriptome profile and RNA sequence mutation profile to the PID prediction equation.

[0094] In one embodiment of the present invention, the reference set further comprises DNA sequence mutation profiles. In a further embodiment of the present invention, the reference set further comprises DNA sequence mutation profiles, and a linear mixed model is used to fit the target transcriptome profile and DNA sequence mutation profile to the PID prediction equation.

[0095] In one embodiment of the present invention, the reference set further includes a metagenomic profile. In a further embodiment of the present invention, the reference set further includes a metagenomic profile, and a linear mixed model is used to fit the target transcriptome profile and metagenomic profile to the PID prediction equation.

[0096] The term “metagenome,” as used herein, refers to all DNA recovered from a sample, including DNA from the sample’s commensal microorganisms or “microbiome.” “Metagenome profile,” as used herein, refers to the overall composition that characterizes the microbial DNA in a sample. “Microbiome,” as used herein, refers to all microorganisms in a sample.

[0097] In one embodiment of the method of the present invention, the metagenomic profile is obtained from an oral swab, nasal swab, pharyngeal swab, saliva, feces, skin, or hair follicle. That is, the metagenomic profile is obtained from a sample containing the microbiome from an oral swab, nasal swab, pharyngeal swab, saliva, fecal sample, skin sample, or hair follicle sample.

[0098] The term “gene sequence mutation,” as used herein, encompasses both RNA and DNA sequence mutations and refers to a change in one or more nucleic acid molecules from their wild-type or reference sequence. “Mutation” includes, without limitation, at least one nucleotide base pair substitution, addition, and deletion of a nucleic acid molecule with a known sequence. Mutant nucleic acids may be expressed or found in one allele (heterozygous) or both alleles (homozygous) of a gene and may be somatic or germline. Therefore, a “gene sequence mutation profile” is the overall composition that characterizes gene sequence mutations in a sample.

[0099] Gene sequence mutations are also, a) If the RNA sequence of the PID gene is shown to have such a known mutation that causes PID; b) When a novel mutation (e.g., a missense mutation causing an amino acid change or a nonsense mutation causing a frameshift) is detected in a known PID gene from an RNA sequence that affects the predicted structure or function of the protein; c) When a dominant mutation is detected in one of the alleles from the RNA sequence; d) When two different mutations occur in the same gene, but on two different alleles; e) When a known mutation in RNA is inferred or imputed from the association with a haplotype marker co-occurring with RNA expressed from the same gene or adjacent genes on the chromosome; f) If the expression of a PID gene sequence that is normally expressed in the blood is not detected in blood RNA (indicating a serious dysregulation or destabilizing mutation); g) If there is a defect in the exon structure of the mutant PID gene as determined by RNA sequencing (indicating a splicing defect); h) If one or more (1-3) additional PID gene mutations are detected in the same patient from RNA / cDNA sequences; and i) When the sequences of several other genes detected in the RNA profile, or the sequences of other genes that are imputed, contribute to the severity of PID. It also includes.

[0100] In other words, in one embodiment of the method of the present invention, the mutation profile is a) RNA sequences of the PID gene, including known mutations that cause PID; b) Novel mutations, or optionally frameshift mutations, that affect the structure or function of a protein encoded by a known gene that causes PID; c) A dominant mutation in one allele that causes PID; d) Two different mutations that cause PID, located in the same gene but on two different alleles; e) Known mutations in RNA that are inferred or imputed by association with co-occurrence markers for mutations that cause PID; f) Absence of expression of a gene normally expressed in a non-PID subject that indicates a dysregulation or destabilizing mutation; g) A defective exon structure that indicates a splicing defect; h) One or more additional mutations, optionally 1 to 3, that cause PID; or i) Sequences of two or more other genes that contribute to PID severity, or sequences into which two or more other genes are imputed. Includes.

[0101] As used herein, “reference set” or “training set” refers to a set of transcriptome profiles, gene sequence mutation profiles, or metagenomic profiles obtained from subjects with and without PIDs, i.e., “reference subjects,” which are used to generate transcriptome relation matrices and subsequently to predict PIDs.

[0102] The term “marker” or “biomarker,” as used herein, refers to a biochemical, genetic (either DNA or RNA) or molecular feature that is a surrogate, and therefore indicates / predicts, a secondary feature, such as a genotype, phenotype, pathological condition, disease, or pathological state.

[0103] In one embodiment of the present invention, the transcriptome profile or sequence mutation profile is obtained from sputum, blood, amniotic fluid, plasma, semen, bone marrow, tissue, urine, ascites, or pleural fluid, and is optionally obtained by fine-needle biopsy. In a further embodiment, the blood includes peripheral blood mononuclear cells.

[0104] When used herein, "subject" can be a human or a non-human animal, such as livestock, zoo animals, or companion animals. In one embodiment, the subject is a mammal. The mammal may be an ungulate and / or, for example, an equid, bovine, sheep, canid, or feline. In one embodiment, the subject is a primate. In one embodiment, the subject is a human. Accordingly, the present invention has human medical applications, as well as veterinary and animal husbandry applications, including the treatment of livestock such as horses, cattle, and sheep, and companion animals such as dogs and cats.

[0105] Throughout this specification and the claims, the phrase “comprise” and its variations, such as “comprising” and “comprises,” mean “comprise, but not limited to,” and are not intended to exclude any other additional elements, components, completes, or steps.

[0106] As used herein, “determining whether a subject has PID or is susceptible to developing PID” means detecting or diagnosing PID in a subject, or predicting or prognosing the likelihood of a subject developing PID. The present invention also includes detecting PID in a subject or detecting a subject's susceptibility to PID. In other words, the present invention includes determining, detecting or diagnosing PID in a subject and / or determining, detecting or diagnosing a subject's susceptibility to PID.

[0107] The term “biological sample,” as used herein, refers to a sample that can be tested for a specific “gene expression profile,” “gene sequence mutation profile,” “transcriptome profile,” or “sequence mutation profile” (wherein sequence mutation profile may be mutations in RNA and / or DNA). A sample may be obtained from an organism (e.g., a human patient) or from a component of an organism (e.g., a cell). A sample may be from any relevant biological tissue or bodily fluid containing RNA and / or DNA. A sample may be a “clinical sample” that originates from a patient. Such samples include, but are not limited to, sputum, blood, blood cells (e.g., leukocytes), amniotic fluid, plasma, semen, bone marrow, and tissue or fine-needle biopsy samples, urine, ascites, and pleural fluid, or cells derived therefrom. Biological samples may also include tissue sections, such as frozen sections taken for histological purposes. A biological sample may also be referred to as a “patient sample.” In one embodiment, the methods of the present invention are not performed on the body of a human or animal, and, for example, the test profile may be determined by analyzing a pre-obtained biological sample.

[0108] As used herein, the term “gene” refers to a nucleic acid sequence containing a coding sequence necessary for the production of a polypeptide or precursor. Regulatory sequences that direct and / or control the expression of a coding sequence may also be included in the term “gene” in some examples. Polypeptides or precursors may be coded by a full-length coding sequence or by a portion of a coding sequence. A gene may contain one or more modifications in either its coding region or uncoding region that may affect the biological activity or chemical structure, expression rate, or mode of expression regulation of the polypeptide or precursor. Such modifications include, but are not limited to, mutations, insertions, deletions, and substitutions of one or more nucleotides, including single nucleotide polymorphisms that occur naturally in populations. A gene may constitute a continuous coding sequence, or it may contain one or more subsequences.

[0109] The term “gene expression level” or “expression level,” as used herein, refers to the amount of “gene expression product” or “gene product” in a sample. “Genetic expression profile” or “gene expression signature,” as used herein, refers to a group of “gene expression products” or “gene products” produced by a particular cell type or tissue type, the collective expression of those genes, or the differences in such gene expression, indicate and / or predict a pathological condition, disease, or pathology, such as an immune disorder. A “gene expression profile” may be qualitative (e.g., presence or absence) or quantitative (e.g., level or mRNA copy number). Therefore, a “gene expression profile” can also be used to determine the number of a particular cell type in a heterogeneous cell sample, such as the number of T cells in a blood sample, based on the amount of cell type-specific “gene expression product” or “gene product.”

[0110] The terms “gene expression product” or “gene product” as used herein refer to the RNA transcript of a gene, including mRNA, and the polypeptide translation product of such RNA transcript. “Genetic expression product” or “gene product” may be, for example, a polynucleotide gene expression product (e.g., unspliced ​​RNA, mRNA, splice variant mRNA, microRNA, fragmented RNA) or a protein expression product (e.g., mature polypeptide, splice variant polypeptide).

[0111] The term “immune cell” as used herein refers to cells such as lymphocytes, including natural killer cells, T cells, B cells, macrophages, and monocytes, dendritic cells, or any other cells capable of producing “immune effector molecules” in response to direct or indirect antigenic stimulation. The term “immune effector molecule” is a molecule produced in response to cell activation or antigenic stimulation, including, but not limited to, interferons (IFNs), interleukins (ILs), such as IL-2, IL-4, IL-10, or IL-12, tumor necrosis factor alpha (TNF-α), colony-stimulating factors (CSFs), such as granulocyte (G)-CSF or granulocyte-macrophage (GM)-CSF, complement, and components within the complement pathway.

[0112] The term "immunodeficiency," as used herein, refers to a pathological condition, disease, or pathological state characterized by dysfunction of the immune system. Examples of "immunodeficiency" include, but are not limited to, autoimmune disorders such as scleroderma, allergies such as allergic rhinitis, and immunodeficiency disorders such as primary immunodeficiency.

[0113] The term “normal immune system” as used herein means an immune system having a normal immune cell composition and in which the immune cells are not functionally abnormal. “Normal” or “healthy” subject as used herein means a subject having a “normal immune system.”

[0114] The term "nucleic acid," as used herein, refers to DNA molecules (e.g., cDNA or genomic DNA), RNA molecules (e.g., mRNA), DNA-RNA hybrids, and analogues of DNA or RNA produced using nucleotide analogues. Nucleic acid molecules may be nucleotides, oligonucleotides, double-stranded DNA, single-stranded DNA, multi-stranded DNA, complementary DNA, genomic DNA, non-coding DNA, messenger RNA (mRNA), microRNA (miRNA), nucleolar small RNA (snoRNA), ribosomal RNA (rRNA), transfer RNA (tRNA), small interfering RNA (siRNA), heteronuclear RNA (hnRNA), or small hairpin RNA (shRNA).

[0115] The method of the present invention may include a further step of treating PID in a subject who has been determined to have PID or to be susceptible to PID.

[0116] Accordingly, the present invention also discloses the treatment of PID in subjects determined to have or be susceptible to PID by the present invention.

[0117] Therefore, this specification includes methods for treating the target PID, Administering antibiotics, immunoglobulins, interferon, growth factors, gene therapy, or enzyme replacement therapy to the subject; or Transplanting hematopoietic stem cells into the target individual. A method including is disclosed, Here, the subject is determined to have PID or to be prone to developing PID by the method of the present invention.

[0118] Furthermore, the use of antibiotics, immunoglobulins, interferons, growth factors, enzymes, genes, or hematopoietic stem cells in the manufacture of a therapeutic drug for the subject PID is also disclosed, wherein the subject is determined to have or be susceptible to developing PID by the method of the present invention.

[0119] Also disclosed are antibiotics, immunoglobulins, interferons, growth factors, enzymes, genes, or hematopoietic stem cells for use in a method for treating a subject PID, wherein the subject is determined to have or be prone to developing PID by the method of the present invention.

[0120] In determining, detecting, or diagnosing PID, RNA sequencing offers three main advantages over DNA sequencing in terms of mutation detection and assessment of immune function: (a) mutations are detected only in expressed genes; (b) the integrity of PID gene transcripts; and (c) immune cell composition and activity.

[0121] In one embodiment, the PID to be treated is selected from combined immunodeficiency disorders such as combined immunodeficiency disorder; combined immunodeficiency disorders with associated or symptomatic findings such as congenital thrombocytopenia; antibody production deficiency-dominant types such as unclassifiable immunodeficiency disorder; complement deficiencies such as C1q deficiency; congenital deficiencies of the number, function, or both of phagocytic cells such as severe congenital neutropenia; innate immunity deficiencies such as anhidrotic ectodermal dysplasia with immunodeficiency and autoinflammatory disorders such as familial Mediterranean fever; and immunomodulatory disorders such as familial hemophagocytic lymphohistiocytosis syndrome.

[0122] Effective treatments for PID include managing infections, boosting the immune system, hematopoietic stem cell transplantation, gene therapy, and enzyme replacement therapy.

[0123] For the management of infectious diseases, • Treat infections with antibiotics, usually quickly and aggressively. Refractory infections may require hospitalization and intravenous (IV) antibiotics. • Preventing infections, for example, with long-term antibiotic treatment, to prevent respiratory infections and associated permanent lung and ear damage, and avoiding vaccination of children with PID using vaccines containing live viruses, such as oral polio and measles-mumps-rubella. Treating symptoms with medicinal substances such as ibuprofen for pain and fever, decongestants for sinus congestion, expectorants for thin mucus in the airways, or postural drainage that applies gravity and light blows to the chest to clear the lungs. It includes.

[0124] To boost the immune system, • Immunoglobulin therapy, usually administered intravenously every few weeks, or subcutaneously once or twice a week. • Gamma interferon therapy, which fights viruses and stimulates immune cells, is usually administered intramuscularly three times a week and is most often used to treat chronic granulomatous disease. • Growth factor therapy to increase white blood cell count It includes.

[0125] Stem cell transplantation offers a permanent cure for several forms of life-threatening PID.

[0126] Those skilled in the art will understand that the precise method of administering therapeutically effective doses of antibiotics, immunoglobulins, interferons, growth factors, hematopoietic stem cells, genes for gene therapy, or enzymes for enzyme replacement therapy to a patient will be at the discretion of the physician, based on the PID being treated or prevented. The mode of administration, including the dosage, combination with other drugs, timing and frequency of administration, may be influenced by the patient's expected response to treatment, as well as the patient's condition and medical history.

[0127] Antibiotics, immunoglobulins, interferons, growth factors, hematopoietic stem cells, genes for gene therapy, or enzymes for enzyme replacement therapy will be formulated, dosage-determined, and administered in accordance with best practices. Factors to be considered in this regard include the details of the PID being treated or prevented, the details of the patient being treated, the patient's clinical condition, the site of administration, the method of administration, the administration schedule, possible side effects, and other factors known to the physician. The effective therapeutic dose of antibiotics, immunoglobulins, interferons, growth factors, hematopoietic stem cells, genes for gene therapy, or enzymes for enzyme replacement therapy to be administered will depend on these considerations.

[0128] Antibiotics, immunoglobulins, interferons, growth factors, hematopoietic stem cells, genes for gene therapy, or enzymes for enzyme replacement therapy may be administered systemically or peripherally by routes including, for example, intravenous (IV), intra-arterial, intramuscular (IM), intraperitoneal, intracerebral (intracerobrospinal), subcutaneous (SC), intra-articular, intra-bursal, intrathecal, intracoronal, transendocardial, surgical transplantation, local, and inhalation (e.g., intrapulmonary).

[0129] The term "therapeutic dose" refers to the amount of antibiotics, immunoglobulins, interferons, growth factors, hematopoietic stem cells, genes for gene therapy, or enzymes for enzyme replacement therapy that is effective in treating the target PID.

[0130] The terms "to treat," "to treat," or "treatment" refer to both therapeutic measures and preventive or protective measures, the goal of which is to prevent or improve the patient's PID, or to slow (minimize) the progression of the patient's PID. Patients requiring treatment include those who already have PID as well as those for whom PID should be prevented.

[0131] The terms "prevention," "prevention," "preventive," or "prophylactic" refer to suppressing, preventing, defending against, or protecting against the occurrence of PID, including abnormalities or symptoms. Individuals requiring prevention may be prone to developing PID.

[0132] The term "improvement" refers to a reduction, decrease, or disappearance of PID, including abnormalities or symptoms. Individuals requiring improvement may already have PID, have a predisposition to developing PID, or be individuals who should be prevented from developing PID.

[0133] Figure 10 provides a block diagram of a computer processing system 500 that can be configured to implement embodiments and / or features described herein. System 500 is a general-purpose computer processing system. It will be understood that Figure 10 does not illustrate all functional or physical components of the computer processing system. For example, a power supply or power interface is not depicted, however, system 500 will either have a power supply or be configured to connect to a power supply (or both). It will also be understood that the specific type of computer processing system will determine the appropriate hardware and architecture, and that alternative computer processing systems suitable for implementing the features of this disclosure may have additional, alternative, or fewer components than those depicted.

[0134] The computer processing system 500 includes at least one processing unit 502 (for example, a general-purpose or central processing unit, a graphics processing unit, or an alternative computing device). The computer processing system 500 may include multiple computer processing units. In some examples, when the computer processing system 500 is described as performing calculations or functions, all processing necessary for the performance of those calculations or functions will be performed by the processing unit 502. In other cases, the processing necessary for the performance of those calculations or functions may also be performed by a remote processing device that is accessible to and available for use by the system 500 (either shared or exclusive).

[0135] The processing unit 502 communicates data with one or more computer-readable storage devices via a communication bus 504, which store instructions and / or data for controlling operations of the processing system 500. In this example, the system 500 includes system memory 506 (e.g., BIOS), volatile memory 508 (e.g., random access memory such as one or more DRAM modules), and non-volatile (or non-temporary) memory 510 (e.g., one or more hard disks or solid-state devices). Such memory devices may also be referred to as computer-readable storage media.

[0136] System 500 also includes one or more interfaces as substantially indicated by 512, through which System 500 interfaces with various devices and / or networks. Generally speaking, other devices may be integrated into System 500 or they may be separate. If a device is separate from System 500, the connection between that device and System 500 may be via wired or wireless hardware communication protocols, and may be direct or indirect (e.g., networked) connections.

[0137] Wired connections to other devices / networks may be made by any suitable standard or proprietary hardware connection protocol, such as Universal Serial Bus (USB), eSATA, Thunderbolt, Ethernet, HDMI, and / or any other wired connection hardware / connection protocol.

[0138] Wireless connectivity with other devices / networks may also be by any suitable standard or proprietary hardware / communication protocol, such as infrared, Bluetooth, Wi-Fi; near-field communication (NFC); ​​Pan-European digital mobile telephone system (GSM), Enhanced Data GSM Environment (EDGE), Long-Term Evolution (LTE), Code Division Multiple Access (CDMA and / or its variants), and / or any other wireless hardware / connection protocol.

[0139] Generally speaking, depending on the details of the system in question, the devices to which system 500 connects—whether by wired or wireless means—include one or more input / output devices (generally indicated by the input / output device interface 514). Input devices are used to input data into system 100 for processing by processing unit 502. Output devices enable system 500 to output data. Exemplary input / output devices are listed below, however, it should be understood that not all computer processing systems will include all the devices mentioned, and additional and alternative devices may also be used.

[0140] For example, System 500 may include, or be connected to, one or more input devices for inputting information / data into (for receiving by System 500). Such input devices may include a keyboard, mouse, trackpad (and / or other touch-sensitive / contact-sensitive devices including a touchscreen display), microphone, accelerometer, proximity sensor, GPS device, touch sensor, and / or other input devices. System 500 may also include, or be connected to, one or more output devices controlled by System 500 for outputting information. Such output devices may include devices such as displays (e.g., cathode ray tube displays, liquid crystal displays, light-emitting diode displays, plasma displays, touchscreen displays), speakers, vibration modules, light-emitting diodes / other lights, and other output devices. System 500 may also include, or be connected to, a device that can function as both an input and output device, such as a memory device / computer-readable medium (e.g., a hard drive, solid drive, disk drive, compact flash card, SD card, and other memory / computer-readable medium device) from which System 500 can read data and / or write data, as well as a touchscreen display that can both display data (output) and receive touch signals (input).

[0141] System 500 also includes one or more communication interfaces 516 for communicating with networks, such as the Internet 100 in the environment. Through the communication interfaces 516, System 500 can communicate data to and receive data from networked devices, which themselves may be other computer processing systems.

[0142] System 500 stores and accesses computer-readable instructions and data that, when executed by a computer application (also referred to as software or a program)—that is, by a processing unit 502—constitutes System 500 to receive, process, and output data. Instructions and data may be stored in a non-temporary computer-readable medium accessible to System 500. For example, instructions and data may be stored in non-temporary memory 510. Instructions and data may be transmitted to / received by data signals on a transmission path enabled by (e.g.) a wired or wireless network connection on an interface such as 512.

[0143] Applications accessible to System 500 typically include operating system applications such as Microsoft Windows®, Apple OSX, Apple iOS, Android, Unix, or Linux.

[0144] In some cases, some or all of a given computer-implemented method may be performed by the system 500 itself, while in other cases, the processing may be performed by other devices communicating data with the system 500.

[0145] Transcriptome differences constitute a genetic signature for a disease, which can be identified using learning software algorithms and proprietary reference databases.

[0146] The genome algorithm generates predictive scores that can be used to identify patients with the disease. The software components include a bioinformatics pipeline that searches for specific gene mutations already established as causing PID.

[0147] The present invention includes the method described herein and a program that can utilize a set of existing bioinformatics tools (including R and GATK for transcriptome relation matrices and BLUP prediction) for detecting mutations causing PID.

[0148] Herein, the present invention will be explained with reference to the following non-limiting examples. [Examples]

[0149] Examples An open-label exovivo study using samples from confirmed cases of primary immunodeficiency and healthy subjects. Research Overview This was an open-label, multicenter exovivo study using biological samples collected from 20 confirmed cases of primary immunodeficiency (PID) and 20 healthy subjects.

[0150] This study demonstrated that PID can be diagnosed using the following: (i) Gene expression data obtained by RNA sequencing, i.e., RNAseq or transcriptome; (ii) Gene expression data combined with gene sequence data, i.e., RNA-seq or transcriptome; (iii) Gene expression data, i.e., RNA-seq or transcriptome, obtained using linear mixed model prediction techniques that can be used alone for the diagnosis of PID and by targeted or non-targeted massively parallel sequencing, combined with microbial metagenomic data; or (iv) Microbial metagenomic data, It is obtained by targeted or non-targeted massively parallel sequencing and using linear mixed model prediction.

[0151] Exclusion criteria Subjects who had previously undergone hematopoietic stem cell transplantation were excluded from this study.

[0152] Sample collection Blood cells were collected from peripheral venous whole blood for RNA extraction. Microbial samples were collected from oral swabs, nasal swabs, pharyngeal swabs, saliva, fecal samples, skin samples, or hair follicle samples for DNA extraction.

[0153] Determining the transcriptome profile RNA sequencing was performed to identify gene sequence mutations that indicate PID, and the gene expression profiles of PID subjects were determined for comparison with normal subjects.

[0154] i) Sample collection and RNA extraction Blood cells were prepared from peripheral venous whole blood using PAXgene® blood RNA tubes (PAXgene Blood RNA Kit (50) - Catalog No. / ID: 762164) according to the manufacturer's instructions. The reagent composition of PAXgene Blood RNA tubes protects RNA molecules from degradation and can stabilize cellular RNA from human whole blood for up to 3 days at 18-25°C, up to 5 days at 2-8°C, or up to 8 years at -20°C / -70°C.

[0155] 2.5 ml of collected blood was placed in a PAXgene blood RNA tube and incubated at room temperature for at least 2 hours to ensure complete lysis of blood cells. When storing the PAXgene Blood RNA tube at 2-8°C, -20°C, or -70°C after blood collection, the sample was first equilibrated to room temperature, then stored at room temperature for 2 hours before proceeding. After preparing the buffer solution, the following steps were performed: (i) Centrifuge the PAXgene Blood RNA tube at 3000-5000 x g for 10 minutes using a swing-out rotor, and remove the supernatant. (ii) Add 4 ml of RN ase-free water to the tube and close it using the fresh BD Hemogard Closure included in the kit. (iii) Vortex until the pellet is visibly dissolved. Centrifuge at 3000-5000 × g for 10 minutes using a swing-out rotor and remove the supernatant completely. (iv) Add 350 μl of Buffer BR1 and vortex until the pellet is visibly dissolved. (v) Transfer the sample to a 1.5 ml Eppendorf test tube. Add 300 μl of lbuffer BR2 and 40 μl of proteinase K in succession. Mix by vortexing for several seconds. (vi) Incubate at 55°C for 10 minutes at 400-1400 rpm using a shaker incubator. (vii) Pipette the lysate directly into a PAXgene Shredder spin column (lilac) placed in a 2 ml collection tube, and centrifuge at maximum speed (but not exceeding 20,000 × g, as this may damage the column) for 3 minutes. (viii) Carefully transfer the entire supernatant of the flow-through fraction into a fresh 1.5 ml tube, taking care not to disturb the pellet in the processing tube. (ix) Add 350 μl of ethanol (96-100%, purity grade PA). Mix by vortexing and gently centrifuge to remove any droplets from the inside of the tube cap. (x) Pipette 700 μl into a PAXgene RNA spin column (red) placed in a 2 ml processing tube, and centrifuge at 16000 × g (8000~20000 × g) for 1 minute. Discard the flow-through. (xi) Pipette the remaining sample onto a PAXgene RNA spin column and centrifuge at 16,000 × g (8,000 to 20,000 × g) for 1 minute. Discard the flow-through. (xii) Wash the column with 350 μl of Buffer BR3. Centrifuge at 16,000 × g (8,000~20,000 × g) for 1 minute. (xiii) Add 80 μl of DNase I mix directly to the center of the PAXgene RNA spin column membrane and incubate at room temperature (20-30°C) for 15 minutes. (xiv) Pipette 350 μl of Buffer BR3 into a PAXgene RNA spin column and centrifuge at 16,000 × g (8,000-20,000 × g) for 1 minute. Discard the flow-through. Wash the (xv) column with 500 μl of BR4 and centrifuge at 16,000 × g (8,000 to 20,000 × g) for 1 minute. Discard the flow-through and centrifuge again at 16,000 × g (8,000 to 20,000 × g) for another minute. (xvi) Add another 500 μl of Buffer BR4 to the column and centrifuge at 16,000 × g (8,000-20,000 × g) for 3 minutes. Discard the processing tube containing the flow-through and place the PAXgene RNA spin column in a new 2 ml processing tube. Centrifuge at 16,000 × g (8,000-20,000 × g) for 2 minutes. Transfer the column to a 1.5 ml tube. (xvii) Add 40 μl of Buffer BR5 directly to the column membrane. Centrifuge at 16,000 × g (8,000-20,000 × g) for 2 minutes to elute the RNA. (Note: To achieve maximum elution efficiency, it is important to center the PAXgene RNA spin column so that the entire membrane is wetted with Buffer BR5.) (xviii) For example, quantify RNA / purity using a NanaDrop 1000 / 2000 or Qubit instrument and an RNA-specific binding fluorescent dye such as Quant-iT® RNA. (xix) For example, determine the integrity of the RNA using a BioAnalyser 2100 or TapeStation 2200 instrument (Agilent Technologies). (xx)If RNA samples are not to be used immediately, store them at -20°C or -70°C.

[0156] ii) RNA sequencing RNA-seq libraries were prepared using the TruSeq RNA sample preparation kit (Illumina) according to the manufacturer's protocol outlined in Figure 2.

[0157] The whole transcriptome sequencing library was prepared using Illumina's "TruSeq Stranded Total RNA Library Prep Kit with Ribo-Zero Globin Set" according to the manufacturer's instructions.

[0158] We pooled multiplexes of libraries, each containing one of 12 indexed adapters. Each pool was sequenced in a 101-cycle pair-end-run on a single flow cell lane of a HiSeq2000 sequencer (Illumina).

[0159] iii) Gene expression profile generation and sequence analysis 100-base paired-end reads generated by a HiSeq2000 sequencer (Illumina) were called using CASAVA v1.8 and output in fastq format. Sequence quality was evaluated using trimmomatic (v0.39), and low-quality bases and sequence reads were trimmed and filtered using a script. Bases with a quality score of less than 20 were trimmed from the 3' end of the reads. Reads with an average quality score of less than 20, or with N greater than 3, or with a final length of less than 35 bases were discarded. Only paired reads were retained for alignment.

[0160] After RNA sequencing, the unprocessed read sequences were trimmed using Trimmomatic software [7] to ensure the 3' end had minimal quality (at least a phred score of 30), adapter traces were removed, and the sequences were finally filtered to a minimum length of 32 bp. Alignment with Ensembl GRCh38.84 was performed using hisat2 (v2.1), or alternatively, UCSC hg19 reference genome (Illumina iGenomes) sequencing was performed using TopHat2 [8]. Lane merging and duplicate marking were performed using gatk (v4.1.2.0), and QC and quantification were performed with RNAseQC (v2.3.4) GENCODE v24 annotation modified according to the GTEx collapse gene model. After quantification of gene expression by counting the number of uniquely mapped reads, gene expression differences were performed using edgeR (v3.26.4) [9].

[0161] The quantification method involved using programs such as GTEx or HTSeq counts to aggregate the number of unprocessed mapped reads, thereby achieving gene-level quantification and exon-level quantification. This and similar alternative sequencing methods are outlined by Conesa et al[7]. Exon read counts with an expression level of at least 2 (counts per million reads (CPM)) were retained in at least one of the 20 samples. RNA profile normalization was performed using Bioconductor resources[8] and the EdgeR Bioconductor package[9] to adjust for sequencing depth and other variables. Figure 3 shows a difference in gene expression analysis comparing 19,521 genes expressed in the blood of PID patients and corresponding normal controls.

[0162] PID is generally a monogenic disorder, and the identification of known harmful homozygous mutations (in addition to clinical symptoms) is sufficient for diagnosing and recommending treatment for PID. Sequence analysis for mutation detection in PID genes was performed by comparing the RNA sequence reads described above with reference human genome or transcript references to identify known harmful mutations. Paired RNA reads were aligned with genomic exons using TOPHAT2[8], and only reads that fall within the gene exon boundaries as indicated by UCSC hg19 were used. Each pair of alignments from each individual was sorted and indexed using SAMtools[10, 11]. Using a list of known or suspected PID genes and known harmful mutations that fall within the gene exon boundaries as indicated by UCSC hg19 genome assembly from the PID Discovery Project

[12] , informative allele variants in individuals were extracted using the SAMtools mpileup function (version 0.1.14). Further methods for variant detection in RNA sequences are becoming available

[13] .

[0163] The RNA analysis pipeline can detect homozygous mutations using a set of PID genes and their known mutations, as well as mutations in other suspected genes

[12] . The pipeline can also detect heterozygous mutations that contribute to the disease phenotype, including dominant mutations, or combinations of different harmful mutations in two alleles of the same gene

[14] . In addition, the RNA analysis pipeline can detect variant SNPs that are closely linked to (and point to the underlying mutant haplotype of) the causative mutation that may contribute to the diagnosis

[15] . In some cases, SNP mutations in other parts of the genome may provide information about the expected severity of disease expression in various individuals caused by PID mutations.

[0164] PID diagnosis of a target using transcriptome best linear unbiased prediction From a reference transcriptome profile set from normal and PID patients, a transcriptome relation matrix was constructed, from which predictive equations were derived to develop predictive equations for PID diagnosis using transcriptome BLUP. The reference transcriptome profile set was used to construct the transcriptome relation matrix, as previously described for microbial molecular signatures

[16] . A transcriptome profile is a vector of sequenced read counts aligned with a collection of human gene (or exon) sequences in the UCSC hg19 genome or a reference human transcriptome database. These reads are generated by non-target sequencing of RNA-derived cDNA. These transcriptome profiles relate to the relative abundances of various mRNA species. The model used assumes a normal distribution, and therefore the transcriptome profiles are logarithmically transformed and standardized. Logarithmically transformed and standardized counts for gene (or exon) j of sample i, component x ij Several transcriptome profiles were combined from an n×m matrix X containing n samples and m genes. Genes with fewer than 10 reads aligned with these profiles were removed from the matrix before standardization. These profiles were compared to create a transcriptome relation matrix (calculated as G=XX' / m). Disease status was predicted using BLUP. A mixed model was fitted to this data: y=1 n μ + Zg + e. In the formula, y is the disease phenotype vector, with one record per sample. n is a vector of 1, μ is the overall mean, Z is the design matrix that assigns records to the sample, and g is the random effect estimator ~N(0,Gσ). 2 g ) The phenotype y was adjusted for other fixed effects such as age and sex before analysis. Using ASReml, σ was extracted from the data. 2 g To estimate the disease state of the sample (

number

number

[0165] Solving this equation yields estimators of the mean and residuals for each transcriptome profile, where

number

number

[0166] PID transcriptome profile prediction was performed using the free statistical software R (version 3.1.2; The R Foundation for Statistical Computing; http: / / www.r-project.org / ) with the package rrBLUP

[17] . The transcriptome relation matrix was fitted to BLUP and validated using a two-fold cross-validation with either PID and non-PID as the training or validation set, and an alternative procedure called the leave-one-out method, in which one individual is sequentially removed from the dataset and the disease predictor is estimated using the remaining data. Individuals that are predicted are always excluded from the training set. Figure 4 shows the results of applying the predictive model using the leave-one-out prediction method. Figure 5 is the ROC curve demonstrating the usefulness of this model. Table 1 lists the 500 predictive mutant genes used in the predictive model. Figure 6 shows four examples of individual genes that are upregulated or downregulated in PID.

[0167] [Table 1]

[0168] [Table 2]

[0169] [Table 3]

[0170] Increasing the reference number of transcriptome samples from both affected and unaffected individuals facilitates further training in transcriptome BLUP, leading to increased accuracy in prediction and diagnosis with each iteration.

[0171] [Table 4]

[0172] [Table 5]

[0173] Determination of metagenomic profiles Non-targeted, massively parallel sequencing of ribosomes or microbial DNA was performed to generate a reference PID metagenomic profile.

[0174] i) Sample collection and DNA extraction To obtain a microbiome profile, DNA was extracted from oral swabs and hair follicles using the DNA extraction kit described below.

[0175] Buccal swab sample collection: 1. Brush your teeth at some point within 4 hours before sample collection, and refrain from eating before (or after) sample collection. 2. The cheek mucosa was wiped by pressing a cotton swab or nylon brush against it. 3. Using sterile forceps, remove the cotton from the swab and place it in a tube containing the lysis buffer, or insert the head of a nylon brush directly into the tube containing the lysis buffer.

[0176] Extract DNA from an oral swab. material: 1. QIAamp DNA Mini Kit (Qiagen, Catalog Number / ID: 51304, Catalog Number / ID: 51306) 2.RNase A solution (R6148-25ml, Sigma) 3. Preparation of lysis buffer: 20 mg lysozyme in 25 mM Tris-HCl, pH 8.0 solution; 2.5 mM EDTA, pH 8.0 and 1% Triton X-100

[0177] protocol: Place the oral swab (cotton) into a 1.5 mL or 2 mL tube. Add 400 μl of lysis buffer (20 mg / ml lysozyme in 25 mM Tris-HCl, pH 8.0 solution; 2.5 mM EDTA, pH 8.0, and 1% Triton X-100). Mix by pressing the cotton several times and pipetting. Incubate at 37°C for 60 minutes. Add 40 μl of proteinase K (20 mg / ml) and 400 μl of Buffer AL. Mix thoroughly by vortexing for 10 seconds (Note: Do not mix proteinase K directly with Buffer AL). Gently centrifuge the tube to remove any droplets from the inside of the cap. Incubate at 55°C for 60 minutes. Vortex occasionally during incubation to disperse the sample. • Incubate at 80°C for an additional 15 minutes to inactivate proteinase K. Transfer the solution to a new tube. Use the tip of the pipette to firmly press the cotton against the tube and remove as much solution as possible. Add 8 μl of RNase A solution (R6148-25 ml, Sigma) at 37°C for 60 minutes. · Add 450 μl of ethanol (96 - 100%) to the sample and mix by pulse vortexing for 15 seconds. Briefly centrifuge the tube to remove the droplets on the inside of the lid. · Apply the mixture (including all precipitates, which may need to be divided into two portions) to a Mini spin column. Close the cap and centrifuge at maximum speed for 1 minute. · Add 500 μl of Buffer AW1 to the column. Close the cap and centrifuge at maximum speed for 1 minute, then discard the filtrate. · Add 500 μl of Buffer AW2 to the column. Close the cap and centrifuge at maximum speed for 1 minute, then discard the filtrate. · Transfer the column to a new 2 ml collection tube and centrifuge at maximum speed for 2 minutes. · Place the column in a 1.5 ml tube and add 50 - 100 μl of Buffer EB (Qiagen) or 10 mM Tris - HCl, pH 8.5 according to the required gDNA concentration. Incubate at room temperature for 1 - 3 minutes, then centrifuge at maximum speed for 2 minutes.

[0178] For skin sample collection, skin preparation instructions include refraining from bathing for 12 hours before all sample collections and refraining from using skin softeners or antibacterial soaps or shampoos. Sampling sites include the post - auricular fold, antecubital fossa, or volar forearm. Obtain bacterial swabs (using Epicentre swabs) and scraping specimens (using a sterile disposable surgical scalpel) from a 4 cm 2 area and incubate in enzyme lysis buffer and lysozyme as described above for oral swab samples.

[0179] Amplification of microbial DNA from oral swabs PCR 16S PCR Primers used: 341F / 806R primers covering the 16S V4 region. The primer sites are targeted by the "forward" primer 341F and the "reverse" primer 806R. In addition, include barcoding primers for Illumina MySeq sequencing (the following shaded part).

[0180] Illumina Multiplexing Read1 sequencing primer was added (806R Ad). [ka]

[0181] Illumina Multiplexing Read2 sequencing primer was added (341F Ad). [ka]

[0182] [Table 6]

[0183] [Table 7]

[0184] Sequencing of microbial DNA from oral swabs The standard Illumina protocol was used for MiSeq sequencing.

[0185] ii) Target and non-target ultra-parallel sequencing Library preparation for sequencing was performed using the indexing protocol with Illumina barcoding primers as described by the manufacturer. The index is a short third read from the sequencing run. Briefly, the DNA is sheared to 300 bp, an adapter is added by ligation, and then the index is added using PCR. The library is then quantified and pooled. Paired-end sequencing of genomic DNA was performed using a HiSeq2000™ sequencer. Sequencing reads were trimmed so that the average Phred quality score for each read was above 20. Reads with a trimmed length below 50 were discarded.

[0186] iii) Metagenome profiling analysis PID diagnosis of a target using best linear unbiased prediction of metagenomics Essentially, as previously described

[16] , a metagenomic relation matrix was constructed using a set of generated reference metagenomic profiles. A metagenomic profile is a vector of sequenced read counts aligned with a collection of 16S rRNA sequences or other available or generated reference sequence sets (referred to here as contigs) in a database. These reads were generated by non-targeted sequencing of microbial DNA or by sequencing of 16S ribosome sequences amplified by PCR from microbial DNA. These metagenomic profiles relate to the relative abundances of various microbial species. The model used assumes a normal distribution, so the metagenomic profiles are logarithmically transformed and standardized.

[0187] Logarithmically transformed and standardized counts for contig j of sample i, component x ij Several metagenomic profiles were combined from an n×m matrix X containing n samples and m contigs. Contigs with fewer than 10 reads to align with these profiles were removed from the matrix before standardization. These profiles were compared to create a microbiome relation matrix (calculated as G=XX' / m). Phenotypes were predicted using the best linear unbiased prediction. A mixture model was fitted to the data: y=1 n μ + Zg + e. In the formula, y is the clinical phenotype vector, with one record per sample. n is a vector of 1, μ is the overall mean, Z is the design matrix that assigns records to the sample, and g is the variation effect estimator ~N(0,Gσ). 2 g ) is. Using ASReml, σ from the data 2 g To estimate the phenotype of the sample (

number

number

[0188] Solving this equation yields estimates of the mean and residuals for each metagenomic profile, where

number

number

[0189] Microbiome profile predictions for PID were performed using the free statistical software R (version 3.1.2; The R Foundation for Statistical Computing; http: / / www.r-project.org / ), specifically the rrBLUP package

[17] . The metagenomic relation matrix was fitted to a best linear regression model (BLUP), validated using a two-fold cross-validation with either PID and non-PID as the training or validation set, and an alternative procedure called the leave-one-out method, which sequentially removes one individual from the dataset and estimates disease predictions using the remaining data. Individuals predicted were always excluded from the training set.

[0190] As described above, updating the reference number of microbiome samples from affected and unaffected individuals to train the metagenomics BLUP (or BayesR) increases predictive accuracy with each iteration. Figure 7 shows an analysis demonstrating an example of significant differences in specific microorganisms between PID patients and age- and sex-matched controls.

[0191] PID diagnosis of a target using a combination of RNA and metagenomic best linear unbiased prediction. Integrated (transcriptomics and metagenomics) predictions were performed using R statistical software. Twenty positive and 20 negative PID diagnoses, along with serum transcriptome and metagenomic profiles, were fitted to a linear regression model.

[0192] The Z matrix describing the RNA transcript abundances above is shown below as the metagenomic Z matrix. 1 We developed an extended relation matrix that can be combined with relation matrices: y=1 n μ + Zg + Z 1 g 1 +e

[0193] The integrated predicted PID disease phenotype was calculated by multiplying these coefficients by the serum transcriptome profile and metagenomic profile, respectively. Predictive accuracy was evaluated by the Pearson correlation "r," i.e., the correlation between the measured value and the predicted value.

[0194] These results demonstrate the following: Firstly, transcriptome profiles were able to predict PID in these situations; secondly, integrating transcriptomes with metagenomics information can improve predictive accuracy. Updating training transcriptomes and microbiome sample references, including both affected and unaffected individuals, will further improve predictive accuracy.

[0195] PID diagnosis based on gene sequence prediction PID is generally a single-gene disorder, and for the diagnosis and treatment recommendations of PID, in addition to clinical symptoms, the identification of known homozygous mutations is sufficient, and this information can be derived from RNA sequences as described above. In addition, when the expression of PID genes that are normally expressed in the blood is not detected in the blood, this also indicates a significant regulatory defect or destabilizing mutation, and RNAseq can directly reveal such significant expression defects, enabling diagnosis without confirmation of the causative mutation.

[0196] Other genomic variants of mRNA, such as SNP variants, having an association with defective immune genes in the population may be useful for prediction. Known mutations in mRNA may also be inferred or imputed from their association with haplotype markers that co-occur in mRNA expressed from the same gene or genes that are in proximity on the chromosome. This genomic information can be obtained from genomic sequences or RNA sequences, and by using this alone or in combination with transcriptome BLUP or transcriptome BayesR, diagnostic information is provided.

[0197] The manifestation of the same PID disease mutation varies among individuals

[18] , and various genomic variants may affect the severity of the disease. Measuring this contributing or predictive (protective) mutation may be useful in helping to predict low-severity or late-onset PID cases where the disease expression is attenuated. If sufficient patient samples are available, this type of genomic mutation detected through RNA sequences (and / or whole-genome or exome sequences) will be even more useful in helping to predict disease severity. These can also help to improve the diagnosis of PID cases with autoimmune findings

[19] or autoimmune symptoms.

[0198] The patients included in this study had known gene mutations causing the disease, identified by DNA sequencing. To determine whether the PID gene was transcribed in the blood at a sufficient level to identify the gene mutation at the mRNA level in RNA-seq data, the expression levels of the PID gene transcript were determined in several individuals. By examining the number of sequence reads covering the PID gene, it is possible to determine the likelihood of detecting the mutation in RNA. Figure 7 demonstrates the detection of sufficient gene expression for several PID genes in PID patients. Using the CXCR4 gene as a successful example of mutation detection, a dominant missense gene mutation causing the disease was identified using RNA sequencing (Figure 9). In whole blood RNA-seq data obtained from 41 PID patients, a total of 183 mRNA sequence reads covered the region where the mutant allele (indicated by the arrow) was observed in the CXCR4 mRNA sequence (83 copies), and a normal allele variant sequence (100 copies) was observed and determined. The A base variant at this location on chromosome 2 is a missense mutation that creates a stop codon in the coding sequence of CXCR4 (from arg to STOP), which is known to cause PID (Figure 9).

[0199] PID diagnosis of the target using various linear mixed model methods Another linear mixed model approach alternative to BLUP, such as BayesR, which is applied to genomic prediction based on whole-genome sequence variation as described by Kemper et al

[20] , can also be applied to transcriptome and / or metagenomics data in the same way as the method described above using BLUP, where the X matrix describes individuals by read counts per contig after normalization. The BayesR method assumes that the true effects of gene expression are derived from a series of normal distributions starting from zero variance and reaching medium to large variance. The advantage of BayesR over BLUP is that the effects of individual genes are not compressed as tightly towards the mean as in BLUP. The BayesR approach can also be extended to include known biological information such as immune system regulatory functions as described by MacLeod et al

[21] (BayesRC).

[0200] PID diagnosis of a subject by machine learning techniques Another approach alternative to the linear mixed model can also be applied to transcriptome and / or metagenomics data in a predictive manner similar to that described above using BLUP, enabling the classification and prediction of PID. Machine learning, support vector machines, and neural networks can provide an alternative approach to the linear mixed model that uses transcriptome and / or metagenomics data from patients and normal controls as input for patient classification and subsequent prediction model training. Similar approaches have been used to classify cancer patients into high-risk or low-risk groups and to develop prediction models to assist in prognosis determination

[22] , and investigations are underway

[23] regarding the use of tumor RNAseq data as input for this purpose.

[0201] The linear mixed model may be combined with a descriptive report from the subject, and such reports can be useful as they can assist clinicians who have obtained the information in evaluating immune system dysregulation, affected cells and pathways, disease states, and possibly preferred treatments.

[0202] To do this, genes with significantly different expression levels in a given PID patient (e.g., 20, 50, or even 100 DE genes) may be identified, and this set of DE genes may be subjected to a pathway overpopulation analysis (i.e., gene set enrichment analysis) such as the DAVID or Reactome program, from which a report on the affected pathways and cellular functions in the subject may be generated.

[0203] Further qualitative reporting to complement the diagnosis could provide clustering reports on how the subject's transcriptome compares to other patients in the database. This could be based on transcriptome relation matrices or other analyses of gene expression differences from the patient in question. It would be understood that patients with mutations in the same gene would form clusters that are close to each other based on their transcriptome. As larger patient databases become available, this clustering could help classify newly diagnosed patients into PID disease subtypes based on their transcriptome (complementing and complementing the mutation detection similarities found).

[0204] References 1. Salem S, Langlais D, Lefebvre F, Bourque G, Bigley V, Haniffa M, Casanova JL, Burk D, Berghuis A, Butler KM et al: Functional characterization of the human dendritic cell immunodeficiency associated with the IRF8(K108E) mutation. Blood 2014, 124(12):1894-1904. 2. Naik S, Bouladoux N, Wilhelm C, Molloy MJ, Salcedo R, Kastenmuller W, Deming C, Quinones M, Koo L, Conlan S et al: Compartmentalized control of skin immunity by resident commensals. Science 2012, 337(6098):1115-1119. 3. Oh J, Freeman AF, Park M, Sokolic R, Candotti F, Holland SM, Segre JA, Kong HH: The altered landscape of the human skin microbiome in patients with primary immunodeficiencies. Genome research 2013, 23(12):2103-2114. 4. Gallo V, Dotta L, Giardino G, Cirillo E, Lougaris V, D’Assante R, Prandini A, Consolini R, Farrow EG, Thiffault I et al: Diagnostics of Primary Immunodeficiencies through Next-Generation Sequencing. Frontiers in immunology 2016, 7:466. 5. Erbe M, Hayes BJ, Matukumalli LK, Goswami S, Bowman PJ, Reich CM, Mason BA, Goddard ME: Improving accuracy of genomic predictions within and between dairy cattle breeds with imputed high-density single nucleotide polymorphism panels. Journal of dairy science 2012, 95(7):4114-4129. 6. Zhang W, Yu Y, Hertwig F, Thierry-Mieg J, Thierry-Mieg D, Wang J, Furlanello C, Devanarayan V, Cheng J, Deng Y et al: Comparison of RNA-seq and microarray-based models for clinical endpoint prediction. Genome biology 2015, 16:133. 7. Conesa A, Madrigal P, Tarazona S, Gomez-Cabrero D, Cervera A, McPherson A, Szczesniak MW, Gaffney DJ, Elo LL, Zhang X et al: A survey of best practices for RNA-seq data analysis. Genome biology 2016, 17:13. 8. Huber W, Carey VJ, Gentleman R, Anders S, Carlson M, Carvalho BS, Bravo HC, Davis S, Gatto L, Girke T et al: Orchestrating high-throughput genomic analysis with Bioconductor. Nature methods 2015, 12(2):115-121. 9. Robinson MD, McCarthy DJ, Smyth GK: edgeR: a Bioconductor package for differential expression analysis of digital gene expression data. Bioinformatics 2010, 26(1):139-140. 10.Etherington GJ, Ramirez-Gonzalez RH, MacLean D: bio-samtools 2: a package for analysis and visualization of sequence and alignment data with SAMtools in Ruby. Bioinformatics 2015, 31(15):2565-2567. 11. Li H, Handsaker B, Wysoker A, Fennell T, Ruan J, Homer N, Marth G, Abecasis G, Durbin R: The Sequence Alignment / Map format and SAMtools. Bioinformatics 2009, 25(16):2078-2079. 12. Itan Y, Casanova JL: Novel primary immunodeficiency candidate genes predicted by the human gene connectome. Frontiers in immunology 2015, 6:142. 13. Sheng Q, Zhao S, Li CI, Shyr Y, Guo Y: Practicability of detecting somatic point mutation from RNA high throughput sequencing data. Genomics 2016, 107(5):163-169. 14. Lionakis MS: Genetic Susceptibility to Fungal Infections in Humans. Current fungal infection reports 2012, 6(1):11-22. 15. Hsu AP, Sampaio EP, Khan J, Calvo KR, Lemieux JE, Patel SY, Frucht DM, Vinh DC, Auth RD, Freeman AF et al: Mutations in GATA2 are associated with the autosomal dominant and sporadic monocytopenia and mycobacterial infection (MonoMAC) syndrome. Blood 2011, 118(10):2653-2655. 16. Ross EM, Moate PJ, Marett LC, Cocks BG, Hayes BJ: Metagenomic predictions: from microbiome to complex health and environmental phenotypes in humans and cattle. PloS one 2013, 8(9):e73056. 17. Endelman JB: Ridge Regression and Other Kernels for Genomic Selection with R Package rrBLUP. The Plant Genome 2011, 4(3):250-255. 18. Alcais A, Quintana-Murci L, Thaler DS, Schurr E, Abel L, Casanova JL: Life-threatening infectious diseases of childhood: single-gene inborn errors of immunity? Annals of the New York Academy of Sciences 2010, 1214:18-33. 19. Carneiro-Sampaio M, Coutinho A: Early-onset autoimmune disease as a manifestation of primary immunodeficiency. Frontiers in immunology 2015, 6:185. 20. Kemper KE, Reich CM, Bowman PJ, Vander Jagt CJ, Chamberlain AJ, Mason BA, Hayes BJ, Goddard ME: Improved precision of QTL mapping using a nonlinear Bayesian method in a multi-breed population leads to greater accuracy of across-breed genomic predictions. Genetics, selection, evolution : GSE 2015, 47:29. 21. MacLeod IM, Bowman PJ, Vander Jagt CJ, Haile-Mariam M, Kemper KE, Chamberlain AJ, Schrooten C, Hayes BJ, Goddard ME: Exploiting biological priors and sequence variants enhances QTL discovery and genomic prediction of complex traits. BMC genomics 2016, 17:144. 22. Kourou K, Exarchos TP, Exarchos KP, Karamouzis MV, Fotiadis DI Machine learning applications in cancer prognosis and prediction. Comput Struct Biotechnol J. 2014 Nov 15;13:8-17. 23. Han H, Liu Y. Transcriptome marker diagnostics using big data. IET Syst Biol. 2016 Feb;10(1):41-8.

Claims

1. A method for determining whether a subject has primary immunodeficiency (PID) or is susceptible to developing PID, comprising fitting the transcriptome profile of the subject to a PID prediction equation created by fitting a transcriptome relation matrix generated from a set of reference transcriptome profiles of reference subjects having and not having PID to a linear mixed model, wherein the result of the prediction equation indicates whether the subject has PID or is susceptible to PID.

2. A method for creating a primary immunodeficiency (PID) predictive equation for determining whether a subject has PID or is susceptible to developing PID, the method comprising fitting a transcriptome relation matrix generated from a reference transcriptome profile set of reference subjects having and not having PID to a linear mixed model to create the PID predictive equation.

3. The method according to claim 1 or 2, further comprising measuring the transcriptome profile of the subject.

4. The method according to any one of claims 1 to 3, further comprising measuring the transcriptome profile of the reference object.

5. The method according to any one of claims 1 to 4, wherein the linear mixed model is Best Linear Unbiased Prediction (BLUP), BayesR, or Random Forest.

6. The method according to any one of claims 1 to 5, wherein the reference set further comprises RNA sequence mutation profiles.

7. The method according to any one of claims 1 to 6, further comprising measuring the RNA sequence mutation profile of the subject, which will be used to determine PID or susceptibility to PID.

8. The method according to any one of claims 1 to 7, wherein the reference set further comprises RNA sequence mutation profiles, and the transcriptome profile and RNA sequence mutation profile of the subject are fitted to the PID prediction equation using the linear mixed model.

9. The method according to any one of claims 1 to 8, wherein the reference set further comprises a DNA sequence mutation profile.

10. The method according to any one of claims 1 to 9, further comprising measuring or determining the DNA sequence mutation profile of the subject on which PID or susceptibility to PID will be determined.

11. The method according to any one of claims 1 to 10, wherein the reference set further comprises DNA sequence mutation profiles, and the transcriptome profile and DNA sequence mutation profile of the subject are fitted to the PID prediction equation using the linear mixed model.

12. The aforementioned mutation profile a) RNA sequences of PID genes containing known mutations that cause PID; b) A novel mutation that affects the structure or function of a protein encoded by a known gene that causes PID; c) A dominant mutation in one allele that causes PID; d) Two different mutations that cause PID, located in the same gene but on two different alleles; e) Known mutations in RNA that are inferred or imputed by association with co-occurrence markers for mutations that cause PID; f) Absence of expression of a gene normally expressed in a non-PID subject that indicates a dysregulation or destabilizing mutation; g) A defective exon structure that indicates a splicing defect; h) One or more additional mutations that cause PID; or i) sequences of two or more other genes that contribute to PID severity, or sequences into which two or more other genes are imputed. The method according to any one of claims 6 to 11, including the method described in any one of claims 6 to 11.

13. The method according to any one of claims 1 to 12, wherein the reference set further comprises a metagenomic profile.

14. The method according to any one of claims 1 to 13, further comprising measuring or determining the metagenomic profile of the subject on which PID or susceptibility to PID will be determined.

15. The method according to any one of claims 1 to 13, wherein the reference set further comprises a metagenomic profile, and the transcriptome profile and metagenomic profile of the subject are fitted to the PID prediction equation using the linear mixed model.

16. The method according to any one of claims 1 to 15, wherein the transcriptome profile or sequence mutation profile is obtained from sputum, blood, amniotic fluid, plasma, semen, bone marrow, tissue, urine, ascites, or pleural fluid.

17. The method according to claim 16, wherein the blood comprises peripheral blood mononuclear cells.

18. The method according to claim 13 or 14, wherein the metagenomic profile is obtained from an oral swab, nasal swab, pharyngeal swab, saliva, feces, or skin.

19. The method according to any one of claims 1 to 18, wherein the subject is a human.

20. The method according to any one of claims 1 to 19, wherein the profile of the subject on which PID or susceptibility to PID is to be determined is determined or measured by analyzing a biological sample previously obtained from the subject.