METHOD FOR DETECTING MEMBRANE VESICLES, METHOD FOR SEARCHING FOR BIOMARKER SEQUENCES, APPARATUS FOR DETECTING MEMBRANE VESICLES, APPARATUS FOR SEARCHING FOR BIOMARKER SEQUENCES, PROGRAM, COMPUTER, AND METHOD FOR SEARCHING FOR DISEASE BIOMARKERS

The qPCR method simplifies the detection and quantification of membrane vesicles and biomarkers by using biomarker-specific primers, addressing the complexity of existing sequencing methods and enabling disease diagnosis.

JP7734991B2Active Publication Date: 2025-09-08NAT INST FOR MATERIALS SCI
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023548403
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-09-16
Filing Date
2022-09-01
Publication Date
2025-09-08
Estimated Expiration
2042-09-01

AI Technical Summary

Technical Problem

Existing methods for sequencing nucleic acids in membrane vesicles are complicated and difficult to quantify, making them unsuitable for disease diagnostic techniques like liquid biopsy.

Method used

A method using quantitative polymerase chain reaction (qPCR) with primers designed to contain biomarker sequences, selecting frequent regions with low similarity to non-host cell sequences, and aligning these regions to detect membrane vesicles and biomarkers.

Benefits of technology

Enables easy quantification and detection of membrane vesicles and biomarkers, facilitating disease diagnosis through liquid biopsy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007734991000001
    Figure 0007734991000001
  • Figure 0007734991000002
    Figure 0007734991000002
  • Figure 0007734991000003
    Figure 0007734991000003
Patent Text Reader

Abstract

The method for detecting membrane vesicles of the present invention comprises detecting membrane vesicles in a sample by qPCR using primers that are designed to contain, in one amplification product, a biomarker sequence that is a part of the base sequence of membrane vesicles produced by host cells. The biomarker sequence is selected in the following manner: obtaining the base sequence of membrane vesicles produced by the host cells; mapping the base sequence to a reference sequence derived from the host cells; among the mapped regions, selecting frequent regions where the detection frequency exceeds a standard; performing a homology search with the use of the frequent regions as query sequences; selecting, as candidate biomarker sequences, frequent regions where the detected similarity is lower to the homologous sequence derived from other than the host cells than the standard; aligning the candidate biomarker sequences to the homologous sequence; detecting a region where sequence conservation is lower than the standard; and then selecting this region as the biomarker sequence. Thus, membrane vesicles in a biological sample can be easily quantified.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method for detecting membrane vesicles, a method for searching for biomarker sequences, a membrane vesicle detection device, a biomarker sequence searching device, a program, a computer, and a method for searching for disease biomarkers. [Background technology]

[0002] In recent years, membrane vesicles secreted from cells have attracted attention as one of the factors that mediate pathogens and various diseases. It has been pointed out that these vesicles, which migrate into the bloodstream and are transported throughout the body, may affect a wide range of diseases, from digestive system diseases and heart diseases to neurological diseases such as Alzheimer's disease. Therefore, there is a demand for a technique for determining the sequence and quantification of nucleic acids contained in membrane vesicles contained in body fluids, etc.

[0003] As an example of such technology, Patent Document 1 describes a method for sequencing nucleic acids in exosomes and extracellular vesicles, which includes the following: "A method for sequencing nucleic acids from a biological sample: (a) preparing a biological sample; (b) contacting the biological sample with a solid capture surface under conditions sufficient to retain cell-free DNA and extracellular vesicles from the biological sample on or within the capture surface; (c) contacting a lysis reagent with the capture surface while the cell-free DNA and extracellular vesicles are present on or within the capture surface, thereby releasing DNA and RNA from the capture surface and producing a homogenate; and (d) extracting DNA, RNA, or both from the homogenate." (e) selectively removing ribosomal DNA or RNA sequences from the homogenate, or from the extracted DNA, the extracted RNA, or both; (f) reverse transcribing the RNA into cDNA; (g) constructing a double-stranded DNA library from the extracted DNA, the reverse-transcribed cDNA, or both the extracted DNA and the reverse-transcribed cDNA; (h) optionally amplifying DNA, RNA, or both DNA and RNA from the library; selectively enriching nucleic acid sequences from the cDNA or double-stranded DNA library; and (i) sequencing the library containing the cDNA, double-stranded DNA, or both cDNA and double-stranded DNA." [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Special Publication No. 2019-535307 Summary of the Invention [Problem to be solved by the invention]

[0005] Although the method described in Patent Document 1 can determine the nucleic acid sequence of extracellular vesicles contained in a sample, the procedure is complicated and it is difficult to quantify the membrane vesicles contained in the sample, making it difficult to apply the method to disease diagnostic techniques (such as liquid biopsy) that use membrane vesicles as an indicator.

[0006] Therefore, an object of the present invention is to provide a method for detecting membrane vesicles that can easily quantify membrane vesicles in biological samples. Another object of the present invention is to provide a method for searching for biomarker sequences, a membrane vesicle detection device, a biomarker sequence searching device, a program, a computer, and a method for searching for disease biomarkers. [Means for solving the problem]

[0007] As a result of extensive research into achieving the above object, the present inventors have found that the above object can be achieved by the following configuration.

[0008] [1] A method for detecting membrane vesicles in a biological sample by quantitative polymerase chain reaction using primers designed to contain a biomarker sequence, which is a part of the base sequence of membrane vesicles produced by a host cell, in one amplification product, wherein the biomarker sequence is selected by the following steps: obtaining the base sequence of the membrane vesicles produced by the host cell; mapping the base sequence to a reference sequence derived from the host cell; selecting a frequent region from the mapped region whose detection frequency exceeds a standard; performing a homology search using the frequent region as a query sequence, and selecting the frequent region whose similarity to the detected homologous sequence derived from a sequence other than the host cell is lower than a standard as a biomarker candidate sequence; and aligning the biomarker candidate sequence to the homologous sequence to detect a region whose sequence conservation is lower than a standard, and using this as the biomarker sequence.

[0009] [2] The method for detecting membrane vesicles described in [1], wherein the reference sequence is a genomic sequence of the host cell or a nucleic acid sequence derived from a virus that has infected the host cell.

[0010] [3] A method for detecting membrane vesicles described in [1] or [2], wherein the biological sample is collected from a subject, and the procedure further includes isolating and obtaining the membrane vesicles from the biological sample before obtaining the base sequence of the membrane vesicles. The type and amount of membrane vesicles contained in a biological sample are likely to reflect the physical condition of the subject from whom the biological sample was collected. In other words, if the subject is suffering from a certain disease (if the subject is a patient with that disease), the biological sample is likely to contain membrane vesicles corresponding to the disease and its progression. Biomarker sequences discovered using membrane vesicles isolated and obtained from such biological samples are more useful as markers for diagnosing the subject's physical condition (the type of disease and its progression), and the detected membrane vesicles can be used as a standard for its quantitative evaluation. The same applies to [6] described below.

[0011] [4] A method for detecting membrane vesicles, comprising: obtaining a base sequence of membrane vesicles produced by a host cell; mapping the base sequence to a reference sequence derived from the host cell; selecting a frequent region from the mapped region whose detection frequency exceeds a standard; performing a homology search using the frequent region as a query sequence and selecting the frequent region whose similarity to a detected homologous sequence derived from a sequence other than the host cell is lower than a standard as a biomarker candidate sequence; aligning the biomarker candidate sequence to the homologous sequence to detect a region whose sequence conservation is lower than a standard and designating it as a biomarker sequence; and detecting the membrane vesicles in a biological sample by quantitative polymerase chain reaction using primers designed to include the biomarker sequence in a single amplification product.

[0012] [5] The method for detecting membrane vesicles described in [4], wherein the reference sequence is a genomic sequence of the host cell or a nucleic acid sequence derived from a virus that has infected the host cell.

[0013] [6] The method for detecting membrane vesicles described in [4] or [5], wherein the biological sample is collected from a subject, and the method further comprises isolating and obtaining the membrane vesicles from the biological sample before obtaining the base sequence of the membrane vesicles.

[0014] [7] A method for detecting membrane vesicles according to any one of [1] to [6], wherein the biological sample is at least one selected from the group consisting of saliva, blood, serum, plasma, buffy coat, lymph, interstitial fluid, body cavity fluid, digestive fluid, sweat, urine, nasal discharge, tears, semen, vaginal fluid, amniotic fluid, milk, sputum, surgical irrigation fluid, feces, and swabs of skin or body mucosa. When the biological sample is selected from the above, it is more likely that membrane vesicles are contained in the biological sample. Also, when the subject suffers from a disease, it is more likely that the above-mentioned types of biological sample contain membrane vesicles specific to that disease. Therefore, when the biological sample is selected from the above, the membrane vesicle detection method of the present invention exhibits better effects as a liquid biopsy method.

[0015] [8] A method for detecting membrane vesicles described in any of [1] to [7], wherein the selection of the biomarker candidate sequence is to select the sequence with the smallest alignment score to the homologous sequence. By selecting the above sequences as candidate biomarker sequences, the sequence is more likely to contain a biomarker sequence that is a region of low sequence conservation, allowing for more efficient biomarker sequence discovery, and ultimately more efficient membrane vesicle detection.

[0016] [9] The method for detecting membrane vesicles according to any one of [1] to [8], wherein the frequently occurring region is a protein coding region. If the frequent regions are designated as protein coding regions, mapping and other processes in the subsequent steps can be carried out more smoothly.

[0017]

[10] The method for detecting membrane vesicles according to any one of [1] to [9], wherein the quantitative polymerase chain reaction is a real-time polymerase chain reaction. According to the above-mentioned method for detecting membrane vesicles, the amount of membrane vesicles in a biological sample can be quantified more accurately and quickly.

[0018]

[11] A method for detecting membrane vesicles according to any one of [1] to

[10] , wherein the standard for the detection frequency is a predetermined standard based on the significance level of a statistical test. According to the above-described method for detecting membrane vesicles, a frequently occurring region can be selected more easily.

[0019]

[12] A method for searching for biomarker sequences, comprising: obtaining a base sequence of membrane vesicles produced by a host cell; mapping the base sequence to a reference sequence derived from the host cell; selecting a frequent region from the mapped region whose detection frequency exceeds a standard; performing a homology search using the frequent region as a query sequence, and selecting the frequent region whose similarity to a detected homologous sequence derived from a sequence other than the host cell is lower than a standard as a biomarker candidate sequence; and aligning the biomarker candidate sequence to the homologous sequence to detect a region whose sequence conservation is lower than a standard, and selecting this as a biomarker sequence.

[0020]

[13] The method for searching for a biomarker sequence according to

[12] , wherein the reference sequence is a genomic sequence of the host cell or a nucleic acid sequence derived from a virus that has infected the host cell.

[0021]

[14] A method for searching for a biomarker sequence described in

[12] or

[13] , further comprising isolating and obtaining the membrane vesicles from a biological sample collected from a subject before obtaining the base sequence of the membrane vesicles. The type and amount of membrane vesicles contained in a biological sample are likely to reflect the physical condition of the subject from whom the biological sample was collected. In other words, if the subject is suffering from a certain disease, the biological sample is likely to contain membrane vesicles corresponding to the disease and its progression. Biomarker sequences discovered using membrane vesicles isolated from such biological samples are more useful as markers for diagnosing the subject's physical condition (the type of disease and its progression).

[0022]

[15] The method for searching for a biomarker sequence described in

[12] , wherein the biological sample is at least one selected from the group consisting of saliva, blood, serum, plasma, buffy coat, lymph, interstitial fluid, body cavity fluid, digestive fluid, sweat, urine, nasal discharge, tears, semen, vaginal fluid, amniotic fluid, milk, sputum, surgical irrigation fluid, feces, and swabs of skin or body mucosa. When the biological sample is selected from the above, it is more likely that membrane vesicles are contained in the biological sample. Also, when the subject is suffering from a certain disease, the above-mentioned types of biological sample are more likely to contain membrane vesicles specific to that disease. Therefore, when the biological sample is selected from the above, the obtained biomarker sequence can be more preferably applied to a liquid biopsy method.

[0023]

[16] The method for searching for a biomarker sequence according to any one of

[12] to

[15] , wherein the selection of the biomarker candidate sequence comprises selecting a sequence having the smallest alignment score to the homologous sequence.

[0024]

[17] The method for searching for a biomarker sequence according to any one of

[12] to

[16] , wherein the frequently occurring region is a protein coding region.

[0025]

[18] The method for searching for a biomarker sequence according to any one of

[12] to

[17] , wherein the criterion for the detection frequency is a criterion determined in advance based on the significance level of a statistical test.

[0026]

[19] A membrane vesicle detection device comprising: a base sequence analysis device for obtaining the base sequence of membrane vesicles produced by a host cell; a quantitative PCR device for detecting the membrane vesicles in a biological sample by quantitative polymerase chain reaction using primers designed to include a biomarker sequence, which is a part of the base sequence, in one amplification product; and a control device, wherein the control device comprises: a map unit that maps the base sequence obtained by the base sequence analysis device to a reference sequence derived from the host cell; a frequent region selection unit that selects, from the mapped region, a frequent region whose detection frequency exceeds a standard; a biomarker candidate sequence selection unit that performs a homology search using the frequent region as a query sequence and selects, as a biomarker candidate sequence, the frequent region whose similarity to the detected homologous sequence derived from a sequence other than the host cell that is lower than a standard; and a biomarker sequence determination unit that aligns the biomarker candidate sequence to the homologous sequence to detect a region whose sequence conservation is lower than a standard and selects it as the biomarker sequence.

[0027]

[20] A program that causes the control device of a membrane vesicle detection device, which has a base sequence analyzer, a quantitative PCR device, and a control device that is a computer, to execute the following steps: controlling the base sequence analyzer to obtain the base sequence of membrane vesicles produced by host cells; mapping the base sequence to a reference sequence derived from the host cells; selecting from the mapped regions a frequent region whose detection frequency exceeds a standard; performing a homology search using the frequent region as a query sequence and selecting as a biomarker candidate sequence the frequent region whose similarity to detected homologous sequences derived from other than the host cells is lower than a standard; aligning the biomarker candidate sequence to the homologous sequence to detect a region whose sequence conservation is lower than a standard and designating it as a biomarker sequence; and controlling the quantitative PCR device to detect the membrane vesicles in a biological sample by quantitative polymerase chain reaction using primers designed to include the biomarker sequence in one amplification product.

[0028]

[21] A biomarker sequence searching device comprising: a base sequence analysis device for obtaining the base sequence of membrane vesicles produced by a host cell; and a control device, wherein the control device comprises: a map unit that maps the base sequence to a reference sequence derived from the host cell; a frequent region selection unit that selects, from the mapped regions, a frequent region whose detection frequency exceeds a standard; a biomarker candidate sequence selection unit that performs a homology search using the frequent region as a query sequence and selects, as a biomarker candidate sequence, the frequent region whose similarity to a detected homologous sequence derived from a sequence other than the host cell is lower than a standard; and a biomarker sequence determination unit that aligns the biomarker candidate sequence to the homologous sequence, detects a region whose sequence conservation is lower than a standard, and selects this as a biomarker sequence.

[0029]

[22] A program for causing the control device of a biomarker sequence searching device having a base sequence analysis device and a control device that is a computer to execute the following steps: controlling the base sequence analysis device to obtain the base sequence of membrane vesicles produced by a host cell; mapping the base sequence to a reference sequence derived from the host cell; selecting from the mapped regions a frequent region whose detection frequency exceeds a standard; performing a homology search using the frequent region as a query sequence and selecting as a biomarker candidate sequence the frequent region whose similarity to a detected homologous sequence derived from a sequence other than the host cell is lower than a standard; and aligning the biomarker candidate sequence to the homologous sequence to detect a region whose sequence conservation is lower than a standard and selecting it as a biomarker sequence.

[0030]

[23] A computer for searching for biomarker sequences, the computer having: a mapping unit that maps the base sequence of membrane vesicles produced by a host cell to a reference sequence derived from the host cell; a frequent region selection unit that selects, from the mapped regions, a frequent region whose detection frequency exceeds a standard; a biomarker candidate sequence selection unit that performs a homology search using the frequent region as a query sequence and selects, as a biomarker candidate sequence, the frequent region whose similarity to a detected homologous sequence derived from a sequence other than the host cell is lower than a standard; and a biomarker sequence determination unit that aligns the biomarker candidate sequence to the homologous sequence, detects a region whose sequence conservation is lower than a standard, and selects the region as the biomarker sequence.

[0031]

[24] A program that causes a computer to execute the following steps: mapping the base sequence of membrane vesicles produced by a host cell to a reference sequence derived from the host cell; selecting from the mapped regions a frequent region whose detection frequency exceeds a standard; performing a homology search using the frequent region as a query sequence and selecting as a biomarker candidate sequence the frequent region whose similarity to a detected homologous sequence derived from a sequence other than the host cell is lower than a standard; and aligning the biomarker candidate sequence to the homologous sequence to detect a region whose sequence conservation is lower than a standard and designating it as a biomarker sequence.

[0032]

[25] A base sequence of a first membrane vesicle obtained from a first biological sample collected from a diseased patient is compared with a base sequence of a known cell or a genome sequence of a known virus to determine host cell A, which is the host cell of the first membrane vesicle; a base sequence of a second membrane vesicle obtained from a second biological sample collected from a healthy subject is compared with a base sequence of a known cell or a genome sequence of a known virus to determine host cell B, which is the host cell of the second membrane vesicle; and by comparing host cell A with host cell B, host cell C, which is the host cell with a high abundance ratio in the first biological sample, is identified. mapping the base sequence of membrane vesicles produced by the host cell C to a reference sequence derived from the host cell C; selecting a frequent region from the mapped region whose detection frequency exceeds a standard; performing a homology search using the frequent region as a query sequence, and selecting the frequent region whose similarity to the detected homologous sequence derived from a cell other than the host cell C is lower than a standard as a biomarker candidate sequence; and aligning the biomarker candidate sequence to the homologous sequence to detect a region whose sequence conservation is lower than a standard and designating it as a disease biomarker. The disease biomarker discovery method described above allows identification of sequences specific to membrane vesicles contained in biological samples obtained from patients with a target disease. The disease biomarkers identified by the method described above can be used in in vitro diagnostics such as liquid biopsy.

[0033]

[26] The method for searching for a disease biomarker described in

[25] , further comprising: selecting a host cell D from the host cells B that is a closely related species to the host cell C; mapping the base sequence or a fragment thereof of a membrane vesicle derived from the host cell D to the reference sequence; and excluding the mapped region from the biomarker candidate sequence. In the above-mentioned method for searching for disease biomarkers, sequence regions derived from membrane vesicles contained in biological samples collected from healthy individuals are excluded from the biomarker candidate sequences, thereby further improving the specificity of the disease biomarkers obtained. [Effects of the Invention]

[0034] The present invention provides a method for detecting membrane vesicles that can easily quantify membrane vesicles in a biological sample, a method for searching for biomarker sequences, a membrane vesicle detection device, a biomarker sequence searching device, a program, a computer, and a disease biomarker. [Brief explanation of the drawings]

[0035] [Figure 1] 1 is a flowchart showing a method for searching for biomarker sequences used in the membrane vesicle detection method of the present invention. [Figure 2] 1 is a flowchart showing the specific steps of the method for detecting membrane vesicles of the present invention. [Figure 3] 10 is a flowchart of a method for detecting membrane vesicles in another embodiment (variant) of the method for detecting membrane vesicles of the present invention. [Figure 4] 1 is a hardware configuration diagram of an embodiment of a membrane vesicle detection device of the present invention. [Figure 5] 1 is a functional block diagram of an embodiment of a membrane vesicle detection device of the present invention. [Figure 6] 10 is an operational flow of a control device of a membrane vesicle detection device. [Figure 7] FIG. 10 is a hardware configuration diagram of a second embodiment of a membrane vesicle detection device. [Figure 8] FIG. 10 is a functional block diagram of a second embodiment of a membrane vesicle detection device. [Figure 9] FIG. 10 is an operational flow diagram of the second embodiment of the membrane vesicle detection device. [Figure 10] FIG. 1 is a hardware configuration diagram of an embodiment of a biomarker sequence searching device of the present invention. [Figure 11] FIG. 1 is a functional block diagram of an embodiment of a biomarker sequence searching device of the present invention. [Figure 12] FIG. 1 is an operational flow diagram of an embodiment of a biomarker sequence searching device of the present invention. [Figure 13]FIG. 2 is a hardware configuration diagram of a computer according to an embodiment of the present invention. [Figure 14] FIG. 2 is a functional block diagram of an embodiment of a computer according to the present invention. [Figure 15] FIG. 1 is a flow diagram of a process for searching biomarker sequences by a processor in a computer embodiment of the present invention. [Figure 16] This is the result of mapping the base sequence within the membrane vesicle particle to the reference sequence. [Figure 17] This is the result of changing the vertical axis in FIG. 16 from the number of membrane vesicle particles to the "detection frequency," that is, the total number of membrane vesicle particles that had the sequence of that region. [Figure 18] A homology search was performed using the screened frequently occurring regions as a query, and the functional information tagged to each protein sequence information in the matching database was extracted using an analysis algorithm. This is a list of the results of detecting gene functions that were detected with statistically significant high frequency. [Figure 19] FIG. 1 is a flow chart of the method for discovering disease biomarkers of the present invention. [Figure 20] This is the result of comparing the base sequence profiles of membrane vesicles (bacterial origin) in samples from healthy individuals and periodontal disease patients. [Figure 21] This is the result of comparing the base sequence profiles of membrane vesicles (virus-derived) in samples from healthy individuals and periodontal disease patients. [Figure 22] FIG. 1 shows a procedure for determining a region (biomarker sequence) that is specifically detected in periodontal disease patients. [Figure 23] FIG. 1 shows a procedure for determining a region (biomarker sequence) that is specifically detected in periodontal disease patients. DETAILED DESCRIPTION OF THE INVENTION

[0036] The present invention will be described in detail below. The following description of the components may be based on a representative embodiment of the present invention, but the present invention is not limited to such an embodiment. In this specification, a numerical range expressed using "to" means a range that includes the numerical values ​​before and after "to" as the lower and upper limits.

[0037] [Terminology] As used herein, "nucleic acid" or "polynucleotide" refers to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. The term refers only to the primary structure of the molecule. Thus, the term includes double- and single-stranded DNA, and RNA.

[0038] As used herein, the term "membrane vesicle" refers, as one form, to "outer membrane vesicles (OMVs)," which are microvesicles produced by constricting a portion of the host cell membrane to the outside of the bacterial cell. Their size is not particularly limited, but is often 10 to 1,000 nm, often 10 to 300 nm in gram-negative bacteria, and often 50 to 150 nm in gram-positive bacteria.

[0039] "Membrane vesicles" also include outer membrane vesicles produced by host bacteria, "exosomes" produced and released by animal cells via the endocytic pathway, and "apoptotic bodies" produced by apoptosis.

[0040] As used herein, the term "host cell" refers to a cell that produced the membrane vesicles to be analyzed. For example, the term "host bacterium" refers to the bacterium that produced the membrane vesicles to be analyzed. The host cell is not particularly limited and may be any of prokaryotic cells, eukaryotic cells, or archaeal cells. Examples of the host cell include bacteria, archaea, yeast, plant cells, insect cells, and animal cells (e.g., human cells, non-human cells, non-mammalian vertebrate cells, and invertebrate cells).

[0041] The host cell may also be in a state of being infected with a virus. Membrane vesicles produced from such host cells may contain a base sequence derived from the virus (e.g., viral DNA). In this specification, there is no distinction between host cells and host cells infected with a virus, and they are simply referred to as "host cells." Generally, cells infected with a virus are sometimes called "hosts," but in this specification, "host" consistently refers to cells that have produced (or will produce) membrane vesicles, and the term "host cells" will be used regardless of whether they are infected with a virus.

[0042] As used herein, "detection of membrane vesicles" means at least determining the presence or absence of membrane vesicles, and preferably includes quantifying the number of membrane vesicles.

[0043] In this specification, "single particle analysis" refers to a method of isolating cells (bacteria), membrane vesicles, etc. one by one and analyzing their base sequences, and is also referred to as "single cell analysis," etc.

[0044] As used herein, the term "biological sample" refers to a specimen that may contain biological substances such as nucleic acids and proteins. In one form, a biological sample includes a body fluid collected from a subject. A biological sample may be a liquid or solid isolated from any location within the subject's body, for example, a peripheral location.

[0045] Examples of biological samples include saliva, blood, serum, plasma, buffy coat, lymph, interstitial fluid, body cavity fluid, digestive fluid, sweat, urine, nasal discharge, tears, semen, vaginal fluid, amniotic fluid, milk, sputum, surgical lavage fluid, feces, and wipes and swabs of skin or body mucosa, or combinations thereof, with saliva being one preferred form.

[0046] One of the excellent features of the membrane vesicle detection method of the present invention is that accurate test results (presence or absence of membrane vesicles and amount of membrane vesicles) can be obtained even if the amount of biological sample is small (i.e., even if the amount of membrane vesicles contained is small). From this perspective, the amount of biological sample required for membrane vesicle detection may be small.

[0047] In methods for determining the presence or absence of membrane vesicles in a sample by analyzing the nucleic acid components of membrane vesicles, a certain amount of membrane vesicles is required in the sample, and a certain amount of each specific type of membrane vesicle is required, resulting in a large amount of sample.

[0048] As will be described in detail later, the method for detecting membrane vesicles of the present invention is characterized by selecting a portion of a frequently occurring region in membrane vesicles as a biomarker sequence and quantifying it by quantitative PCR (qPCR, PCR stands for "Polymerase Chain Reaction"). Therefore, the amount of membrane vesicles contained in a biological sample can be significantly smaller than in conventional methods.

[0049] Specifically, the biological sample may be 0.01 to 30 mL. This amount is affected by the type of biological sample; for example, serum or plasma may be 0.01 to 5 mL, and urine may be 1 to 30 mL.

[0050] As used herein, the term "subject" refers to a human or an animal. In one embodiment, the subject is preferably a human or an animal with a disease (hereinafter also referred to as a "specific disease") whose onset, progression, etc. are associated with or are expected to be associated with the presence of specific host cells. By targeting a human or an animal with a specific disease, the amount of detected membrane vesicles may provide information for determining the progression, etc., of the specific disease. The "specific disease" is not particularly limited, but is preferably a bacterial and / or viral infection, and more preferably periodontal disease.

[0051] [Membrane vesicle detection method] Next, the method for detecting membrane vesicles of the present invention will be described. A first embodiment of the method for detecting membrane vesicles of the present invention comprises detecting membrane vesicles in a biological sample by quantitative polymerase chain reaction using primers designed to contain, in one amplification product, a biomarker sequence, which is a part of the base sequence of the membrane vesicles produced by a host cell. The biomarker sequence is selected by the following steps: obtaining the base sequence of the membrane vesicles produced by the host cell; mapping the base sequence to a reference sequence derived from the host cell; selecting from the mapped region a frequent region whose detection frequency exceeds a standard; performing a homology search using the frequent region as a query sequence; selecting, as a candidate biomarker sequence, the frequent region whose similarity to a detected homologous sequence derived from a sequence other than the host cell is lower than a standard; and aligning the candidate biomarker sequence to the homologous sequence to detect a region whose sequence conservation is lower than a standard, and using this as the biomarker sequence.

[0052] First, the biomarker sequences used in this method will be described. FIG. 1 is a flowchart showing a method for searching for biomarker sequences used in the present method. First, in step S11, a biological sample is collected from a subject. If the subject has a specific disease (if the subject is a patient with a specific disease), the obtained biological sample is likely to contain membrane vesicles that are closely related to the progression of the specific disease, among membrane vesicles that can be produced by host cells (typically host bacteria or virus-infected host cells).

[0053] When searching for biomarker sequences using the flow shown in Figure 1, collecting and using biological samples from subjects with a specific disease is advantageous in that it allows obtaining biomarker sequences for detecting the presence or absence and quantity of membrane vesicles that are likely to be associated with the specific disease.

[0054] It is generally believed that the types and amounts of membrane vesicles produced by a given host cell may vary depending on its growth environment, etc. However, there are still many unknowns regarding the conditions under which a given host cell produces what types of membrane vesicles. On the other hand, membrane vesicles isolated and obtained from biological samples of subjects with a specific disease are likely to contain membrane vesicles related to the progression of the specific disease.

[0055] One of the features of the present invention is that frequently occurring sequence regions are selected from all of a plurality of membrane vesicles according to the flow chart shown in Figure 1, and the biomarker sequences selected from these are used to detect membrane vesicles. According to the membrane vesicle detection method of the present invention, specific membrane vesicles corresponding to biomarker sequences contained in a biological sample can be detected even if the expression conditions or functions of individual membrane vesicles are unknown. This indicates the possibility that this method can be used to utilize membrane vesicles as markers for determining the progression of specific diseases. The membrane vesicle detection method of the present invention can be used as a liquid biopsy method. This point will be specifically explained in the examples below.

[0056] On the other hand, this screening method does not necessarily have to include step S11. In that case, the biological sample may typically be a cell population containing host cells, a culture medium of host cells, etc. In this case, step S11 may be omitted, and the subsequent step S12 may be performed.

[0057] Next, in step S12, membrane vesicles are separated and obtained from the biological sample. By separating and obtaining membrane vesicles from the biological sample, it becomes easier to obtain base sequences derived from the membrane vesicles (in other words, inside the membrane vesicles) more accurately. As mentioned above, nucleic acids in membrane vesicles (or derived from membrane vesicles) have great potential as indicators for selectively analyzing specific diseases. However, if we try to extract them directly from biological samples without a separation procedure, for example, nucleic acid molecules other than those from membrane vesicles may be mixed in, which can easily compromise the yield and integrity of the nucleic acids derived from membrane vesicles. In addition, isolating and obtaining membrane vesicles is preferable because it leads to the concentration of nucleic acids, contributes to improving detection sensitivity, and at the same time removes contaminants that may inhibit subsequent processes (e.g., amplification reactions).

[0058] The method for isolating and obtaining membrane vesicles from a biological sample is not particularly limited, and known methods can be used as appropriate. Specific methods are described, for example, in paragraph 0076 of JP-A-2019-513391.

[0059] For example, when the host cells are bacteria and the biological sample is a suspension containing a group of bacteria that may or may not contain the host bacteria and membrane vesicles, an example of a method for isolating and obtaining membrane vesicles will be described below. First, the bacterial cells can be roughly separated by removing them from the suspension by centrifugation.

[0060] Next, the separated centrifugal supernatant is filtered using a membrane filter (for example, pore size 0.1 to 0.9 μm), whereby the bacterial cells remaining in the supernatant can be easily removed. Since the centrifugal supernatant may contain contaminants such as proteins in addition to bacterial cells, it is also possible to use a method in which the supernatant is concentrated using a 100-200 kDa cutoff filter, and then the membrane vesicles are precipitated and collected by ultracentrifugation.

[0061] The precipitate obtained in this manner may contain structures of bacteria and the like (e.g., ciliates and flagella), and in order to remove these, further purification may be performed using, for example, density gradient centrifugation using sucrose or the like.

[0062] Furthermore, the present inventors have found that nucleic acid components not derived from membrane vesicles are often attached to the surface of membrane vesicles. Therefore, the separated membrane vesicles may be further treated with a nuclease to decompose these nucleic acid components. Known enzymes, such as deoxyribonuclease, can be used as such enzymes.

[0063] More specifically, a method for isolating and obtaining membrane vesicles from the culture medium of a single cell line will be described. First, once a desired growth phase has been achieved under predetermined conditions, the culture medium is placed in a centrifuge tube with a volume of 1 to 100 mL and centrifuged at 3000 to 9000 rpm and 1 to 20°C for 1 to 30 minutes.

[0064] Once the supernatant and precipitate have been separated, discard the precipitate, pass the supernatant through a 0.1 to 0.9 μm filter, place it in another centrifuge tube, and store it at 1 to 10°C.

[0065] The supernatant is then placed in an ultracentrifuge tube and ultracentrifuged, for example, at a centrifugal force of 150,000 to 300,000 × g at 1 to 10°C for 1 to 3 hours. After ultracentrifugation, the supernatant is discarded, and the precipitated OMVs are resuspended in a buffer solution (e.g., PBS buffer at pH 7.4) and stored at 1 to 10°C. This ultracentrifugation step may be repeated multiple times.

[0066] Deoxyribonuclease (DNase) may be added to this suspension to decompose the nucleic acids attached to the membrane vesicles. By performing the deoxyribonuclease treatment, the incorporation of contaminants (nucleic acid components) is suppressed, allowing for more accurate measurements. The amount of deoxyribonuclease to be added is not particularly limited, but for example, 10 μL of the above suspension may be diluted 10-fold, and 2 μL of DNase (2000 U) may be added thereto.

[0067] The suspension may also be diluted to adjust the number (concentration) of membrane vesicles contained per unit volume of the suspension. By diluting the suspension and adjusting the concentration of membrane vesicles, the possibility of two or more membrane vesicles being encapsulated in one droplet becomes lower. In other words, it becomes easier to encapsulate each individual vesicle in a droplet more reliably.

[0068] The method for diluting the suspension is not particularly limited, but for example, the suspension may be diluted to an appropriate concentration while observing the concentration of membrane vesicles in the suspension by dynamic light scattering or the like. A preferred method for observing the concentration of membrane vesicles is to irradiate the particles with a laser, track the Brownian motion of each particle from the scattered light (tracking method), and calculate the particle diameter and number from the diffusion rate based on the Stokes-Einstein equation.

[0069] Next, in step S13, the base sequence of the separated and obtained membrane vesicles (derived from the membrane vesicles) is obtained. There are no particular limitations on the method for obtaining the base sequence derived from the membrane vesicles, and known methods can be applied. Typically, membrane vesicles are lysed to extract polynucleotides, which are then amplified as necessary, followed by library preparation, sequencing, and analysis to obtain the base sequence. In this case, the method for obtaining the base sequence of the membrane vesicles is not particularly limited, and known methods such as single particle analysis and shotgun metagenomic analysis can be used.

[0070] Next, in step S14, the obtained nucleotide sequence is mapped to a reference sequence derived from the host cell. Mapping can be performed using known mapping software. Examples of such software include "Bowtie2" and "BWA."

[0071] The degree of identity of the obtained base sequence to the reference sequence is not particularly limited, but when the maximum value of the mapping quality score provided by the software (e.g., as a SAM file) is taken as 100%, it is preferably 80% or more, more preferably 90% or more, and even more preferably 95% or more.

[0072] The reference sequence is preferably, for example, the genomic sequence of the host cell or a nucleic acid sequence derived from a virus that has infected the host cell. The nucleic acid sequence derived from a virus may be the viral genome (DNA or RNA).

[0073] When a biological sample is obtained from a subject, the host cells of the membrane vesicles isolated from the biological sample may be unknown, whereas the host cells of the membrane vesicles isolated from the culture medium of the host cells are known.

[0074] When the host cell is unknown, mapping can be performed using the base sequence of the host cell genome and / or the viral genome sequence recorded in a public database, etc. as a reference sequence. This makes it possible to determine whether a cell having the mapped sequence region (as one form, the region with the highest degree of identity) in its genome is the host cell, or whether a cell infected with a virus having the mapped sequence region in its genome is the host cell. Examples of databases that can be used include NCBI RefSeq, NCBI GenBank, UCSC Genome Browser, GRCh37 reference primary assembly, and JRGv2.

[0075] On the other hand, if the host cell is known, the nucleotide sequence derived therefrom may be used as the reference sequence. In this case, the genomic sequence of the host cell and the sequence of the viral genome may be obtained from the databases mentioned above.

[0076] Alternatively, when a biological sample contains host cells, polynucleotides may be extracted from the residue (residue) after separating and obtaining membrane vesicles from the biological sample, and their base sequences may be analyzed to provide a reference sequence. In this case, the reference sequence may contain sequences other than the genomic sequence of the host cell. However, in this search method, the reference sequence only needs to contain the genomic sequence of the host cell and / or a sequence derived from the virus; the inclusion of other base sequences has little effect on subsequent steps. This is because, as in the case of using a public database as a reference sequence, the base sequence of the membrane vesicles is mapped to the genomic sequence of the host cell (or the genomic sequence of the virus that has infected the host cell) with a high degree of identity.

[0077] Mapping may also be performed in predetermined units, for example, in units of protein coding regions of the base sequences inside membrane vesicles. In this case, protein coding regions can be predicted and detected for the base sequences derived from membrane vesicles using, for example, software tools for genome annotation such as "prokka," and for each region, a homology search can be performed using the base sequence of the host cell genome as a reference, followed by mapping.

[0078] Next, in step S15, frequent regions whose detection frequency exceeds a standard are selected from the mapped regions. The method for selecting frequent regions is not particularly limited, but for example, when base sequences derived from membrane vesicles are individually obtained by single particle analysis, the number of membrane vesicles that have the sequence of each region in the genome sequence of each membrane vesicle and host cell for each mapped region can be added up to determine the detection frequency for each region. On the other hand, if the base sequence is obtained by shotgun metagenomics or the like without separating the membrane vesicles isolated from the biological sample individually, the base sequence can be determined by accumulating the number of sequence reads that contain the sequence of each mapped region.

[0079] In step S14, if mapping is performed based on protein coding regions, the above regions may be protein coding sequences, i.e., the detection frequency may be calculated for each protein coding sequence.

[0080] The criteria for detection frequency are not particularly limited as long as it is possible to identify frequently occurring regions in each region of the reference sequence to which base sequences derived from membrane vesicles are mapped. However, in one embodiment, it is preferable that the criteria be predetermined based on the significance level of a statistical test.

[0081] For example, when mapping is performed using protein coding regions as units, a statistical test is performed to determine whether the detection frequency of the target protein coding region is equivalent to the detection frequency of other protein coding regions. The significance probability (p-value) obtained by this statistical test is compared with a predetermined significance level. In this case, the significance level is preset to, for example, 0.01, and if the p-value is less than this value, it can be determined to be "more frequent than the standard."

[0082] Next, in step S16, a homology search is performed using the frequently occurring region as a query sequence, and frequently occurring regions whose similarity to detected homologous sequences derived from other than the host cell is lower than a standard are selected as biomarker candidate sequences. The frequent regions are regions that are detected at a specifically high frequency among the base sequences derived from membrane vesicles produced by host cells. Among these frequent regions, those specific to host cells are selected as candidate biomarker sequences.

[0083] Here, the base length of each sequence will be explained. The base sequence of a membrane vesicle is, in one form, several thousand to several tens of thousands of bp. The frequently occurring regions vary depending on the mapping method, etc., but are shorter than the base sequence of a membrane vesicle. The biomarker sequence is even shorter, in one form, being 50 to 250 bp. The base length of the candidate biomarker sequence is selected from a frequently occurring region, and therefore, in one embodiment, the base length may be the same as that of the frequently occurring region, and in one embodiment, the base length is 1000 bp or more.

[0084] The method of homology search is not particularly limited, but it is preferable to use the genome sequence of each cell in the public database already described as a reference. The software tool used for the homology search is not particularly limited, and any known software can be used, including, but not limited to, "diamond."

[0085] Preferably, candidate biomarker sequences have low similarity to homologous sequences, which increases the likelihood of finding a biomarker sequence. One example of a method for selecting such candidate biomarker sequences is to align each frequently occurring region with homologous sequences detected by homology search, and select the candidate biomarker sequence that has the smallest alignment score. Note that sequences derived from host cells are excluded from the homologous sequences.

[0086] In one embodiment, the detected homologous sequences (excluding those derived from the host cell) are arranged in descending order of alignment score, and the alignment scores of the highest-ranked homologous sequences are added up and used as an index. The index is compared for each frequently occurring region, and frequently occurring regions that meet predetermined criteria are selected as candidate biomarker sequences.

[0087] Examples of predetermined criteria include the ranking of scores. Specifically, when the alignment scores of each frequent region to the homologous sequence (or the total value of the alignment scores up to a certain rank in order from the highest alignment score) are sorted in ascending order, the frequent regions up to a certain rank are selected as candidate biomarker sequences.

[0088] The biomarker candidate sequences selected in this manner are regions (sequences) that are detected at significantly higher frequencies in base sequences derived from membrane vesicles and have lower similarity to the reference.

[0089] Next, in step S17, the biomarker candidate sequence is aligned with the homologous sequence to detect a region with sequence conservation lower than the standard, and this is designated as the biomarker sequence.

[0090] By aligning the candidate biomarker sequence with a homologous sequence (derived from a source other than the host cell), a region with particularly low sequence conservation (specifically, approximately 50 to 250 bp) is detected and used as the biomarker sequence.

[0091] The algorithm used for alignment is not particularly limited, and examples include "BLAST," "ClustalW," "Kalign," "MAFFT," "MUSCLE," and "T-Coffee," all of which are well known.

[0092] The criteria for sequence conservation are not particularly limited, but one example is a method in which, when homologous sequences are multiple aligned, entropy is calculated based on the frequency of occurrence of the four bases ATGC for each position in the sequence, and this is used as the degree of sequence conservation. This value is calculated for 50 to 250 consecutive bases, and if the average exceeds a predetermined range, the region is designated as a biomarker sequence.

[0093] The biomarker sequence determined in this manner is a region of the base sequence of membrane vesicles in a biological sample that is detected with high frequency and is specific to host cells, so by quantifying this biomarker sequence, the amount of membrane vesicles in a biological sample can be quantified.

[0094] Furthermore, when a biological sample is collected from a subject, by appropriately selecting the subject (e.g., a patient suffering from a particular disease), membrane vesicles that can be indicators of the subject's body and the progression of the disease can be detected using the above biomarker sequence, making it applicable to liquid biopsy methods.

[0095] Next, a method for detecting membrane vesicles in a biological sample using the above biomarker sequences will be described. Figure 2 is a flow chart showing the specific steps of the method for detecting membrane vesicles of the present invention. In FIG. 2, steps S11 to S17 are the same as the biomarker sequence search method already explained, and therefore explanations thereof will be omitted.

[0096] The method for detecting membrane vesicles includes, in step S18, detecting membrane vesicles by quantitative PCR using a primer pair designed to contain the biomarker sequence in a single amplification product, based on the information on the biomarker sequence selected in step S17. Note that step S18 may further include a reverse transcription step, and the quantitative PCR may be quantitative reverse transcription PCR.

[0097] Target sequences amplified by quantitative PCR can typically be detected using fluorescent dyes. There are at least two common methods for detecting them. One method uses fluorescent dyes that bind to double-stranded DNA (e.g., "SYBR Green I" (trade name), "TB Green" (trade name), etc.). As the PCR amplification product accumulates, more dye binds to the amplification product, increasing the fluorescent signal output in proportion to the amount of amplification product produced.

[0098] The second method uses fluorogenic oligonucleotide probes modified to contain a fluorescent reporter dye and a quencher dye at the 5' and 3' ends, respectively, that emit a fluorescent signal proportional to the amount of amplification product produced.

[0099] Examples of fluorescent dyes include "6-FAM," "HEX," "TET," "TAMRA," "JOE," "ROX," "Cyanine 3," "Cyanine 5," "Cyanine 5.5," "Cal Fluor Gold 540," "Cal Fluor Orange 560," "Cal Fluor Red 590," "Quasar 570," "Quasar 670," and "TxRd (Sulforhodamine 101-X)" (all trade names). Examples of quencher dyes include "TAMRA," "DABCYL dT," "BHQ-1," "BHQ-2," "BHQ-3," "OQ," "Iowa Black FQ," and "Iowa Black RQ" (all trade names).

[0100] The difference between these two methods is that the former detects all double-stranded DNA regardless of sequence (e.g., non-specific reaction products), whereas the latter fluorogenic probe approach is specific to the target sequence. Either method can be applied to the membrane vesicle detection method of the present invention.

[0101] There are no particular limitations on the method for designing primers (pairs) used in quantitative PCR, and known methods can be used. In particular, primers can be easily designed using computer-generated algorithms. Examples of such algorithms include "Primer-BLAST" and "DINAMelt." Furthermore, primers can also be selected according to guidelines other than those described above that are known to those skilled in the art.

[0102] In one embodiment, the forward primer and reverse primer for quantitative PCR may each be 18 to 25 bases in length (preferably 20 to 24 bases), have a Tm of 57 to 61°C, and have a GC content of 40 to 60%.

[0103] The base length of the amplification product (amplification product) is not particularly limited, but an example is 50 to 250 bp.

[0104] When a fluorescent probe is used in quantitative PCR, the annealing region of the fluorescent probe preferably does not overlap with the annealing region of the primer, and is preferably designed to anneal to a sequence between the annealing regions of the primer pair. In one embodiment, the nucleotide distance between the end of the forward primer and the start of the probe is preferably 0 to 60 base pairs (bp).

[0105] The probe for quantitative PCR may have a melting temperature (Tm) that is approximately 5 to 10°C higher than the melting temperature of the primers, and in this case, when the thermal cycle is reduced from the denaturation temperature to the annealing and extension temperature, the fluorescent probe anneals to the target sequence before any of the primers anneal. The Tm of the fluorescent probe is not particularly limited, but in one embodiment, it may be 55 to 80°C.

[0106] The base length of the probe is not particularly limited, but may be, for example, 8 to 45 bp. More specifically, if it is a TaqMan probe, it may be 20 to 30 bases long.

[0107] As PCR progresses, the fluorescent signal output increases proportionally with the amount of amplified product. Measurement of the amount of membrane vesicles in a biological sample can be performed by first performing parallel PCR using a biological sample containing a high amount of membrane vesicles and serial dilutions of a standard sample containing known concentrations of the biomarker sequence. Serially diluted samples generate a standard curve relating copy number of the biomarker sequence to the cycle number at which fluorescence above background is first detected. By determining the cycle data for a biological sample and comparing it to a standard curve, the copy number of the biomarker sequence (or its complementary sequence) in the biological sample can be calculated.

[0108] By quantifying the concentration (copy number) of the biomarker sequence in the standard material, for example, by absorbance, absolute quantification of the concentration of the biomarker sequence in the biological sample can be performed from the standard curve.

[0109] In another embodiment (variant) of the membrane vesicle detection method of the present invention, step S19 may be further included between steps S17 and S18, in which primers are designed and synthesized so that the biomarker sequence is contained in one PCR amplification product. The method for designing and synthesizing the primers is not particularly limited, and in addition to those already described, known methods can be used, and an automated nucleic acid synthesizer, as described below, may also be used. Figure 3 is a flowchart of the membrane vesicle detection method in the variant.

[0110] Thus, this method allows for the easy quantification of membrane vesicles in biological samples.

[0111] [Membrane vesicle detection device] Next, the membrane vesicle detection device of the present invention will be described. A first embodiment of the apparatus of the present invention is a membrane vesicle detection apparatus comprising: a base sequence analysis apparatus for obtaining the base sequence of membrane vesicles produced by host cells; a quantitative PCR apparatus for detecting the membrane vesicles in a biological sample by quantitative polymerase chain reaction using primers designed to contain a biomarker sequence, which is a part of the base sequence, in one amplification product; and a control device. The control device comprises: a map unit that maps the base sequence obtained by the base sequence analysis apparatus to a reference sequence derived from the host cell; a frequent region selection unit that selects, from the mapped regions, frequent regions whose detection frequency exceeds a standard; a biomarker candidate sequence selection unit that performs a homology search using the frequent region as a query sequence and selects, as biomarker candidate sequences, the frequent regions that show lower similarity to detected homologous sequences derived from sequences other than the host cell than a standard; and a biomarker sequencing unit that aligns the biomarker candidate sequences to the homologous sequences to detect regions with lower sequence conservation than a standard and selects these regions as the biomarker sequences.

[0112] The above device will be described in detail with reference to the drawings. Figure 4 shows the hardware configuration of this device. The membrane vesicle detection device 20 has a base sequence analysis device 21, a quantitative PCR device 22, and a control device 23, and each device exchanges data with each other via a bus so that the control device 23 can control the base sequence analysis device 21 and the quantitative PCR device 22.

[0113] The base sequence analyzer 21 is equipped with a pre-processing device 28 that can extract nucleic acids from the separated and obtained membrane vesicles, fragment DNA, select the chain length of the fragments, and amplify them, and a sequencer 29 that performs sequence analysis. The control device 23 is a computer that includes a processor 24 , a storage device 25 , a display device 26 , and an input device 27 .

[0114] The processor 24 is, for example, a microprocessor, a processor core, a multiprocessor, an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), and a general-purpose computing on graphics processing unit (GPGPU).

[0115] The storage device 25 has the function of temporarily and / or non-temporarily storing various programs and data, and provides a working area for the processor 24 . The storage device 25 is, for example, a read only memory (ROM), a random access memory (RAM), a hard disk drive (HDD), a flash memory, or a solid state drive (SSD).

[0116] The display device 26 can display the analysis results, the sample name, the operating procedure, etc. The display device 26 may be a liquid crystal display, an organic EL (Electro Luminescence) display, or the like. Furthermore, the display device 26 may be configured integrally with the input device 27. In this case, the display device 26 may be a touch panel display that provides a GUI (Graphical User Interface).

[0117] The input device 27 can accept input of measurement conditions, sample names, etc., and can also accept input of instructions to the membrane vesicle detection device 20. The input device 27 may be a keyboard, a mouse, a scanner, a touch panel, or the like.

[0118] The base sequence analyzer 21 is an apparatus for obtaining base sequences derived from the separated and obtained membrane vesicles. The base sequence analyzer 21 includes a pre-processing device 28 that performs pre-processing of the separated and obtained membrane vesicles, and a sequencer 29 that performs sequence analysis.

[0119] The pretreatment device 28 is configured to automatically perform nucleic acid extraction from membrane vesicles, reverse transcription as needed, fragmentation of purified DNA (or cDNA), selection of DNA chain length to be applied to the sequencer 29, and amplification of DNA as needed; a commercially available device can also be incorporated into this device. Examples of such pretreatment devices include the "Biomek" pretreatment system manufactured by Beckman Coulter and the "Bravo" NGS automated system manufactured by Agilent. Note that the base sequence analyzer 21 does not necessarily have to include the pretreatment device 28. The sequencer 29 is not particularly limited, and any known next-generation sequencer can be used without any particular limitation.

[0120] A quantitative PCR device is a device for detecting membrane vesicles in a biological sample by quantitative polymerase chain reaction using a primer pair designed to contain a membrane vesicle biomarker sequence in a single amplification product, and any known real-time PCR device can be used without any particular restrictions. Examples of such devices include the "AriaMx" manufactured by Agilent and the "qTOWER" manufactured by Analytik Jena. 3", "QIAquant" manufactured by Qiagen, and "QuantStudio" manufactured by Thermo Fisher.

[0121] 5 is a functional block diagram of the present apparatus. The membrane vesicle detection apparatus 20 includes a base sequence analyzer 21, a quantitative PCR apparatus 22, and a control device 23. The control device 23 includes a mapper 31, a frequent region selector 32, a biomarker candidate sequence selector 33, and a biomarker sequence determiner 34.

[0122] The mapping unit 31 is a function realized by the processor 24 of the control device 23 executing a program stored in the storage device 25. The mapping unit 31 maps the base sequence of the membrane vesicle obtained by the base sequence analysis device 21 to a reference sequence (or a set of sequences including the reference sequence) derived from a host cell obtained (35) from a public database or the like via a network using the input device 27 or a communication function (not shown) of the control device 23. The mapping method has already been described in the description of the membrane vesicle detection method of the present invention.

[0123] The frequent region selection unit 32 is a function realized by the processor 24 of the control device 23 executing a program stored in the storage device 25. Of the base sequences derived from membrane vesicles, regions mapped to the genome sequence of the host cell by the mapping unit 31 are selected by the frequent region selection unit 32 as frequent regions whose detection frequency exceeds a criterion. The detection frequency criterion may be set in advance and stored in the storage device 25. The method for selecting frequent regions is as already described in the explanation of the first embodiment of the membrane vesicle detection method of the present invention.

[0124] The biomarker candidate sequence selection unit 33 is a function realized by the processor 24 of the control device 23 executing a program stored in the storage device 25. The biomarker candidate sequence selection unit 33 performs a homology search using a frequently occurring region as a query sequence and a genomic sequence of a cell other than the host cell obtained (36) from a public database or the like via a network using the input device 27 or the communication function of the control device 23 as a reference, and selects, as a biomarker candidate sequence, the frequently occurring region whose similarity to the detected homologous sequence derived from a non-host cell is lower than a standard. The method for selecting biomarker candidate sequences is as already described in the description of the first embodiment of the membrane vesicle detection method of the present invention.

[0125] The biomarker sequence determination unit 34 is a function realized by the processor 24 of the control device 23 executing a program stored in the storage device 25. The biomarker sequence determination unit 34 aligns the biomarker candidate sequences selected by the biomarker candidate sequence selection unit 33 with the homologous sequences, detects regions where sequence conservation is lower than a standard, and designates these as biomarker sequences. The method for determining the biomarker sequence is as already described in the description of the first embodiment of the membrane vesicle detection method of the present invention.

[0126] The quantitative PCR device 22 detects membrane vesicles in a biological sample using primers designed to contain in one amplification product the biomarker sequence (or its complementary sequence) determined by the biomarker sequence determination unit 34. That is, by detecting the amplification of the biomarker sequence determined by the biomarker sequence determination unit 34, the presence / absence and / or amount of the biomarker in the original biological sample is measured.

[0127] Next, the operation of the membrane vesicle detection device will be described. Fig. 6 shows the operation flow of the control device 23 of the membrane vesicle detection device 20. First, the control device 23 controls the base sequence analyzer 21 to obtain the base sequence of the membrane vesicle produced by the host cell (step S41).

[0128] Specifically, the control device 23 first controls the pretreatment device 28 to separate and obtain membrane vesicles from the biological sample, and then performs pretreatment such as extraction of nucleic acid components, etc. Next, the control device 23 controls the sequencer 29 to obtain base sequences derived from the membrane vesicles. The membrane vesicle detection device does not need to have the pretreatment device 28, and in this case, the device may be configured to control the sequencer 29 to obtain base sequences derived from membrane vesicles from pretreated samples.

[0129] Next, the control device 23 controls the map unit 31 to map the base sequence obtained by the base sequence analyzer 21 to a reference sequence derived from the host cell (step S42). In one embodiment, the control device 23 maps the base sequence derived from the membrane vesicle to a reference sequence derived from the host cell obtained (35) from a public database or the like.

[0130] Next, the control device 23 controls the frequent region selection unit 32 to select a frequent region whose detection frequency exceeds a standard from among the regions mapped to the reference sequence (step S43).

[0131] Next, the control device 23 controls the biomarker candidate sequence selection unit 33 to perform a homology search using the frequent region selected by the frequent region selection unit 32 as a query sequence against sequences derived from other than host cells that have been acquired (36) from a public database or the like via a network using the input device 27 or the communication function of the control device 23, and selects, as biomarker candidate sequences, frequent regions whose similarity to the detected homologous sequences derived from other than host cells is lower than a criterion (step S44).

[0132] Next, the control device 23 controls the biomarker sequence determination unit 34 to align the biomarker candidate sequences selected by the biomarker candidate sequence selection unit 33 with homologous sequences derived from cells other than the host cell detected by the homology search, and detects regions with sequence conservation lower than the standard, which are then designated as biomarker sequences (step S45).

[0133] Next, the control device 23 controls the quantitative PCR device 22 to detect membrane vesicles in the biological sample by quantitative polymerase chain reaction using primers designed based on the biomarker sequence and designed to include the biomarker sequence (or its complementary sequence) in one amplification product (step S46).

[0134] This device allows for easy quantification of membrane vesicles in biological samples.

[0135] [Membrane vesicle detection device (second embodiment)] A second embodiment of the membrane vesicle detection device of the present invention is a membrane vesicle detection device comprising: a base sequence analysis device for obtaining the base sequence of membrane vesicles produced by host cells; a nucleic acid synthesis device for designing and synthesizing a primer pair that contains a biomarker sequence, which is a part of the base sequence, in one amplification product; a quantitative PCR device for detecting the membrane vesicles in a biological sample by quantitative polymerase chain reaction using the primers; and a control device. The control device comprises: a map unit that maps the base sequence obtained by the base sequence analysis device to a reference sequence derived from the host cell; a frequent region selection unit that selects, from the mapped regions, frequent regions whose detection frequency exceeds a standard; a biomarker candidate sequence selection unit that performs a homology search using the frequent region as a query sequence and selects, as biomarker candidate sequences, the frequent regions that show lower similarity to detected homologous sequences derived from sequences other than the host cell than a standard; and a biomarker sequencing unit that aligns the biomarker candidate sequences to the homologous sequences to detect regions with lower sequence conservation than a standard and selects these regions as the biomarker sequences.

[0136] Since this device has the same configuration as the first embodiment except for the inclusion of a nucleic acid synthesizer, the following description will focus on the above differences.

[0137] Fig. 7 is a hardware configuration diagram of the device, Fig. 8 is a functional block diagram of the device, and Fig. 9 is an operation flow diagram of the control device of the device.

[0138] The membrane vesicle detection device 50 includes a nucleic acid synthesizer 51. The nucleic acid synthesizer 51 designs and synthesizes a primer pair that contains the biomarker sequence (or its complementary sequence) in one amplification product, based on the biomarker sequence determined by the biomarker sequence determination unit 34. Furthermore, a fluorescent probe may also be synthesized, if necessary.

[0139] Such nucleic acid synthesizers can be well known, such as the "NTS M" series from Nippon Techno Service, the "Dr. Oligo" series from Biolytic Lab Performance, and DNA / RNA synthesizers from Oligomaker.

[0140] The operation flow of the control device 23 differs from the operation flow of the control device in the first embodiment in that it has the following two steps instead of step S46.

[0141] First, the method includes a step in which the control device 23 controls the nucleic acid synthesizer 51 to design and synthesize a primer (and a probe, if necessary) that contains the biomarker sequence (or its complementary sequence) in one amplification product based on the biomarker sequence determined by the biomarker sequence determination unit 34 (step S47).

[0142] The second feature is that the control device 23 controls the quantitative PCR device 22 to detect membrane vesicles in a biological sample using the primers synthesized by the nucleic acid synthesizer (and, if necessary, the probes synthesized) (step S48).

[0143] Since this device has a nucleic acid synthesizer 51, it can automatically determine the biomarker sequence, synthesize primers, and detect membrane vesicles by quantitative PCR, making it easier to quantify membrane vesicles in biological samples.

[0144] [Biomarker sequence search device] An embodiment of the biomarker sequence searching device of the present invention is a biomarker sequence searching device comprising: a base sequence analyzing device for obtaining the base sequence of a membrane vesicle produced by a host cell; and a control device, wherein the control device comprises: a map unit that maps the base sequence to a reference sequence derived from the host cell; a frequent region selection unit that selects, from the mapped regions, a frequent region whose detection frequency exceeds a standard; a biomarker candidate sequence selection unit that performs a homology search using the frequent region as a query sequence and selects, as a biomarker candidate sequence, the frequent region whose similarity to a detected homologous sequence derived from a sequence other than the host cell is lower than a standard; and a biomarker sequencing unit that aligns the biomarker candidate sequence to the homologous sequence, detects a region whose sequence conservation is lower than a standard, and selects this as a biomarker sequence.

[0145] FIG. 10 is a hardware configuration diagram of the present device. FIG. 11 is a functional block diagram of the present device. FIG. 12 is an operational flow diagram of the control device of the present device. As shown in FIGS. 10 and 11, the biomarker sequence searching device 30 has a hardware configuration similar to that of the first embodiment of the membrane vesicle detection device already described, except that the quantitative PCR device has been removed. The configuration of the device other than the above and the functions of each part are the same as those of the first embodiment of the membrane vesicle detection device, and therefore will not be described again.

[0146] Furthermore, as shown in the operational flow diagram of the control device in Figure 12, the operational flow of the control device of this device differs from the operational flow of the control device of the first embodiment of the membrane vesicle detection device only in that it does not include step S46 (a step in which the control device 23 controls the quantitative PCR device 22 to detect membrane vesicles in a biological sample by quantitative polymerase chain reaction using primers designed to include a biomarker sequence (or its complementary sequence) in one amplification product); all other steps are the same.

[0147] This device makes it possible to search for biomarker sequences that can be used to easily quantify membrane vesicles in biological samples.

[0148] [computer] The computer of the present invention is a computer that searches for biomarker sequences, and has: a mapping unit that maps the base sequence of membrane vesicles produced by a host cell to a reference sequence derived from the host cell; a frequent region selection unit that selects, from the mapped regions, a frequent region whose detection frequency exceeds a standard; a biomarker candidate sequence selection unit that performs a homology search using the frequent region as a query sequence and selects, as biomarker candidate sequences, the frequent region whose similarity to detected homologous sequences derived from sequences other than the host cell is lower than a standard; and a biomarker sequencing unit that aligns the biomarker candidate sequences to the homologous sequences, detects regions whose sequence conservation is lower than a standard, and selects these regions as the biomarker sequences.

[0149] FIG. 13 is a diagram showing the hardware configuration of this computer. The computer 60 includes at least a processor 24, a storage device 25, a display device 26, and an input device 27, and may further include a communication device (not shown).

[0150] The processor 24 is, for example, a microprocessor, a processor core, a multiprocessor, an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), and a general-purpose computing on graphics processing unit (GPGPU).

[0151] The storage device 25 has the function of temporarily and / or non-temporarily storing various programs and data, and provides a working area for the processor 24 . The storage device 25 is, for example, a read only memory (ROM), a random access memory (RAM), a hard disk drive (HDD), a flash memory, or a solid state drive (SSD).

[0152] The display device 26 can display the analysis results, the sample name, the operating procedure, etc. The display device 26 may be a liquid crystal display, an organic EL (Electro Luminescence) display, or the like. Furthermore, the display device 26 may be configured integrally with the input device 27. In this case, the display device 26 may be a touch panel display that provides a GUI (Graphical User Interface).

[0153] The input device 27 can accept input of measurement conditions, sample names, etc., and can also accept input of instructions to the membrane vesicle detection device 20. The input device 27 may be a keyboard, a mouse, a scanner, a touch panel, or the like.

[0154] 14 is a functional block diagram of the computer 60. The computer 60 has a map section 31, a frequent region selection section 32, a biomarker candidate sequence selection section 33, and a biomarker sequence determination section .

[0155] The mapping unit 31 is a function realized by the processor 24 executing a program stored in the storage device 25. The mapping unit 31 maps the base sequence of a membrane vesicle, which is input (61) from the input device 27 or externally via a network using a communication function (not shown), to the genome sequence of a host cell (or a set of sequences including the same), which is also obtained (35) from a public database or the like. The mapping method has already been described in the description of the membrane vesicle detection method of the present invention.

[0156] The frequent region selection unit 32 is a function realized by the processor 24 executing a program stored in the storage device 25. Of the base sequences derived from membrane vesicles, regions mapped by the mapping unit 31 to a reference sequence derived from a host cell are selected by the frequent region selection unit 32 as frequent regions whose detection frequency exceeds a criterion. The detection frequency criterion may be set in advance and stored in the storage device 25. The method for selecting frequent regions is as already described in the description of the first embodiment of the membrane vesicle detection method of the present invention.

[0157] The biomarker candidate sequence selection unit 33 is a function realized by the processor 24 executing a program stored in the storage device 25. The biomarker candidate sequence selection unit 33 performs a homology search using a frequently occurring region as a query sequence against sequences derived from sources other than host cells that have been obtained (36) from the input device 27 or a public database or the like via a network using the communication function, and selects, as biomarker candidate sequences, the frequently occurring region whose similarity to the detected homologous sequence derived from sources other than host cells is lower than a criterion. The method for selecting biomarker candidate sequences is as already described in the description of the first embodiment of the membrane vesicle detection method of the present invention.

[0158] The biomarker sequence determination unit 34 is a function realized by the processor 24 executing a program stored in the storage device 25. The biomarker sequence determination unit 34 aligns the biomarker candidate sequences selected by the biomarker sequence candidate selection unit 33 with the homologous sequences, detects regions where sequence conservation is lower than a standard, and designates these as biomarker sequences. The method for determining the biomarker sequence is as already described in the description of the first embodiment of the membrane vesicle detection method of the present invention.

[0159] FIG. 15 is a flow diagram of the process of searching for biomarker sequences by the processor of the computer.

[0160] First, the processor 24 performs a process of mapping the base sequence of the membrane vesicle obtained from outside to a reference sequence derived from the host cell by the mapping unit 31 (step S71). In one embodiment, the processor 24 maps the base sequence derived from the membrane vesicle to a reference sequence, which is a genomic sequence of the host cell obtained (35) from a public database or the like.

[0161] Next, the processor 24 causes the frequent region selection unit 32 to perform a process of selecting frequent regions whose detection frequency exceeds a standard from among the regions mapped to the genome sequence of the host cell (step S72).

[0162] Next, the processor 24 performs a homology search using the frequent region selected by the frequent region selection unit 32 as a query sequence against (using this as a reference) a sequence derived from a cell other than the host cell (for example, a genomic sequence of a cell other than the host cell) obtained from the input device 27 or a public database or the like via a network using the communication function (36), and selects, as a biomarker candidate sequence, a frequent region whose similarity to the detected homologous sequence derived from a cell other than the host cell is lower than a standard (step S73).

[0163] Next, the processor 24 causes the biomarker sequence determination unit 34 to align the biomarker candidate sequences selected by the biomarker candidate sequence selection unit 33 with homologous sequences derived from cells other than the host cell detected by the homology search, detect regions with sequence conservation lower than the standard, and designate these as biomarker sequences (step S74).

[0164] This computer makes it possible to search for biomarker sequences that can be used to easily quantify membrane vesicles in biological samples from the base sequences of membrane vesicles input from an external source.

[0165] [Discovery methods for disease biomarkers] The disease biomarker discovery method of the present invention is one example of application in which the already-described biomarker sequence discovery method is used to discover biomarkers for detecting specific diseases. Figure 19 is a flow chart of the disease biomarker discovery method of the present invention. The disease biomarker discovery method will be described in detail below using the flow chart of Figure 19.

[0166] First, in step S81, the base sequence of a first membrane vesicle obtained from a first biological sample collected from a diseased patient is compared with the base sequence of a known cell or the genome sequence of a known virus to determine the "host cell A" that is the host cell of the first membrane vesicle.

[0167] The type of "disease" is not particularly limited, but is preferably the above-mentioned "specific disease," and typically, diseases caused by bacteria and / or viruses are more preferred, with bacterial and / or viral infections being even more preferred. One form is preferably periodontal disease.

[0168] The first biological sample collected from a diseased patient is not particularly limited, but is preferably at least one selected from the group consisting of saliva, blood, serum, plasma, buffy coat, lymph, interstitial fluid, body cavity fluid, digestive fluid, sweat, urine, nasal discharge, tears, semen, vaginal fluid, amniotic fluid, milk, sputum, surgical irrigation fluid, feces, and swabs of skin or body mucosa, and more preferably saliva. Furthermore, the collection method for these is not particularly limited, and they may be collected by known methods.

[0169] The first biological sample may be a single sample collected from a single patient, multiple samples (multiple types of samples, samples collected at multiple times) collected from a single patient, or single / multiple samples collected from multiple patients (patient groups). In particular, samples collected from multiple patients are preferred from the viewpoint of more accurately determining the "host cell C" described below, i.e., as being specific to a patient group. In this case, step S81 may be performed multiple times using the samples collected from the multiple patients as they are, or the samples collected from the multiple patients may be mixed and used as a new first biological sample to perform step S81.

[0170] This step may further include the step of isolating membrane vesicles contained in the first biological sample collected from the patient(s) and obtaining the base sequence of the membrane vesicles. The method of isolating membrane vesicles from the biological sample and obtaining the base sequence can be the same as steps S11 to S13 in the flow of the biomarker sequence search method already described (FIG. 1).

[0171] Generally, the host cell of the first membrane vesicle contained in the first biological sample collected from a diseased patient is unknown. In this step, the base sequence of the first membrane vesicle is compared with the base sequence of a known cell or the genome sequence of a known virus to identify the host cell of the first membrane vesicle, "host cell A." As already explained, "host cell A" may be a cell itself or a cell infected with a virus.

[0172] The first membrane vesicle may be one type or two or more types. When there are two or more types, the first membrane vesicle may be produced by one type of host cell or two or more types of host cells. Therefore, the host cell A may be one type or two or more types.

[0173] The method for comparing the base sequence of the first membrane vesicle with the base sequence of a known cell or the genome sequence of a known virus is not particularly limited, and examples include a method of performing a homology search using BLAST or the like, using the base sequence of a genome of a known cell or a genome sequence of a known virus stored in a public database, etc., as a reference. The comparison in this step can be performed using a method similar to step 14 in the flow of the biomarker sequence search method already described (Figure 1).

[0174] The method for organizing the data obtained during the implementation of this step is not particularly limited, but examples include a table that compiles multiple records in which sequences obtained from individual membrane vesicles obtained from the first biological sample are attributed to one or more host cells A (the cells themselves or viruses). Furthermore, the table may further include base length data for each host cell, and such data can be visualized using a heat map or the like (e.g., Figure 21).

[0175] Next, in step S82, the base sequence of the second membrane vesicle obtained from the second biological sample collected from the healthy individual is compared with the base sequence of a known cell or the genome sequence of a known virus to determine the host cell of the second membrane vesicle, "host cell B." As already explained, "host cell B" may be a cell itself or a cell infected with a virus. The procedures in this step are the same as those in step S81 except that a second biological sample is used instead of the first biological sample, and the preferred form is also the same.

[0176] Next, in step S83, host cell A is compared with host cell B to identify "host cell C" as the host cell with a high abundance ratio in the first biological sample. First, the "abundance ratio" is defined as the ratio of the number of membrane vesicles contained in a biological sample when the membrane vesicles are attributed to a host cell or a virus that has infected the host cell. In one embodiment, it is preferably defined as the number of membrane vesicles attributed to a certain host cell (virus) relative to the total number of membrane vesicles detected from a certain biological sample. Note that the sequence of an individual membrane vesicle does not necessarily belong to one type of host cell (virus), and may belong to two or more types of host cells. In such cases, the abundance ratio may be calculated as described above.

[0177] In this step, the above comparison excludes cases where the abundance ratio is simply high (membrane vesicles detected at high frequency) but is non-specific, i.e., cases where the vesicles originate from a common host cell. The term "high abundance ratio" may refer to the highest abundance ratio or the abundance ratio equal to or greater than a predetermined reference value. The above comparison identifies "host cell C," which is the host cell of membrane vesicles detected more frequently in the first biological sample. Host cell C is included in host cell A. It is preferable that host cell C is not included in host cell B. When host cell C is included in host cell B, it is preferable that the abundance ratio of host cell C in the second biological sample is equal to or less than a predetermined reference value.

[0178] Specific comparison methods for selecting host cell C include, for example, the following method. First, the host cells (or viruses) of the membrane vesicles contained in each sample are grouped by closely related species, and a group of host cells (or viruses) specific to the first membrane vesicle is identified. Next, the abundance ratio (attributed frequency of membrane vesicle units) within the identified group is compared, and host cells whose abundance ratio is equal to or greater than a predetermined standard are designated as "host cell C." Examples of grouping methods include, for example, grouping by "phylum" in the taxonomic hierarchy if the host cells are bacteria, and grouping by "family" if the host cells are infected by viruses. The present disease biomarker discovery method includes this step, making it possible to discover markers with higher specificity.

[0179] Next, in step S84, the base sequence of the membrane vesicles produced by host cell C is mapped to a reference sequence derived from host cell C (typically, the genome sequence of host cell C or the genome sequence of a virus that has infected host cell C).Further, in step S85, from among the mapped regions, regions whose detection frequency exceeds a standard are selected. The "membrane vesicles produced by host cell C" used in the above step may typically be membrane vesicles obtained from the first biological sample that are determined to have been produced by host cell C.

[0180] For example, the present inventors have found that even if there is one type of host cell C, the membrane vesicles produced from that host cell C may not be of one type, but may be of two or more types. In the above steps S84 and S85, regions that are detected with high frequency are identified in membrane vesicles, which may be of multiple types, produced from the host cell C.

[0181] In the detailed procedure described above, step S84 is similar to step S14 in the flow of the biomarker sequence search method (FIG. 1) already described, and the preferred embodiment is also similar. Also, step S85 is similar to step S15 in the flow of the biomarker sequence search method (FIG. 1) already described, and the preferred embodiment is also similar.

[0182] Next, in step S86, a homology search is performed using the frequently occurring region as a query sequence, and the detected "frequent region" whose similarity to homologous sequences derived from "other than host cell C" is lower than the standard is selected as a biomarker candidate sequence. This step is similar to step S16 in the flow of the biomarker sequence search method already described (FIG. 1), and the preferred embodiment is also similar. In this step, sequences that are particularly low in similarity to references and the like are selected from the "frequent regions" identified in step S15.

[0183] Next, in step S87, host cell D, which is a closely related species of host cell C, is selected from host cell B, and then, in step S88, the base sequence or a fragment thereof of membrane vesicles derived from host cell D is mapped to a reference sequence (reference sequence derived from host cell C), and then, in step S89, the mapped region is excluded from the biomarker candidate sequence.

[0184] Host cell D is a close relative of host cell C (a specific host cell among host cell A of a first membrane vesicle contained in a first biological sample collected from a patient) and is a host cell contained in host cell B (a host cell of a second membrane vesicle contained in a second biological sample collected from a healthy individual). The membrane vesicles of host cell C and its closely related host cell D may share common regions in their nucleotide sequences. In particular, the existence of common regions is inferred at the fragment (e.g., sequence read) level.

[0185] In this step, the base sequence of the membrane vesicles of host cell D is first mapped to a reference sequence derived from host cell C, i.e., the genome sequence of host cell C or the genome sequence of a virus that has infected host cell C. If any of the mapped regions overlaps with the biomarker candidate sequence, this region is excluded from the biomarker candidate sequence. This further enhances the specificity of the biomarker candidate sequence. Note that the above steps S87 to S89 do not necessarily have to be included.

[0186] The range of closely related species may be selected as appropriate. For example, one embodiment may involve a method in which the method is performed for each "phylum" of the taxonomic rank if the host cells are bacteria, or a method in which the method is performed for each "family" of the virus if the host cells are virus-infected host cells.

[0187] Next, in step S90, the biomarker candidate sequences are aligned with homologous sequences to detect regions with sequence conservation lower than a standard, which are designated as disease biomarkers. Step S90 is similar to step S17 in the flow of the biomarker sequence search method already described (FIG. 1), and the preferred embodiment is also similar.

[0188] The disease biomarkers determined in this manner are regions that are detected with high frequency in the base sequences of membrane vesicles in biological samples collected from diseased patients and are specific to host cell C. Therefore, by quantifying the disease biomarkers in membrane vesicles, information such as the progression of the disease can be obtained. [Example]

[0189] The present invention will be described below with reference to examples, but the present invention is not limited to these examples.

[0190] [Example 1] Streptococcus mutans NG8 was used as the bacterial strain and cultured in DSMZ medium or Gifu medium at 37°C. To create anaerobic conditions, a N2 / CO2 (80:20 v / v) mixed gas was first introduced into the vial containing the bacterial strain and medium for 20 minutes. The culture was then continued until the optical density (OD600) of the culture reached 1.0. After cultivation, the bacterial culture was centrifuged at 7800 rpm at 4°C for 10 minutes. The supernatant containing membrane vesicles was passed through a 0.22 μm filter to remove cell debris, and membrane vesicles were purified from the filtered supernatant.

[0191] The filtered supernatant was ultracentrifuged at 126,000 × g for 2 hours at 4°C. After ultracentrifugation, the supernatant was removed, and the pelleted membrane vesicles were resuspended in phosphate-buffered saline (PBS) and stored at 4°C. Further purification of membrane vesicles was performed using an iodoxanol (OptiPrep) density gradient. The 60% iodoxanol stock solution was diluted with PBS to 35%, 30%, 25%, 20%, 15%, and 10%. The stored suspension containing membrane vesicles was ultracentrifuged at 140,000 × g for 2 hours at 4°C. The resulting pellet was resuspended in 35% iodoxanol solution. The six prepared iodoxanol solutions were loaded into ultracentrifuge tubes in descending order of density. The 35% suspension containing membrane vesicles was placed at the bottom, and the 10% solution was placed at the top. The tubes were ultracentrifuged in a swing-out rotor at 140,000 × g for 16 hours at 4°C. After ultracentrifugation, 1 mL fractions were collected from each layer and stored in separate Eppendorf tubes. The size and number of particles contained in each fraction were measured using a nanoparticle tracking device, and the fraction containing the most particles with diameters of approximately 100 to 200 nm was determined to be the "fraction concentrated in membrane vesicles" and used for further analysis.

[0192] The DNA fragments derived from the membrane vesicles were analyzed using a next-generation sequencer to obtain short-read sequences. Among these short-read sequences (approximately 150 bp), sequences with low sequence quality were removed using fastp (Chen et al., fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics 34, 17, 2018).

[0193] These short read sequences were then assembled using SPAdes (Bankevic et al., SPAdes: A New Genome Assembly Algorithm and Its Applications to Single-Cell Sequencing. J. Comput. Biol., 19, 5, 2012) to create long base sequences called contigs, which are several kbp to several tens of kbp in length.

[0194] Protein-coding regions (CDS) within the base sequences inside each membrane vesicle were predicted and detected using Prokka (Seemann, Prokka: rapid prokaryotic genome annotation. Bioinformatics, 30, 14, 2014). For each CDS, a homology search was performed using the reference genome of a cultured bacterial strain registered with NCBI (Streptococcus mutans strain NG8 NZ_CP013237.1) to identify regions mapped to the genome.

[0195] For homology searches, we used diamond (Buchfunk et al., Sensitive protein alignments at tree-of-life scale using DIAMOND, Nat. Methods., 18, 2021). Regions detected in a statistically significantly higher number of membrane vesicles than other CDS regions were defined as frequent regions (hyper-geometric distribution test: p-value < 0.01).

[0196] Figure 16 shows the results of mapping the base sequence within a membrane vesicle particle to a reference sequence. 96 membrane vesicle particles were analyzed, and the regions in each particle where the base sequence was mapped to the reference sequence are shown in black. Figure 17 shows the results of changing the vertical axis in Figure 16 from the number of membrane vesicle particles to the "detection frequency," i.e., the total number of membrane vesicle particles that contained the sequence of that region. The above results indicate that there are regions that many membrane vesicle particles share in common.

[0197] A homology search (using diamond) was performed on the screened frequently occurring regions using protein sequences registered in the Swiss-Prot database as reference sequences, and the functional information (Gene Ontology) tagged to each protein sequence information in the matched database was extracted using an analysis algorithm to detect gene functions that were detected with statistically significant frequency (Hyper-geometric distribution test: p-value < 0.05).

[0198] Figure 18 lists the gene functions. "SM bacteria" refers to the results for Streptococcus mutans. "PAO" and "PG bacteria" refer to the functions of the frequently occurring regions (protein-coding regions) in Pseudomonas aeruginosa PAO1 and Porphyromonas gingivalis W83 (PG), respectively, which were selected using the same procedure as for SM bacteria.

[0199] Next, a homology search was performed using the frequently occurring regions as query sequences and the protein sequences (nr) of organisms across all domains registered in NCBI as reference sequences (diamond). Among the protein-coding sequences in the database that matched, the sum of the alignment scores (Max Score) of the top 20 protein-coding sequences, excluding sequences derived from the same strain, was calculated, and only the sequences with the smallest value were selected as biomarker candidate sequences. The base sequences of these biomarker candidate sequences and their homologous sequences (the top 20 protein-coding sequences mentioned above) were aligned using MAFFT. If a region of low sequence conservation over approximately 100 bases was detected as a result of the alignment, that region was designated as the biomarker sequence.

[0200] The 18-25 base pairs at either end of the least conserved biomarker sequence were used as primers. The Tm values ​​of the primer sequences were calculated using a TmCalculator, and the sequence length was adjusted to be suitable for detection by qPCR.

[0201] By using this primer pair in combination with, for example, Takara's "TB Green Fast qPCR Mix" and Promega's "GoTaq Probe qPCR Master Mix," it is possible to quantify membrane vesicles produced by the host bacterium Streptococcus mutans in biological samples.

[0202] [Example 2] A search for periodontal disease biomarkers was carried out using the following method.

[0203] (Collection of biological samples) First, saliva samples were obtained from three healthy volunteers and six periodontal disease patients. The periodontal disease patients were classified as stage III, grade C (according to the American Academy of Periodontology and the European Federation of Periodontology). These patients had not received antibiotics for three months, were non-smokers, and had no diabetes or other systemic diseases. Saliva samples were collected using a Saliva Collection Aid (Salimetrics LLC, Carlsbad, CA). A mixture of three healthy volunteer samples was used as the healthy volunteer sample. Three samples from each of the six periodontal disease patients were combined and designated periodontal disease patient sample 1 and periodontal disease patient sample 2. These samples were used for subsequent analysis. In the following, the healthy volunteer sample, periodontal disease patient sample 1, and periodontal disease patient sample 2 are collectively referred to as the "saliva sample."

[0204] (Isolation of membrane vesicles) The saliva samples were centrifuged at 7800 rpm at 4°C for 10 minutes to separate the precipitate and supernatant. The supernatant was then filtered through a membrane filter (pore size 0.22 μm) to remove any remaining bacterial matter. From the filtered supernatant, membrane vesicles were isolated using an ExoBacteria OMV Isolation Kit (System Biosciences, CA, USA).

[0205] (Removal of nucleic acid components outside membrane vesicles) The separated membrane vesicle samples were treated with deoxyribonuclease (DNase) to degrade nucleic acid components present outside the membrane vesicles. Specifically, 2 μL of DNase (13 units (U) / μL) was added to 100 μL of purified membrane vesicle sample, and the mixture was treated at 37°C for 30 minutes and then at 80°C for 10 minutes.

[0206] (Concentration adjustment and particle number determination) The membrane vesicle sample was diluted to a concentration of 40,000 particles / μL. The number of membrane vesicles was quantified by irradiating the particles with a laser, tracking the Brownian motion of each particle from the scattered light (tracking method), and calculating the particle diameter and number based on the Stokes-Einstein equation from their diffusion rate. Zetaview (DKSH, Germany) was used for measurements and calculations.

[0207] (Single particle analysis) The concentration was adjusted so that each membrane vesicle was encapsulated individually, and the membrane vesicle sample was mixed with agarose gel, and each membrane vesicle was encapsulated in one gel bead.

[0208] The encapsulated membrane vesicles were dissolved by adding the following reagents. 50U / μL Ready-lyse Lysozyme Solution (Epicentre), 2U / mL Zymolyase (Zymo research), 22U / mL lysostaphin (MERCK), 250U / mL mutanolysin (MERCK), This treatment was carried out overnight at 37°C in DPBS (Dulbecco's Phosphate-Buffered Saline).

[0209] Next, the cells were treated with 0.5 mg / mL achromopeptidase (MERCK) in DPBS at 37°C for 8 hours. Finally, the sample was treated overnight at 40°C with a solution of 1 mg / mL protease (Proteinase K (Promega)) and 0.5% SDS (Sodium dodecylsulfate) dissolved in DPBS.

[0210] The DNA in the dissolved gel beads was amplified using the REPLI-g Single Cell Kit (QIAGEN). After DNA amplification, the gel bead DNA molecules were stained with a nucleic acid staining reagent (1x SYBR Green). Gel beads that showed a certain level of fluorescence were separated and collected individually into a 96-well plate using a cell sorter (FACSMelody cell sorter (BD Bioscience)).

[0211] The samples were then aliquoted onto well plates and subjected to a second amplification using the REPLI-g Single Cell Kit. An Illumina library was then prepared using the Nextera XT DNA Library Prep Kit (Illumina), followed by sequencing using the Illumina Miseq or Hiseq to obtain short-read sequences. These procedures were performed on 192 particles for the healthy volunteer sample and 576 particles for the patient sample, and sequences were obtained for each particle.

[0212] (Data Processing) Among the obtained short read sequences (approximately 150 bp), sequences with low sequence quality were removed using fastp (Chen et al., fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics 34, 17, 2018).

[0213] These short read sequences were then assembled using SPAdes (Bankevic et al., SPAdes: A New Genome Assembly Algorithm and Its Applications to Single-Cell Sequencing. J. Comput. Biol., 19, 5, 2012) to create long base sequences called contigs, which are several kbp to several tens of kbp in length.

[0214] The protein-coding sequences (CDS) within the base sequences inside each membrane vesicle were predicted and detected using Prokka (Seemann, Prokka: rapid prokaryotic genome annotation. Bioinformatics, 30, 14, 2014). For each CDS, a homology search was performed using the NCBI all protein sequence nr database and the GTDB database as references. The bacterium with the most CDS regions in its genome with the highest degree of match was determined to be the host bacterium for that particle.

[0215] For homology searches, we used diamond (Buchfunk et al., Sensitive protein alignments at tree-of-life scale using DIAMOND, Nat. Methods., 18, 2021). Since membrane vesicles containing virus-derived CDS were also detected, we obtained information on the viral DNA affiliation (viral species) for these membrane vesicles.

[0216] The host bacterial profiles of membrane vesicles obtained as a result of whole particle analysis were compared between healthy and patient samples, and the phylum Patescibacteria was found to be the bacterial taxonomic group with the highest abundance in the patient samples.

[0217] Figure 20 shows the results of a comparison of the base sequence profiles of membrane vesicles in samples from healthy individuals and samples from periodontal disease patients. Based on the information on the detected base sequences, the bacterial taxonomic group from which the base sequences inside each membrane vesicle originated was identified. Figure 20 shows a heat map showing the frequency of the length of the detected DNA region when each bacterial strain is divided into the phylum level (phylum) and genus level (genus). It was found that base sequences derived from TM7x, which belongs to the Patescibacteria phylum, were frequently detected in samples from periodontal disease patients. Among the host bacterial strains belonging to the above categories, bacterial strain A (Patescibacteria TM7x sp900555265), which was detected most frequently (had the highest abundance ratio), was determined to be the membrane vesicle host bacterial strain specific to periodontal disease patients.

[0218] We also performed a similar comparison of virus-derived sequences. Figure 21 shows the results, showing the region lengths of the virus-derived base sequences detected in each membrane vesicle by species. Based on this comparison, virus species B (Podviridae ctUiB3), which was detected most frequently in periodontal disease patients compared to healthy individuals, was determined to be a virus specific to periodontal disease patients.

[0219] The sequence reads from membrane vesicles determined to be derived from bacterial strain A (or viral species B) were mapped to the whole genome sequence (draft genome sequence) of the bacterial strain A (or viral species B), and the genome regions that were detected in significantly more membrane vesicles were determined to be frequently occurring DNA regions in those membrane vesicles (hyper-geometric distribution test: p-value < 1.0 -6 ).

[0220] Next, a homology search was performed using the frequently occurring regions as query sequences and the protein sequences (nr) of organisms across all domains registered in NCBI as reference sequences (diamond). Among the protein-coding sequences in the database that matched the hits, the sum of the alignment scores (Max Score) of the top 20 protein-coding sequences, excluding sequences from the taxonomic group (phylum for bacteria, family for viruses) to which the same strain or species belongs, was calculated, and only the sequences with the smallest value were selected as biomarker candidate sequences.

[0221] Next, a homology search was performed on the candidate biomarker sequences against contigs obtained from membrane vesicles of healthy individuals, and the sum of the alignment scores (Max Score) of the top 20 protein-coding sequences was calculated, and the sequence with the smallest value was selected.

[0222] Furthermore, among the host bacteria of membrane vesicles detected in healthy individuals, sequence reads derived from membrane vesicles hosting bacterial strains belonging to the same taxonomic group (phylum level) as the above-mentioned bacterial strain A were mapped to the genome of bacterial strain A. By excluding the regions detected in the sequence reads of these membrane vesicles derived from healthy individuals from the above-mentioned biomarker candidate sequences, the final biomarker sequences for detecting periodontal disease were obtained.

[0223] FIG. 22 is a diagram showing the procedure for determining a region (biomarker sequence) that is specifically detected in periodontal disease patients. Figure 22(A) shows an example of the results of mapping the internal base sequence of membrane vesicles determined to be derived from TM7x sp900555265 to the reference sequence (the genome sequence of TM7x sp900555265) among membrane vesicles derived from periodontal disease patient samples. The vertical axis of Figure 22(A) represents individual membrane vesicles, and the horizontal axis represents their position on the genome sequence. The heat map in Figure 22(B) shows the genomic regions detected with high frequency in Figure 22(A). The horizontal axis represents the position on the genome sequence.

[0224] In contrast, Figure 22(C) shows membrane vesicles derived from healthy individuals that were determined to be derived from Patescibacteria, including TM7x sp900555265, mapped to the reference sequence (the genome sequence of TM7x sp900555265). The vertical and horizontal axes are the same as those in Figure 22(A). The heat map in Figure 22(D) shows the genomic regions detected with high frequency in Figure 22(C). The horizontal axis represents the position on the genome sequence.

[0225] Figure 22(E) shows a map of genomic regions detected specifically in periodontal disease patients, based on a comparison of Figures 22(B) and 22(D). Each of these regions represents a biomarker sequence.

[0226] Similarly, for viral sequences, membrane vesicle-derived sequence reads containing CDS derived from viruses belonging to the same taxonomic group (family level) as virus B among the virus-derived sequences detected in healthy individuals were mapped to the virus B genome. Since no regions were detected in the sequence reads of membrane vesicles derived from healthy individuals, the biomarker candidate sequences obtained above were used as biomarker sequences. Of these sequences, contiguous DNA regions of 1000 bp or more were targeted.

[0227] FIG. 23 is a diagram showing the procedure for determining a region (biomarker sequence) that is specifically detected in periodontal disease patients. Figure 23(B) shows an example of the results of mapping the internal base sequence of membrane vesicles containing a CDS region determined to be derived from Podviridae ctUiB3 to the reference sequence (the genomic sequence of ctUiB3) among membrane vesicles derived from a periodontal disease patient sample. The vertical axis in Figure 23(B) represents individual membrane vesicles, and the horizontal axis represents their position on the genomic sequence. Figure 23(A) shows the detection frequency for each region, and the heat map in Figure 23(C) shows the genomic regions detected with high frequency in Figure 23(A). In both cases, the horizontal axis represents the position on the genomic sequence. This genomic region becomes the biomarker sequence.

[0228] (Selection of primer sequences and oligonucleotide probe sequences) From the obtained biomarker sequences, several sequences of approximately 100 bp suitable for detection by qPCR using the probe method (5'-nuclease method) were selected. The Primer Quest (registered trademark) Tool (https: / / sg.idtdna.com / pages / tools / primerquest) provided by Integrated DNA Technologies was used to select these target sequences and design primer and oligonucleotide probe sequences.

[0229] By using this primer pair and probe in combination with the "TaqMan (registered trademark) Gene Expression Master Mix" provided by ThermoFisher Scientific, an Applied Biosystems brand, it is possible to detect membrane vesicle-derived base sequences specific to periodontal disease patients, i.e., periodontal disease, using saliva samples. [Industrial Applicability]

[0230] In recent years, membrane vesicles secreted by bacteria have attracted attention as one of the factors mediating pathogenic bacteria and various diseases. It has been pointed out that these vesicles, when migrated into the bloodstream and transported throughout the body, may potentially affect a wide range of diseases, from digestive system disorders and heart disease to neurological disorders such as Alzheimer's. In other words, identification and detection technologies for blood membrane vesicles, which have a strong influence on various diseases, will serve as the foundation for innovative disease prevention and diagnostic technologies. By conducting (q)PCR tests (or RT-qPCR) using the bacterial membrane vesicle marker sequences identified by this invention as indicators, specific pathogenic bacterial species and diseases can be easily detected.

[0231] Although there have been technologies to determine and analyze the base sequences inside membrane vesicles, there has been no technology to detect membrane vesicles derived from specific bacteria by amplifying only the highly specific base sequences contained inside. The reason for this is that there have been no algorithms to determine high-frequency and highly specific base sequence regions among the base sequences inside membrane vesicles, and there have been no examples of detecting membrane vesicles contained in human samples using base sequences determined using such algorithms. The present invention achieves the above two goals. [Explanation of symbols]

[0232] 20, 50: membrane vesicle detection device, 21: base sequence analysis device, 22: quantitative PCR device, 23: control device, 24: processor, 25: storage device, 26: display device, 27: input device, 28: preprocessing device, 29: sequencer, 31: map unit, 32: frequent region selection unit, 33: biomarker candidate sequence selection unit, 34: biomarker sequence determination unit, 51: nucleic acid synthesizer, 60: computer

Claims

1. detecting membrane vesicles in a biological sample by quantitative polymerase chain reaction using primers designed to amplify a biomarker sequence that is part of the base sequence of membrane vesicles produced by the host cell; The biomarker sequences are prepared by the following steps: Isolating and obtaining the membrane vesicles produced by the host cells from the biological sample; Obtaining the base sequence of the separated membrane vesicles; mapping the nucleic acid sequence to a reference sequence derived from the host cell; Among the mapped regions of the reference sequence, a statistical test is performed to select a region whose detection frequency is statistically significantly higher than that of other regions as a frequent region; performing a homology search using the frequently occurring region as a query sequence, and selecting the frequently occurring regions up to a predetermined rank as biomarker candidate sequences when the alignment scores to the detected homologous sequences derived from other than the host cell are sorted in ascending order; aligning the biomarker candidate sequence with the homologous sequence, calculating sequence conservation in each region of 50 to 250 consecutive bases, detecting regions where the obtained sequence conservation is lower than a standard, and designating the regions as the biomarker sequences; A method for detecting membrane vesicles, selected by

2. The method for detecting membrane vesicles according to claim 1 , wherein the reference sequence is a genomic sequence of the host cell or a nucleic acid sequence derived from a virus that has infected the host cell.

3. The method for detecting membrane vesicles according to claim 1, wherein the biological sample is collected from a subject.

4. Isolating and obtaining membrane vesicles produced by host cells from a biological sample; Obtaining the base sequence of the separated membrane vesicles; mapping the nucleotide sequence to a reference sequence derived from the host cell; Among the mapped regions of the reference sequence, a statistical test is performed to select a region whose detection frequency is statistically significantly higher than that of other regions as a frequent region; performing a homology search using the frequently occurring region as a query sequence, and selecting the frequently occurring regions up to a predetermined rank as biomarker candidate sequences when the alignment scores to the detected homologous sequences derived from other than the host cell are sorted in ascending order; aligning the biomarker candidate sequence with the homologous sequence, calculating sequence conservation in each region of 50 to 250 consecutive bases, detecting regions where the obtained sequence conservation is lower than a standard, and designating the regions as the biomarker sequences; and detecting said membrane vesicles in said biological sample by quantitative polymerase chain reaction using primers designed to amplify said biomarker sequences.

5. The method for detecting membrane vesicles according to claim 4 , wherein the reference sequence is a genomic sequence of the host cell or a nucleic acid sequence derived from a virus that has infected the host cell.

6. The method for detecting membrane vesicles according to claim 4, wherein the biological sample is collected from a subject.

7. The method for detecting membrane vesicles according to any one of claims 1 to 6, wherein the biological sample is at least one selected from the group consisting of saliva, blood, serum, plasma, buffy coat, lymph, interstitial fluid, body cavity fluid, digestive fluid, sweat, urine, nasal discharge, tears, semen, vaginal fluid, amniotic fluid, milk, sputum, surgical irrigation fluid, feces, and swabs and swabs of skin or body mucosa.

8. The method for detecting membrane vesicles according to any one of claims 1 to 6, wherein the selection of the candidate biomarker sequence is performed by selecting the sequence that has the smallest alignment score to the homologous sequence.

9. The method for detecting membrane vesicles according to any one of claims 1 to 6, wherein the frequent region is a protein coding region.

10. The method for detecting membrane vesicles according to any one of claims 1 to 6, wherein the quantitative polymerase chain reaction is a real-time polymerase chain reaction.

11. In the selection of the frequent region, A method for detecting membrane vesicles described in any one of claims 1 to 6, wherein a statistical test is performed to determine whether the detection frequency of a target region among the mapped regions of the reference sequence is equivalent to the detection frequency of other regions, and if the significance probability (p-value) obtained by the statistical test is less than 0.01, the target region is selected as the frequently occurring region.

12. Isolating and obtaining membrane vesicles produced by host cells from a biological sample; Obtaining the base sequence of the separated membrane vesicles; mapping the nucleotide sequence to a reference sequence derived from the host cell; Among the mapped regions of the reference sequence, a statistical test is performed to select a region whose detection frequency is statistically significantly higher than that of other regions as a frequent region; performing a homology search using the frequently occurring region as a query sequence, and selecting the frequently occurring regions up to a predetermined rank as biomarker candidate sequences when the alignment scores to the detected homologous sequences derived from other than the host cell are sorted in ascending order; A method for searching for biomarker sequences, comprising: aligning the biomarker candidate sequence to the homologous sequence; calculating sequence conservation in each region of a continuous 50 to 250 base region; detecting a region where the obtained sequence conservation is lower than a standard; and designating the region as the biomarker sequence.

13. The method for searching for a biomarker sequence according to claim 12, wherein the reference sequence is a genomic sequence of the host cell or a nucleic acid sequence derived from a virus that has infected the host cell.

14. The method for searching for a biomarker sequence according to claim 12, wherein the biological sample is collected from a subject.

15. The method for searching for a biomarker sequence according to claim 14, wherein the biological sample is at least one selected from the group consisting of saliva, blood, serum, plasma, buffy coat, lymph, interstitial fluid, body cavity fluid, digestive fluid, sweat, urine, nasal discharge, tears, semen, vaginal fluid, amniotic fluid, milk, sputum, surgical lavage fluid, feces, and swabs and wipes of skin or body mucosa.

16. The method for searching for a biomarker sequence according to any one of claims 12 to 15, wherein the selection of the biomarker candidate sequence is performed by selecting a sequence having the smallest alignment score to the homologous sequence.

17. The method for searching for a biomarker sequence according to any one of claims 12 to 15, wherein the frequently occurring region is a protein coding region.

18. In the selection of the frequent region, The method for searching for a biomarker sequence according to any one of claims 12 to 15, wherein the statistical test is performed to determine whether the detection frequency of a target region among the mapped regions of the reference sequence is equivalent to the detection frequency of other regions, and if the significance probability (p-value) obtained by the statistical test is less than 0.01, the target region is selected as the frequent region.

19. a base sequence analyzer for obtaining the base sequence of the membrane vesicles produced by the host cell; a quantitative PCR device for detecting the membrane vesicles in a biological sample by quantitative polymerase chain reaction using primers designed to amplify a biomarker sequence that is a part of the base sequence; a control device; The control device a mapping unit that maps the base sequence obtained by the base sequence analyzing device to a reference sequence derived from the host cell; a frequent region selection unit that performs a statistical test to select, from the mapped regions of the reference sequence, a region whose detection frequency is statistically significantly higher than that of other regions, as a frequent region; a biomarker candidate sequence selection unit that performs a homology search using the frequently occurring region as a query sequence, and selects the frequently occurring regions up to a predetermined rank as biomarker candidate sequences when the alignment scores to detected homologous sequences derived from sequences other than the host cell are sorted in ascending order; and a biomarker sequence determination unit that aligns the biomarker candidate sequence with the homologous sequence, calculates sequence conservation in each region of a continuous 50 to 250 base region, detects regions where the obtained sequence conservation is lower than a standard, and defines these regions as the biomarker sequence.

20. A membrane vesicle detection device having a base sequence analyzer, a quantitative PCR device, and a control device that is a computer, Controlling the base sequence analyzer to obtain the base sequence of the membrane vesicles produced by the host cell; mapping the nucleotide sequence to a reference sequence derived from the host cell; A step of selecting, from the mapped regions of the reference sequence, a region whose detection frequency is statistically significantly higher than that of other regions by performing a statistical test as a frequent region; a step of performing a homology search using the frequently occurring region as a query sequence, and selecting the frequently occurring regions up to a predetermined rank as biomarker candidate sequences when the alignment scores to the detected homologous sequences derived from other than the host cell are sorted in ascending order; aligning the biomarker candidate sequence with the homologous sequence, calculating sequence conservation in each region of 50 to 250 consecutive bases, detecting regions where the obtained sequence conservation is lower than a standard, and designating the regions as biomarker sequences; and a program for controlling the quantitative PCR instrument to detect the membrane vesicles in a biological sample by quantitative polymerase chain reaction using primers designed to amplify the biomarker sequence.

21. a base sequence analyzer for obtaining the base sequence of the membrane vesicles produced by the host cell; a control device; The control device includes a mapping unit that maps the base sequence to a reference sequence derived from the host cell; a frequent region selection unit that performs a statistical test to select, from the mapped regions of the reference sequence, a region whose detection frequency is statistically significantly higher than that of other regions, as a frequent region; a biomarker candidate sequence selection unit that performs a homology search using the frequently occurring region as a query sequence, and selects the frequently occurring regions up to a predetermined rank as biomarker candidate sequences when the alignment scores to detected homologous sequences derived from sequences other than the host cell are sorted in ascending order; and a biomarker sequence determination unit that aligns the biomarker candidate sequence to the homologous sequence, calculates sequence conservation in each region of a continuous 50 to 250 base region, detects regions where the obtained sequence conservation is lower than a standard, and sets the detected regions as biomarker sequences.

22. A biomarker sequence searching device having a base sequence analyzing device and a control device which is a computer, Controlling the base sequence analyzer to obtain the base sequence of the membrane vesicles produced by the host cell; mapping the nucleotide sequence to a reference sequence derived from the host cell; A step of selecting, from the mapped regions of the reference sequence, a region whose detection frequency is statistically significantly higher than that of other regions by performing a statistical test as a frequent region; a step of performing a homology search using the frequently occurring region as a query sequence, and selecting the frequently occurring regions up to a predetermined rank as biomarker candidate sequences when the alignment scores to the detected homologous sequences derived from other than the host cell are sorted in ascending order; a step of aligning the biomarker candidate sequence with the homologous sequence, calculating sequence conservation in each region of a continuous 50 to 250 base pair, detecting a region where the obtained sequence conservation is lower than a standard, and designating the region as a biomarker sequence.

23. 1. A computer for searching biomarker sequences, comprising: a mapping unit that maps the base sequence of a membrane vesicle produced by a host cell to a reference sequence derived from the host cell; a frequent region selection unit that performs a statistical test to select, from the mapped regions of the reference sequence, a region whose detection frequency is statistically significantly higher than that of other regions, as a frequent region; a biomarker candidate sequence selection unit that performs a homology search using the frequently occurring region as a query sequence, and selects the frequently occurring regions up to a predetermined rank as biomarker candidate sequences when the alignment scores to detected homologous sequences derived from sequences other than the host cell are sorted in ascending order; and a biomarker sequence determination unit that aligns the biomarker candidate sequence to the homologous sequence, calculates sequence conservation in each region of a continuous 50 to 250 base region, detects regions where the obtained sequence conservation is lower than a standard, and designates the detected regions as biomarker sequences.

24. On the computer, mapping the nucleotide sequence of membrane vesicles produced by the host cell to a reference sequence derived from said host cell; A step of selecting, from the mapped regions of the reference sequence, a region whose detection frequency is statistically significantly higher than that of other regions by performing a statistical test as a frequent region; a step of performing a homology search using the frequently occurring region as a query sequence, and selecting the frequently occurring regions up to a predetermined rank as biomarker candidate sequences when the alignment scores to the detected homologous sequences derived from other than the host cell are sorted in ascending order; a step of aligning the biomarker candidate sequence with the homologous sequence, calculating sequence conservation in each region of a continuous 50 to 250 base pair, detecting a region where the obtained sequence conservation is lower than a standard, and designating the region as a biomarker sequence.

25. comparing the base sequence of a first membrane vesicle obtained from a first biological sample collected from a diseased patient with a base sequence of a known cell or a genome sequence of a known virus to determine host cell A, which is the host cell of the first membrane vesicle; comparing the base sequence of the second membrane vesicle obtained from a second biological sample collected from a healthy individual with the base sequence of a known cell or the genome sequence of a known virus to determine host cell B, which is the host cell of the second membrane vesicle; Identifying a host cell C that is the host cell having a higher abundance ratio in the first biological sample by comparing the host cell A with the host cell B; Mapping the base sequence of the membrane vesicles produced by the host cell C to a reference sequence derived from the host cell C; Among the mapped regions of the reference sequence, a statistical test is performed to select a region whose detection frequency is statistically significantly higher than that of other regions as a frequent region; performing a homology search using the frequently occurring region as a query sequence, and selecting the frequently occurring regions up to a predetermined rank as biomarker candidate sequences when the alignment scores to the detected homologous sequences derived from other than the host cell are sorted in ascending order; a method for searching for a disease biomarker, the method comprising: aligning the biomarker candidate sequence with the homologous sequence; calculating sequence conservation in each region of a continuous 50 to 250 base region; detecting a region where the obtained sequence conservation is lower than a standard; and designating the region as a disease biomarker.

26. selecting a host cell D from the host cells B, the host cell D being a closely related species to the host cell C; Mapping the base sequence or a fragment thereof of the membrane vesicle derived from the host cell D to the reference sequence; The method for discovering disease biomarkers according to claim 25, further comprising: excluding the mapped region from the biomarker candidate sequence.

27. The method of claim 26, wherein the biomarker sequence is determined by: The method for detecting membrane vesicles according to any one of claims 1 to 6, further comprising: aligning the biomarker candidate sequence with the homologous sequence; calculating entropy based on the frequency of appearance of the four bases adenine (A), thymine (T), guanine (G), and cytosine (C) for each position on the sequence in the region of 50 to 250 consecutive bases; determining an average value of the entropy in the region; and determining that the region has lower sequence conservation than a standard if the average value exceeds a predetermined value, and designating the region as the biomarker sequence.

28. In selecting the biomarker candidate sequence, The method for detecting membrane vesicles according to any one of claims 1 to 6, wherein the homology search is performed on one database selected from the group consisting of NCBI RefSeq, NCBI GenBank, UCSC Genome Browser, GRCh37 reference primary assembly, and JRGv2.

Citation Information

Patent Citations

  • Biomarker for use in the diagnosis of intraductal papillary mucinous neoplasm or pancreatic cancer, and method of inspecting intraductal papillary mucinous neoplasm or pancreatic cancer using the biomarker

    JP2017158545A

  • Sequencing and analysis of exosome-associated nucleic acids

    JP2019535307A

  • Lung cancer diagnosis method using bacterial metagenomic analysis

    JP2020503028A

  • Microparticles from streptococcus pneumoniae as vaccine antigens

    US20190328861A1

  • Method for non-invasive prenatal diagnosis based on exosomal DNA and application thereof

    US20190345482A1