Nucleic acid analysis method, analysis method, nucleic acid analysis program, information recording medium, and information processing system
By forming circular nucleic acids through random DNA fragment linking and decoding their sequences, the method addresses amplification bias, enabling the analysis of low-frequency nucleic acids with enhanced accuracy.
Patent Information
- Application Number
- PCT/JP2025/009650
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-29
- Filing Date
- 2025-03-13
- Publication Date
- 2025-10-02
Smart Images

Figure JP2025009650_02102025_PF_FP_ABST
Abstract
Description
Nucleic acid analysis method, analysis method, nucleic acid analysis program, information recording medium, and information processing system
[0001] The present technology relates to a nucleic acid analysis method, an analysis method, a nucleic acid analysis program, an information recording medium, and an information processing system. More specifically, the present technology relates to a method for suitably detecting nucleic acids that are present at a relatively low frequency relative to all nucleic acids contained in a subject without being affected by amplification bias based on the expression frequency of the nucleic acid.
[0002] Conventionally, methods for determining base sequences derived from mRNA have been known.
[0003] For example, Non-Patent Document 1 below discloses a method for determining the mRNA-derived base sequence contained in a circular nucleic acid by binding a group of cDNAs derived from mRNA to form a circular nucleic acid.
[0004] A method for increasing throughput of single molecule sequencing by concatenating short DNA fragments, Scientific Reports, 2017
[0005] For example, when analyzing the expression status of nucleic acids, it is desirable to be able to analyze all nucleic acids, including those expressed at low frequencies.
[0006] Therefore, an object of the present technology is to provide a technology that can suitably analyze all nucleic acids without being affected by amplification bias based on the expression frequency of nucleic acids.
[0007] As a result of extensive research, the present inventors have found that all nucleic acids, including nucleic acids expressed at low frequencies, can be suitably analyzed by randomly combining multiple types of DNA fragments to form multiple circular nucleic acids, determining the base sequences of the circular nucleic acids, and identifying the repeating units from the base sequences.
[0008] That is, the present technology provides a nucleic acid analysis method including a base sequence decoding step for decoding the base sequences of multiple types of circular nucleic acids formed by randomly linking multiple types of DNA fragments, a repeat unit identification step for identifying the multiple types of circular nucleic acids by determining the number of bases of the circular nucleic acid from the base sequence decoded by the base sequence decoding step, and a DNA fragment identification step for identifying the regions of DNA fragments contained in the multiple types of circular nucleic acids identified by the repeat unit identification step. In this case, the multiple types of DNA fragments may be composed of all DNA fragments contained in the target of analysis, or the multiple types of DNA fragments may be cDNA fragments derived from mRNA. Furthermore, if the multiple types of DNA fragments are cDNA fragments derived from mRNA, the mRNA may be derived from a single cell. Furthermore, the nucleic acid analysis method of the present technology may include a duplicate elimination step after the DNA fragment identification step for comparing the order of the DNA fragments contained in the multiple types of circular nucleic acids to eliminate circular nucleic acids with the same order. In this case, the identification of the number of bases in the repeat unit identifying step may be performed by analyzing the period of a maximum value occurring in an autocorrelation sequence of the decoded base sequence, and the autocorrelation sequence may be weighted using a transition of an average value calculated from data within a specific range. In the nucleic acid analysis method of the present technology, the identification of the number of bases in the repeat unit identifying step may be performed by a maximum value occurring in a component distribution of a period of a spectral sequence obtained by spectral analysis of the decoded base sequence, and the spectral sequence may be weighted using a transition of an average value calculated from data within a specific range. In the nucleic acid analysis method of the present technology, the identification of the number of bases in the repeat unit identifying step may be performed by a maximum value occurring in an autocorrelation sequence of the decoded base sequence that coincides with a maximum value occurring in a component distribution of a period of a spectral sequence obtained by spectral analysis of the decoded base sequence. The nucleic acid analysis method of the present technology may include, after the DNA fragment identifying step, a base sequence determination step of comparing the base sequences of multiple DNA fragments identified as identical DNA fragments to determine the base sequences of the identical DNA fragments.In the nucleic acid analysis method of the present technology, in the DNA fragment identification step, the region of the DNA fragment may be identified by identifying the sequence of a linking portion that links the DNA fragments in the circular nucleic acid. Next, the present technology provides a method for analyzing the expression level of mRNA in a single cell using the nucleic acid analysis method of the present technology. In this case, the single cell may be a cell sorted by a biological sample analyzer.
[0009] Furthermore, the present technology provides a nucleic acid analysis program that performs a base sequence decoding step to decode the base sequences of multiple types of circular nucleic acids formed by randomly linking multiple types of DNA fragments, a repeat unit identification step to identify the multiple types of circular nucleic acids by determining the number of bases of the circular nucleic acid from the base sequence decoded by the base sequence decoding step, and a DNA fragment identification step to identify the regions of DNA fragments contained in the multiple types of circular nucleic acids identified by the repeat unit identification step. After performing the DNA fragment identification step, the nucleic acid analysis program of the present technology may also perform a duplicate elimination step to compare the arrangement order of the DNA fragments contained in the multiple types of circular nucleic acids and eliminate circular nucleic acids with the same order. Next, the present technology provides an information recording medium that implements the nucleic acid analysis program of the present technology. Furthermore, the present technology provides an information processing system that implements the nucleic acid analysis program of the present technology.
[0010] 1 shows an example of a flow diagram of a nucleic acid analysis method according to the present technology. FIG. 2 shows a modified example of the flow diagram of the nucleic acid analysis method according to the present technology. FIG. 3 shows an image of the flow of a base sequence decoding process. FIG. 4 shows a schematic diagram illustrating an example of the configuration of a nucleic acid linker. FIG. 5 shows a schematic diagram of a modified example of the configuration of a nucleic acid linker. FIG. 6 shows a schematic diagram of a modified example of the configuration of a nucleic acid linker. FIG. 7 is a schematic diagram for explaining an example of a process for generating a circular nucleic acid using a nucleic acid linker. FIG. 8 is a schematic diagram for explaining the structure of mRNA. FIG. 9 is a schematic diagram for explaining an example of sequencing. FIG. 10 is a schematic diagram for explaining bias in exponential amplification. FIG. 11 is an image diagram of a nucleic acid linker etc. used in a first modified example of a process for generating a circular nucleic acid using a nucleic acid linker. FIG. 12 is a schematic diagram for explaining a first modified example of a process for generating a circular nucleic acid using a nucleic acid linker. FIG. 13 shows an example of a flow diagram for identifying a repeat unit of a circular nucleic acid based on the autocorrelation sequence of a base sequence. FIG. 14 is a schematic diagram for explaining the relationship between deviation from an original signal and a calculated autocorrelation coefficient using an example of a periodic function. FIG. 15 is a schematic diagram for explaining the relationship between deviation from an original signal and a calculated autocorrelation coefficient using an example of a periodic function. FIG. 1 is a schematic diagram illustrating the relationship between the deviation from the original signal and the calculated autocorrelation coefficient, using an example of a periodic function. FIG. 2 is a schematic diagram illustrating the relationship between the deviation from the original signal and the calculated autocorrelation coefficient, using an example of a periodic function. FIG. 3 is a schematic diagram illustrating the relationship between the deviation from the original signal and the calculated autocorrelation coefficient, using an example of a periodic function. FIG. 4 is an image of a graph of autocorrelation calculated from the deviation of a repeating unit of a circular nucleic acid from the signal of its base sequence. FIG. 5 is a schematic diagram illustrating the influence of the deviation of a repeating unit of a circular nucleic acid from the signal of its base sequence on the calculated autocorrelation. FIG. 6 shows a specific example of identifying the number of bases in a circular nucleic acid from its base sequence information. FIG. 7 is an image of calculating the weight of the original signal and weighting the original signal with the calculated weight. FIG. 8 is an example of calculating the weight of the original signal and weighting the original signal with the calculated weight. FIG. 9 is an example of a flow diagram for identifying a repeating unit of a circular nucleic acid based on the component distribution of the period of a spectral sequence obtained by spectral analysis of a base sequence. FIG. 10 is an image of identifying a repeating unit based on the component distribution of the period of a spectral sequence obtained by spectral analysis, using an example of a periodic function.1 is an example of identifying a repeat unit from an original signal based on the component distribution of the period of a spectral sequence obtained by spectral analysis.
[0023] FIG. 1 shows an example of a flow diagram for identifying a repeat unit of a circular nucleic acid based on the autocorrelation sequence of a base sequence and the component distribution of the period of a spectral sequence obtained by spectral analysis of the base sequence.
[0024] FIG. 1 shows data measuring the effect of the number of RCAs on identifying a repeat unit.
[0025] FIG. 1 is an image diagram of dividing the base sequence information of an identified circular nucleic acid into units of DNA fragments.
[0026] FIG. 1 is an image diagram of narrowing down comparison targets for circular nucleic acids based on the identified DNA fragments.
[0027] FIG. 1 is an image diagram of eliminating duplicate circular nucleic acids based on the arrangement order of the identified DNA fragments.
[0028] FIG. 1 is an image diagram of determining the bases of a base sequence.
[0029] FIG. 2 is an image diagram of determining the base sequence of each DNA fragment unit.
[0030] FIG. 2 is an image diagram of assigning an identifier to each cell in single-cell analysis.
[0031] FIG. 2 is an image diagram of an example of a microchannel for forming an emulsion.
[0032] FIG. 3 is an image diagram of an example of a bioparticle sorting device used to form an emulsion.
[0033] FIG. 4 is a schematic diagram showing the formation of emulsion particles in a microchip and the isolation of bioparticles within the formed emulsion particles. FIG. 1 is a diagram showing an enlarged view of a particle sorting unit included in a bioparticle sorting device. FIG. 2 is a schematic diagram for explaining an example of decomposing cells in an emulsion and modifying and extracting mRNA with an identifier. FIG. 3 is a schematic diagram for explaining an example of analyzing the expression level of mRNA for each cell. FIG. 4 is an image diagram of an antibody-nucleic acid molecule complex used when simultaneously performing single-cell analysis and analysis of secreted molecules on the cell surface. FIG. 5 is a diagram showing the overall configuration of a biological sample analyzing device.
[0011] Preferred embodiments of the present technology will be described below. However, the embodiments shown below are examples of typical embodiments of the present technology, and the present technology is not limited to only the preferred embodiments below and can be freely modified within the scope of the present technology.
[0012] [Nucleic acid analysis method] The nucleic acid analysis method of the present technology includes a base sequence decoding step of decoding the base sequences of multiple types of circular nucleic acids formed by randomly linking multiple types of DNA fragments into a circular shape, a repeat unit identifying step of identifying the multiple types of circular nucleic acids by determining the number of bases of the circular nucleic acid from the base sequence decoded by the base sequence decoding step, and a DNA fragment identifying step of identifying the regions of DNA fragments contained in the multiple types of circular nucleic acids identified by the repeat unit identifying step.This makes it possible to identify nucleic acid fragments that appear relatively low in frequency among all the multiple types of nucleic acid fragments contained in a subject without being affected by amplification bias based on the expression frequency of nucleic acid fragments such as DNA, and to preferably detect the expression level.
[0013] Figure 1 shows an example of a flow diagram of the nucleic acid analysis method of the present technology. In the nucleic acid analysis method of the present technology, as shown in Figure 1, operations proceed in the order of a base sequence decoding step S101, a repeat unit identification step S102, and a DNA fragment identification step S103. As shown in a modified example of the flow diagram in Figure 2, the nucleic acid analysis method of the present technology may further include a duplication removal step S014 and a base sequence determination step S105 after the DNA fragment identification step S103. Each step will be described in more detail.
[0014] <Base sequence decoding step S101> In the nucleic acid analysis method of the present technology, the base sequence decoding step involves randomly linking multiple types of DNA fragments to form multiple types of circular nucleic acids, and then decoding the base sequences of the resulting multiple types of circular nucleic acids.
[0015] Fig. 3 is a conceptual diagram of the flow of the base sequence decoding process. Note that Fig. 3 shows an example in which the target DNA fragment is cDNA synthesized using mRNA as a template, but the target DNA fragment is not limited to this.
[0016] First, <3A> in Fig. 3 shows an object in which multiple types of mRNA are expressed. As mentioned above, the object in this case is not limited, but for example, by using mRNA extracted by decomposing a single cell, it is possible to analyze the expression level of mRNA in each single cell.
[0017] Next, in <3B>, reverse transcription using the mRNA as a template is performed to form cDNA DNA fragments derived from the mRNA. For example, as described above, when mRNA extracted by disassembling a single cell is used, the transcriptome information of the cell can be analyzed from the multiple DNA fragments obtained. In this case, any method can be used for reverse transcription from mRNA, such as a method using reverse transcriptase.
[0018] Next, as shown in <3C>, the obtained multiple types of DNA fragments are randomly ligated to form multiple types of circular nucleic acids as shown in <3D>.
[0019] Here, "randomly linking DNA fragments" refers to randomly linking multiple types of DNA fragments. In particular, in this technology, it is preferable to randomly link all DNA fragments contained in the analysis target without being affected by the base sequence of the DNA fragments. In this technology, by randomly linking multiple types of DNA fragments contained in the analysis target, the DNA fragments can be arranged in a unique order for each of the obtained multiple types of circular nucleic acids. Note that the method for randomly linking multiple types of DNA fragments is not particularly limited, and any method can be used as long as it can randomly link multiple types of DNA fragments.
[0020] Next, as shown in <3E>, the base sequences of the obtained multiple types of circular nucleic acids are decoded. The method for decoding the base sequences in this case is not particularly limited, and any method can be used for decoding. The nucleic acid analysis method of the present technology can identify the nucleic acid fragments expressed in the target object and their amounts based on the base sequence information obtained by this decoding. In particular, by using mRNA extracted by decomposing a single cell as the target object and analyzing the base sequence information derived from the mRNA using the nucleic acid analysis method of the present technology, the expression level of mRNA in each single cell can be suitably analyzed.
[0021] (1) Example of a process for generating circular nucleic acid As described above, in the nucleic acid analysis method of the present technology, the method for randomly linking multiple types of DNA fragments is not particularly limited as long as it is a method that can randomly link multiple types of DNA fragments, but for example, by using a nucleic acid fragment (hereinafter referred to as a "linking portion") involved in linking, such as a nucleic acid linking portion described below, multiple types of DNA fragments can be randomly linked via the nucleic acid linking portion, etc. This will be specifically described below.
[0022] 4A is a schematic diagram showing an example of the configuration of a nucleic acid linker. The portion of the nucleic acid linker that is involved in linking two or more DNA fragments is also referred to as a concatenator in this specification. The nucleic acid linker 10 shown in this figure includes a target nucleic acid capturer 11 configured to capture the 3'-terminal region of a target nucleic acid, a complementary strand capturer 12 configured to capture the 3'-terminal region of the complementary strand of the target nucleic acid, and a double-stranded portion 13 that connects the target nucleic acid capturer 11 and the complementary strand capturer.
[0023] The nucleic acid linker may also have the configuration shown in FIG. 4B or 4C . Similar to the configuration example of FIG. 4A , the nucleic acid linker 10 shown in these figures includes a target nucleic acid capturer 11 configured to capture the 3′-terminal region of the target nucleic acid and a complementary strand capturer 12 configured to capture a terminal region of a sequence derived from the target nucleic acid, such as the 3′-terminal region of the complementary strand of the target nucleic acid, but does not have a double-stranded portion 13. The nucleic acid linker 10 shown in FIG. 4C includes a sequence addition portion 14 in addition to the configuration of FIG. 4B . Using a nucleic acid linker configured in this manner makes it possible to shorten the sequence introduced into the target nucleic acid or a circular nucleic acid having a sequence derived from the target nucleic acid, while randomly linking multiple types of DNA fragments via this nucleic acid linker, etc.
[0024] The target nucleic acid capturing unit 11 is configured to capture, for example, the 3'-terminal region of the target nucleic acid. For example, when the target nucleic acid is mRNA, since the mRNA has a poly-A tail at the 3'-terminal region, the target nucleic acid capturing unit 11 may be configured to capture the poly-A tail, and in particular, may have a base sequence configured to capture the poly-A tail.
[0025] In order for the target nucleic acid capture unit 11 to capture the poly-A tail, the target nucleic acid capture unit may have, for example, a poly-T sequence. The length of the poly-T sequence may be, for example, 10 to 50 bases, preferably 15 to 30 bases. That is, the poly-T sequence may be composed of, for example, 10 to 50 T bases, preferably 15 to 30 T bases. The target nucleic acid capture unit 11 may be composed solely of a poly-T sequence. Alternatively, the target nucleic acid capture unit 11 may be single-stranded DNA or RNA. This facilitates binding to the target nucleic acid, particularly complementary binding. The length of the poly-T sequence may be longer or shorter, and may be varied, for example, depending on the length of the poly-A tail of the target nucleic acid. Furthermore, when the sequence of RNA or target RNA without a poly-T sequence is known, a random sequence (random primer, random hexamer, etc.) or a sequence that specifically binds to the target RNA may be used as the target nucleic acid capture unit. The length of the random sequence may be, for example, 6 to 20 bases, preferably 6 to 10 bases, and the length of the sequence that specifically binds to the target RNA may be 10 to 40 bases, preferably 15 to 35 bases.
[0026] The complementary strand capture unit 12 is configured to capture, for example, the terminal region of a target nucleic acid or a nucleic acid having a sequence derived from the target nucleic acid (e.g., the 3'-terminal region of the complementary strand of such a nucleic acid). In particular, the complementary capture unit of one nucleic acid linker can capture the 3'-terminal region of a complementary strand of a target nucleic acid (a complementary strand generated by cDNA synthesis of the other target nucleic acid) other than the target nucleic acid captured by the target nucleic acid capture unit of that nucleic acid linker in the process of generating a circular nucleic acid according to the present disclosure. This allows two or more target nucleic acids or complementary strands of sequences derived from the target nucleic acids to be linked via nucleic acid linkers. Here, "linking via nucleic acid linkers" is not limited to linking via nucleic acid linkers, but also includes linking via a sequence derived from a nucleic acid linker.
[0027] In the process for generating circular nucleic acids according to the present disclosure, the complementary nucleic acid generation step is a step of generating two or more complementary strands to be linked in the subsequent circular nucleic acid generation step. More specifically, complementary strands of one or more target nucleic acids are generated with a nucleic acid linker bound to one end of each of the one or more target nucleic acids. Note that, as described below, the complementary nucleic acid generation step may use the complementary strand of the generated target nucleic acid as a template to generate a nucleic acid having a sequence derived from the target nucleic acid, or may provide a sequence to be captured by a complementary strand capture unit provided in the nucleic acid linker.
[0028] For example, when the target nucleic acid is mRNA, a CCC sequence (C: cytosine) is generated at the 3' end of a complementary strand produced by reverse transcription of the mRNA by the reverse transcriptase that performs the reverse transcription. Therefore, the complementary strand capture unit 12 may be configured to capture the CCC sequence, and may particularly have a base sequence configured to capture the CCC sequence. Thus, the complementary strand capture unit may have a base sequence complementary to the base sequence added to the 3' end during reverse transcription by the reverse transcriptase.
[0029] On the other hand, when a nucleic acid linker not having a double-stranded portion is used, the complementary strand capture portion of the nucleic acid linker may have any sequence, and a primer having a sequence complementary to the sequence of the complementary strand capture portion on the 5' side may be used to generate a nucleic acid having a sequence derived from a target nucleic acid. In this case, the complementary strand capture portion of the nucleic acid linker present at the 5' end of a complementary strand (cDNA) generated by reverse transcribing mRNA can capture a sequence complementary to the sequence of the complementary strand capture portion derived from the primer present at the 5' end of a nucleic acid having a sequence derived from another target nucleic acid, and two or more complementary strands are linked via the nucleic acid linker.
[0030] When such a nucleic acid linker is used, the sequence of the complementary strand capture part is not limited to the GGG sequence, but can be a sequence of any length by combining the four bases A, T, C, and G. If the sequence of the complementary strand capture part is made longer, when two or more generated complementary strands are linked via the nucleic acid linker to form a circular nucleic acid, it is possible to form a circular nucleic acid while reducing the loss of the generated complementary strand. This is expected to increase the probability of detecting even a complementary strand derived from the sequence of an mRNA that is expressed at a low frequency, and improve analytical accuracy.
[0031] The complementary strand capture unit 12 that captures the CCC sequence includes, for example, a GGG sequence. The GGG sequence may be DNA or RNA. That is, the GGG sequence may be GGG or rGrGrG (r: ribonucleotide, G: guanine). The complementary strand capture unit 12 may also be single-stranded DNA or RNA. This facilitates binding to a complementary strand.
[0032] Furthermore, the complementary strand capture part may contain a self-binding inhibition sequence Hn or Nn in addition to the GGG sequence. Here, H is a base other than G, i.e., A, T, or C. N is A, T, G, or C. n is the number of H or N and may be, for example, an integer of 1 or greater. n may be, for example, an integer from 1 to 8. It is said that there are approximately 20,000 types of mRNA, and such a numerical range can cover such mRNAs. In some embodiments, n may be, for example, 1, 2, 3, 4, or 5, or even 1 or 2. When n is 2 or greater, each H or N constituting the self-binding inhibition sequence may be selected independently and randomly. The complementary strand capture part may have, for example, the base sequence GGGH, GGGN, GGGHN, GGGNH, GGGHH, or GGGNN. The complementary strand capture part may be DNA or RNA.
[0033] The double-stranded portion 13 is a portion that connects the target nucleic acid capture portion 11 and the complementary strand capture portion 12, and may be formed from double-stranded DNA, double-stranded RNA, or a hybrid of DNA and RNA. Preferably, the double-stranded portion 13 is DNA. This prevents degradation in the RNA digestion process described below and facilitates the formation of circular nucleic acids. It can also be easily used as a primer in nucleic acid amplification.
[0034] A target nucleic acid capture part 11 is linked to the 3' end of one strand of the double-stranded part 13, and a complementary strand capture part 12 is linked to the 3' end of the other strand. With this structure, the generated complementary strands can be linked to form a circular nucleic acid, as described below.
[0035] The double-stranded portion 13 may include a sequence composed of a random combination of the four bases A, T, C, and G. Preferably, the random sequence is composed of a sequence group with error correction functionality. Examples of sequences with error correction functionality include Sequence-Levenshtein code and filled / truncated right end edit (FREE) barcodes. Methods for generating these sequence groups and error correction mechanisms using these sequence groups are described in Buschmann and Bystrykh BMC Bioinformatics 2013, 14:272 and Proc Natl Acad Sci US A. 2018 Jul 3; 115(27): E6217-E6226. Those skilled in the art can refer to these documents to appropriately generate and use sequences with error correction functionality. In addition to the above two types of sequences, other error-correcting sequences, such as Levenshtein codes, Hamming codes, or Reed-Solomon codes, are known in the art, and any of these may be used in the present disclosure. Sequences with error correction capabilities may also be referred to as indel-correcting DNA barcodes. Software that can be used to generate these error-correcting sequences and perform error correction using these sequences is also known to those skilled in the art, and those skilled in the art can use such software to prepare and use sequences with error correction capabilities. This allows for the extraction of sequences that can identify the original sequence even if a reading error (insertion, deletion, or substitution) of several bases occurs during sequencing.
[0036] The DNA fragments constituting the obtained circular nucleic acids are randomly linked, and may be different from each other in the regions where the nucleic acid linkages are arranged. That is, the multiple nucleic acid linkages possessed by the obtained multiple types of circular nucleic acids all have the same base sequence, but the sequences of the regions derived from the DNA fragments arranged between the multiple nucleic acid linkages are different. Because the DNA fragments can be arranged in a unique order for each of the obtained multiple types of circular nucleic acids, this order of the DNA fragments can be used as an identifier such as a unique molecular identifier (UMI).
[0037] Preferably, of the two base sequence strands constituting the double-stranded portion 13, the strand connected to the target nucleic acid capture portion 11 (particularly the strand having polyT) is phosphorylated at its 5' end. This allows the ligation process described below to be carried out more reliably. Preferably, of the two base sequence strands constituting the double-stranded portion 13, the strand not connected to the target nucleic acid capture portion 11 (particularly the strand not having polyT) has the complementary strand capture portion connected to its 3' end.
[0038] The double-stranded portion may preferably have a non-natural base sequence, and in particular, the random sequence may be composed of a non-natural base sequence. A non-natural base sequence refers to a base sequence that does not exist in nature. Those skilled in the art can appropriately design such base sequences. By using such a non-natural base sequence, for example, unnecessary double-strand formation and sequence detection errors can be suppressed.
[0039] The double-stranded portion 13 may further have a priming sequence. The priming sequence may be, for example, a base sequence that functions as a primer in the nucleic acid amplification process described below, and the sequence can be appropriately selected by those skilled in the art depending on, for example, the type of nucleic acid amplification process and / or the type of enzyme used in the nucleic acid amplification process.
[0040] As described above, in the process for generating circular nucleic acids according to the present disclosure, multiple regions in which nucleic acid linkers are arranged are used, and the priming sequences of the multiple nucleic acid linkers arranged in one region all have the same base sequence, and further, the priming sequences may be the same between the regions, thereby allowing simultaneous amplification reactions to occur from multiple circular nucleic acids having nucleic acid linkers with the same priming sequence in a single amplification treatment.
[0041] When the nucleic acid linker includes a sequence addition section 14 as in the example shown in Fig. 4C, the sequence addition section 14 is disposed between the target nucleic acid capture section 11 and the complementary strand capture section 12. The sequence addition section 14 may be formed from DNA, RNA, or a hybrid of DNA and RNA.
[0042] The sequence addition unit 14 may include a sequence composed of a random combination of four bases, A, T, C, and G. For example, the random sequence may be added as a cell identifier sequence (cell barcode) that identifies the origin of the nucleic acid. Alternatively, the sequence addition unit 14 may be added with a group of sequences having a function that can be added to the double-stranded portion.
[0043] The nucleic acid linker 10 (particularly the double-stranded portion 13 and the sequence addition portion 14) may contain a UMI, but does not have to contain a UMI. By not containing a UMI, the base sequence of the nucleic acid linker can be shortened.
[0044] (1-1) Specific Example of a Process for Generating Circular Nucleic Acids Using a Nucleic Acid Linker Below, with reference to Figures 5A and 5B, a process for generating circular nucleic acids using the nucleic acid linker is described, and nucleic acid amplification using the circular nucleic acid is also described. More specifically, in the circular nucleic acid generation process shown in the figures, two types of mRNA are reverse transcribed to generate two types of cDNA complementary to each mRNA, and these two types of cDNA are then linked to generate a circular nucleic acid. The steps performed in the circular nucleic acid generation process are as follows. Note that in this example, two different types of mRNA are linked, but in the present disclosure, two or more of the same mRNA may also be linked.
[0045] As shown in the upper part of Figure 5A, assume that there are two types of mRNA (mRNA1 and mRNA2), where mRNA1 and mRNA2 are target nucleic acids.
[0046] As shown in Figure 5B, mRNA1 contains a coding region CDS1 and further contains a polyA sequence polyA1 at the 3' end. Similarly, mRNA2 contains a coding region CDS2 and further contains a polyA sequence polyA2 at the 3' end.
[0047] In step S11 shown in Figure 5A, the 3' end of mRNA1 is captured by the nucleic acid linker 10. The 3' end of mRNA1 has a polyA tail, and the polyT sequence 11 of the nucleic acid linker captures, particularly complementarily binds to, the polyA tail, thereby capturing the 3' end of the mRNA by the nucleic acid linker (particularly the target nucleic acid capturer). The 3' end of mRNA2 is also captured by the nucleic acid linker (particularly the target nucleic acid capturer), similar to mRNA1.
[0048] In step S12 of the figure, cDNA is synthesized from mRNA1. Synthesis of the cDNA may be performed, for example, by reverse transcriptase. The cDNA synthesis is performed using the nucleic acid linker as a primer. As a result, a hybrid H1 between mRNA1 and cDNA1 is formed, as shown in the figure. Similarly, for mRNA2, cDNA2 is synthesized by reverse transcriptase, forming a hybrid H2 between mRNA2 and cDNA2. When cDNA synthesis is performed by reverse transcriptase in step S2, a CCC sequence is formed at the 3' end of the synthesized cDNA, as shown in the figure. As described below, this CCC sequence is used in the subsequent formation of circular nucleic acids. Thus, in the method according to the present disclosure, a complementary nucleic acid generation step may be performed in which complementary strands of one or more target nucleic acids are generated with a nucleic acid linker bound to one end of each of the one or more target nucleic acids. In the complementary strand generation step, complementary strands of the target nucleic acids may be generated using the nucleic acid linker as a primer. In the complementary strand generating step, a double strand may be formed between each target nucleic acid and its complementary strand, and as described above, for example, a hybrid between mRNA and cDNA may be formed.
[0049] In step S13 of the figure, hybrid H1 and hybrid H2 generated in step S2 are ligated by complementary binding between the rGrGrG sequence of hybrid H1 and the CCC sequence of hybrid H2.
[0050] In step S14 of the figure, a single-stranded circular nucleic acid is formed using hybrid H1 and hybrid H2. Specifically, the following steps are carried out to form the single-stranded circular nucleic acid.
[0051] First, hybrid H1 has a CCC sequence at the end opposite the end where the rGrGrG sequence used for ligation in step S13 is present, and this CCC sequence is a single-stranded portion. Hybrid H2 has an rGrGrG sequence at the end opposite the end where the CCC sequence used for ligation in step S13 is present, and this CCC sequence is also a single-stranded portion. Therefore, the CCC sequence of hybrid H1 and the rGrGrG sequence of hybrid H2 complementarily bind to form a double-stranded circular nucleic acid. Thus, in the circular nucleic acid generation step included in the method according to the present disclosure, the double strands may be linked via the nucleic acid linker. Furthermore, in the circular nucleic acid generation step, the 5'-end of one complementary strand and the 3'-end of another complementary strand may be linked, and this linkage may be performed in a state where the complementary strand capture portion of the nucleic acid linker bound to one complementary strand is bound to the 3'-terminal region of the other complementary strand.
[0052] In the double-stranded circular nucleic acid, a nick exists between the 5' end of the strand linked to cDNA1 in the double-stranded nucleic acid linker that was bound to mRNA1 and the 3' end of cDNA2 generated by reverse transcription of mRNA2. Similarly, a nick exists between the 5' end of the strand linked to cDNA2 in the double-stranded nucleic acid linker that was bound to mRNA2 and the 3' end of cDNA1 generated by reverse transcription of mRNA1. Therefore, in step S14, ligation is performed to eliminate these nicks while the double-stranded circular nucleic acid is formed. This leaves cDNA1 and cDNA2 bound together, and the state of the circular nucleic acid is maintained even after mRNA degradation, as described below.
[0053] In step S14, after the ligation, the mRNA is degraded. This degradation may be performed using, for example, RNase, particularly RNase H. As a result, a single-stranded circular nucleic acid RN is formed as shown in the figure. In this way, in the circular nucleic acid generation step, a double-stranded circular nucleic acid is formed, and the target nucleic acid (e.g., mRNA) is removed from the double-stranded circular nucleic acid to obtain a single-stranded circular nucleic acid to which a complementary strand is ligated.
[0054] In step S15, a primer is added to the single-stranded circular nucleic acid RN. In this figure, a primer is used that is configured to bind to the portion of the nucleic acid linker that constituted the double-stranded portion and the CCC sequence. In other words, the primer is configured to cover the ligation point.
[0055] In step S16, rolling circle amplification (RCA) is performed using a DNA polymerase starting from the primer, thereby synthesizing a single-stranded linear nucleic acid complementary to the single-stranded circular nucleic acid, for example.
[0056] The base sequence of the single-stranded linear nucleic acid thus formed is then decoded. This decoding may be performed using techniques known in the art. Examples of such sequence analysis techniques include Long Read Sequence technology. For example, the single-stranded cDNA is sequenced using nanopore sequencing (Oxford Nanopore Technologies). Alternatively, the base sequence may be sequenced using HiFi sequencing (Pacific Biosciences). In this case, as shown in FIG. 5C , a double-stranded structure is generated after a step of adding a polyA or polyC tail to the single-stranded cDNA extended by RCA. An SMRTbell® adapter is then added to the double-stranded nucleic acid to produce a circular nucleic acid, which is then sequenced. Here, since the sequence of the nucleic acid linker is known, after the sequencing, the sequenced sequence can be separated by the known nucleic acid linker, thereby identifying the original mRNA sequence and counting the number of original mRNAs. By decoding the repeated mRNA sequence, even in the case of multi-RCA, the molecules resulting from replication can be considered identical and counted. In other words, whether the number is the original mRNA or the number resulting from replication can be distinguished without using UMI. Furthermore, when rGrGrGNN or GGGNN is used, the location of the random NN can also be used to distinguish. Furthermore, by performing long read sequencing of the repeated mRNA complementary sequence, even if a read error occurs during sequencing, the error can be corrected. In other words, linking two or more mRNAs as described above enables error correction, which can improve decoding accuracy.
[0057] When rGrGrG is used as the complementary strand capture portion of the nucleic acid linker, this rGrGrG is degraded by RNase H treatment. Furthermore, when the complementary strand (i.e., the strand connected to the complementary strand capture portion) of the double-stranded portion in the nucleic acid linker is synthesized with ribonucleotides, the complementary strand is also degraded by RNase H treatment. Therefore, in step S15 (i.e., before the RCA in step S16), a primer is added. This primer may be, for example, a primer having the same sequence as the complementary strand of the double-stranded portion. Furthermore, when GGG is used as the complementary strand capture portion and the complementary strand of the double-stranded portion in the nucleic acid linker is synthesized with nucleotides, they are not degraded by RNase H treatment. In this case, the GGG and the complementary strand may be used as primers.
[0058] In the example shown in Figure 5A, the circular nucleic acid is formed using two mRNAs, but it may also be formed using three or more mRNAs. That is, circular nucleic acids containing cDNAs complementary to each of the three or more mRNAs may be similarly formed. The circular nucleic acid may then be subjected to nucleic acid amplification treatment such as RCA as described above.
[0059] By generating a circular nucleic acid as shown in FIG. 5A above, two or more cDNAs are linked to the circular nucleic acid. For example, a cDNA corresponding to a low-expressed mRNA and a cDNA corresponding to a high-expressed mRNA can be linked. This improves the detection efficiency of low-expressed mRNA. Furthermore, in the case of PCR, nucleic acids are amplified exponentially, so mRNAs present in large numbers before amplification are amplified to a much greater extent than mRNAs present in small numbers before amplification. As a result, mRNAs present in small numbers before amplification are difficult to detect. In other words, PCR introduces amplification bias (also known as amplification bias). On the other hand, nucleic acid amplification by RCA is not exponential, so this bias is reduced. That is, in one embodiment, nucleic acid amplification using the circular nucleic acid may be performed by RCA.
[0060] In some embodiments according to the present disclosure, PCR may be used. Even when PCR is used, the bias can be reduced by performing PCR on the circular nucleic acid. This is explained below. Typically, after cDNA synthesis, a base sequence derived from the nucleic acid fragment involved in the linkage, such as a nucleic acid linker, is added to the nucleic acid amplified by PCR. In PCR, sequences are generated exponentially, and unique molecular identifiers (UMIs) are typically assigned during reverse transcription to identify the origin of the generated sequences. Here, the number of UMI types must be equal to or greater than the number of molecular types to be distinguished (e.g., the number of molecular types originally present in each cell). Therefore, UMIs typically consist of six or more bases (nucleotides) (e.g., 6 to 10, or even 10 or more). Although the UMI sequences are decoded during sequencing, these sequences are not contained in the target nucleic acid and do not necessarily need to be sequenced. Because the number of reads that can be decoded during sequencing is limited, it is preferable to be able to eliminate them in order to increase the number of reads of sequences derived from the target nucleic acid.
[0061] Furthermore, in the PCR method, nucleic acids are exponentially amplified, so molecules with high expression levels tend to increase, resulting in a bias in which these molecules are more likely to be detected. For example, as shown in Figure 6A, when PCR is performed on a sample in which there is one target nucleic acid M1 and the number of target nucleic acids M2 is m (seven in the figure), the reverse transcription product of target nucleic acid M1 is amplified to 2 n whereas the target nucleic acid M2 is amplified to (2 n ) × m.
[0062] When PCR is performed on a nucleic acid to which two or more target nucleic acids are linked according to the present disclosure, for example, a molecule with a low expression level is linked to a molecule with a high expression level, and PCR is performed on the linked nucleic acid. This is expected to improve the detection efficiency of low-expression molecules. For example, by performing the linking process according to the present disclosure on the sample described with reference to Figure 6A, it is assumed that, as shown in Figure 6B, a linkage product is generated in which one reverse transcription product of target nucleic acid M1 is linked to three reverse transcription products of target nucleic acid M2, and a linkage product is generated in which four reverse transcription products of target nucleic acid M2 are linked. In this case, each linkage product is 2 n 6A, the difference between the number of amplification products of the target nucleic acid M1 and the number of amplification products of the target nucleic acid M2 is smaller than in the case of A in FIG.
[0063] In this way, even when the PCR method is used instead of the RCA method in the nucleic acid analysis method according to the present technology, the amplification bias is reduced. As described above, in the nucleic acid analysis method according to the present technology, a nucleic acid amplification reaction using the circular nucleic acid may be performed, and the nucleic acid amplification reaction may be, for example, RCA or PCR.
[0064] (1-2) Modified Examples of the Process for Producing Circular Nucleic Acid Using a Nucleic Acid Linker In the specific example of (1-1), a process for producing a circular nucleic acid using a nucleic acid linker having a double-stranded portion was described, but a circular nucleic acid can also be produced by randomly linking multiple types of DNA fragments using a nucleic acid linker that does not have a double-stranded portion. Below, a process for producing a circular nucleic acid using a nucleic acid linker that does not have a double-stranded portion will be described.
[0065] In this case, multiple types of DNA fragments can be randomly linked by combining a nucleic acid linker that does not have a double-stranded portion with a primer having a sequence complementary to the complementary strand capture portion of the nucleic acid linker. Note that the process for generating circular nucleic acids described here is merely an example and is not limited to the form described in this specification. Multiple types of DNA fragments can be randomly linked by arbitrarily designing the nucleic acid linker and primer and combining them in any manner.
[0066] 7A is an image diagram of a nucleic acid linker and the like used in a first modified example of the process for generating a circular nucleic acid. In FIG. 7A, <7A-α> is a nucleic acid linker used in the first modified example of the process for generating a circular nucleic acid. <7A-β> is a primer used when generating a nucleic acid having a sequence derived from a target nucleic acid in the process for generating a circular nucleic acid in this modified example. <7A-γ> is an RNA oligomer for protecting the sequence of the complementary strand capture portion of the nucleic acid linker when generating a nucleic acid having a sequence derived from a target nucleic acid.
[0067] The nucleic acid linker of <7A-α> has a target nucleic acid capture portion (poly T in the figure) capable of capturing the poly A sequence present at the 3' end of mRNA, and a complementary strand capture portion (concatenator in the figure). Although not shown in the example shown in FIG. 7A, the nucleic acid linker may have a sequence addition portion between the target nucleic acid capture portion and the complementary strand capture portion. The sequence of the complementary strand capture portion provided in the nucleic acid linker used in this modification may be any sequence. Furthermore, when the nucleic acid linker has a sequence addition portion, any sequence can be added to the sequence addition portion depending on the function to be imparted to the sequence addition portion.
[0068] The primer shown in <7A-β> has a poly-T sequence complementary to the poly-A sequence added to the cDNA by terminal transferase in step S213 described below, and a sequence complementary to the sequence of the complementary strand capture part provided in the nucleic acid linker of <7A-α>. This allows it to serve as a starting point for generating a nucleic acid having a sequence derived from a target nucleic acid using the cDNA generated in the circular nucleic acid generation process of this modified example as a template, and the sequence complementary to the sequence of the complementary strand capture part allows the generated two or more nucleic acids to be suitably linked via the nucleic acid linker.
[0069] The RNA oligomer <7A-γ> is formed from RNA. The sequence of the RNA oligomer is complementary to the sequence of the complementary strand capture portion of the nucleic acid linker <7A-α>. This allows the RNA oligomer to hybridize favorably to the 5' end of the cDNA produced by the circular nucleic acid production process of this modified example.
[0070] In a first variant of the process for generating a circular nucleic acid, the polyA tailing method is used in the complementary nucleic acid generation step, and a primer having a sequence that can be captured by a complementary strand capture portion provided in a nucleic acid linker is used to generate a nucleic acid having a sequence derived from the target nucleic acid using the complementary strand of the generated target nucleic acid as a template. This will be explained below with reference to Figure 7B.
[0071] 7B, in step S211, a target nucleic acid capturer included in the nucleic acid linker shown in <7A-α> captures the 3' end of mRNA. Step S211 is the same as in the specific example of the (1-1) circular nucleic acid production process described above, except for the nucleic acid linker used. Conditions similar to those applicable to the circular nucleic acid production process can also be suitably applied to this variant.
[0072] Regarding step S212 in FIG. 7B in which cDNA is generated using the mRNA sequence as a template, the same conditions that can be applied to the specific example of the (1-1) circular nucleic acid generation process described above can also be suitably applied to this modified example.
[0073] In step S213 of Figure 7B, the RNA-DNA complex produced in step S212 is treated with RNase H to degrade the mRNA that serves as a template for the cDNA. The complex is then treated with terminal transferase to add a poly(A) sequence to the 3' end of the cDNA. As a result, even if reverse transcription by reverse transcriptase does not extend all the way to the 5' end of the mRNA and a CCC sequence is not added, the poly(T) sequence contained in the primer shown in <2D-β> allows the primer to hybridize favorably to the cDNA in the next step S214.
[0074] 7B, the <7A-β> primer is hybridized to the 3'-end of the cDNA generated in step S213, and the <7A-γ> RNA oligomer is hybridized to the 5'-end. The order of hybridization is not limited, and for example, the RNA oligomer may be hybridized first, or the concentration of the RNA oligomer may be higher than that of the primer and hybridized simultaneously.
[0075] 7B, a DNA polymerase without strand displacement activity is used to generate a nucleic acid having a sequence derived from the target nucleic acid, starting from the primer <7A-β> and using the cDNA as a template. Note that, since the RNA oligomer <7A-γ> is hybridized to the 5'-end of the cDNA, the nucleic acid generation reaction by the DNA polymerase without strand displacement activity terminates at the position where the RNA oligomer is hybridized.
[0076] In step S216 of FIG. 7B, the RNA-DNA complex produced in step S215 is treated with RNase H to degrade the RNA oligomer and obtain a DNA-DNA complex containing a complementary strand of the target nucleic acid.
[0077] 7B, two or more DNA-DNA complexes containing complementary strands of the target nucleic acid are linked via the complementary strand capture portion of the nucleic acid linker and a sequence complementary to the sequence of the complementary strand capture portion derived from the primer <7A-β>. Conditions for this step that are similar to those applicable to the specific example of the process for producing circular nucleic acid (1-1) can also be suitably applied to this modified example.
[0078] As in the specific example of the process for producing circular nucleic acid (1-1) described above, the ligation reaction shown in step S217 forms a single-stranded circular nucleic acid in the same manner as shown in step S14 in Fig. 5A. Note that any conditions can be suitably applied as conditions for forming this single-stranded circular nucleic acid.
[0079] In a first variant of the process for generating a circular nucleic acid, the polyA tailing method is used in the complementary nucleic acid generation step, and a primer is used to which a sequence that is captured by a complementary strand capture portion of a nucleic acid linker is added. For example, even if reverse transcription by reverse transcriptase is not carried out to the 5' end of the mRNA and a CCC sequence is not added, a sequence complementary to the sequence of the complementary strand capture portion of the nucleic acid linker can be added. Therefore, when two or more generated complementary strands are linked via the nucleic acid linker to form a circular nucleic acid, the generated complementary strands can be reduced in their absence, and the circular nucleic acid can be formed. This increases the probability of detecting complementary strands derived from sequences of mRNAs that are expressed at low frequencies, and is expected to improve analytical accuracy.
[0080] Since the combination of DNA-DNA complexes to be ligated in the ligation reaction shown in step S217 cannot be selected, the single-stranded circular nucleic acid formed in step S217 is a single-stranded circular nucleic acid in which sequences derived from cDNAs generated using the mRNA sequence as a template are ligated randomly. When there are many types of cDNAs to be ligated and the order of the ligated cDNAs does not substantially overlap, the single-stranded circular nucleic acid can be identified by the order of the ligated cDNAs.
[0081] Thereafter, although not shown in FIG. 7B , a primer such as <7A-β> is added to the formed single-stranded circular nucleic acid, as in the specific example of the process for generating (1-1) circular nucleic acid described above. Starting from the primer, a single-stranded linear nucleic acid complementary to the single-stranded circular nucleic acid is synthesized by RCA using a DNA polymerase. While <7A-β> is preferably used as the primer used in this process, the primer used is not limited thereto. Furthermore, the same conditions as those applicable to the specific example of the process for generating (1-1) circular nucleic acid described above can also be suitably applied to this variant. The base sequence of the resulting single-stranded linear nucleic acid can also be deciphered, for example, by the same method as that applicable to the process for generating (1-1) circular nucleic acid described above.
[0082] <Repeat unit identification step S102> In the nucleic acid analysis method of the present technology, the repeat unit identification step S102 identifies the number of bases in each circular nucleic acid from the information on the base sequence decoded in the aforementioned base sequence decoding step S101, thereby identifying the multiple types of circular nucleic acids formed.
[0083] In the base sequence decoding step S101, the base sequence of the circular nucleic acid is decoded while rotating around its circular structure, resulting in repeated occurrence of identical base sequences. The interval until the same base sequence appears due to this repetition corresponds to the number of bases in the circular nucleic acid. In other words, by determining the interval until the same base sequence appears, the number of bases in each circular nucleic acid can be determined. The method will be specifically described below.
[0084] (2-1) Identification of Repeat Units Using Autocorrelation Sequences Figure 8 shows the flow of identifying repeat units using autocorrelation sequences from the obtained base sequence information of multiple types of circular nucleic acids. First, the obtained base sequence information is converted into numerical sequence information (S111 in Figure 8). Specifically, an individual number is assigned to each base in A, C, G, etc. For example, A is assigned to 0, T to 1, C to 2, G to 3, etc. Note that this assignment is just an example, and any number can be assigned to each base. In this way, the base sequence information can be converted into numerical sequence information.
[0085] Next, the autocorrelation coefficient is calculated from the obtained sequence (S112 in FIG. 8). Here, the autocorrelation coefficient is a measure of how similar a shifted signal is to the original signal when the signal is shifted, and can be calculated using the following formula (1). The sequence showing the change in the autocorrelation coefficient with respect to the signal shift is defined as the autocorrelation sequence.
[0086]
[0087] The properties of autocorrelation will be explained in detail below using figures. Figures 9A to 9E are schematic diagrams illustrating the relationship between the deviation from the original signal and the calculated autocorrelation coefficients, using an example of a periodic function. In Figures 9A to 9E, the upper graphs (<9A-α> to <9E-α>) show graphs of the original periodic functions, and the lower graphs (<9A-β> to <9E-β>) show changes in the values of the autocorrelation coefficients when the deviation (lag) from the original signal is plotted on the horizontal axis.
[0088] If the range of the horizontal axis of the periodic function is finite, the greater the deviation from the original signal, the smaller the overlap with the original signal, and therefore the value of the autocorrelation coefficient will attenuate as shown in graph <9A-β>.
[0089] In <9A-α> to <9E-α>, the original signal is shown by a solid line, and the shifted signal is shown by a dashed line. In the <9A-α> state, the lag, which indicates the signal shift, is lag = 0.00. In the <9A-α> state, there is no shift from the original signal, so the autocorrelation coefficient is 1.0, as shown when lag = 0.00 in the <9A-β> graph.
[0090] Next, as shown in <9B-α> in Figure 9B, if the original periodic function is shifted to the right and lag = 0.25, the shift reduces the similarity from the original signal, and the autocorrelation coefficient becomes close to 0, as shown in the graph <9B-β>.
[0091] <9C-α> in Figure 9C shows the case where the original periodic function is shifted further to the right with lag = 0.50. When the phase of the original periodic function is changed to match the position where it becomes a cosine function when the original periodic function is converted to a sine function, the similarity to the original signal further decreases, and the autocorrelation coefficient becomes a minimum, as shown in the graph <9C-β>.
[0092] If the original periodic function is then shifted further to the right, with lag = 0.50, as shown in <9D-α> in Figure 9D, the similarity to the original signal now increases, and the autocorrelation coefficient increases to near 0, as shown in the graph <9D-β>.
[0093] Furthermore, <9E-α> in Fig. 9E shows the case where the original periodic function is shifted further to the right with lag = 1.00. When the phase of the original periodic function is changed to match the position where it becomes a sine function, the similarity to the original signal increases further, and the autocorrelation coefficient reaches a local maximum, as shown in the graph <9E-β>.
[0094] Then, as shown in graphs <9A-β> to <9E-β>, as the periodic signal related to the original periodic function is sequentially shifted in the direction of the horizontal axis (time), minimum and maximum values of the autocorrelation coefficient between the original signal and the shifted signal appear periodically. Note that, as mentioned above, since the range of the horizontal axis of the periodic function is finite, the greater the deviation from the original periodic function, the smaller the overlap with the original signal, and therefore the periodically appearing maximum and minimum values attenuate.
[0095] Based on the above, a method for identifying the repeat unit using an autocorrelation sequence calculated from a sequence obtained by converting a base sequence and identifying the number of bases in a circular nucleic acid will be specifically described (S113 and S114 in FIG. 8).
[0096] 10A is an image of a graph of autocorrelation calculated by the deviation of the repeating unit of a circular nucleic acid from the signal of the base sequence. In the example shown in FIG. 10A, a circular nucleic acid having the base sequence "GCTCATCGTAC" and 11 bases is used.
[0097] When the base sequence of the circular nucleic acid is decoded in the base sequence decoding step S101, the base sequence of the circular nucleic acid is decoded while going around its circular structure, and therefore, as shown in the base sequence in the upper part of Figure 10A, the above base sequence "GCTCATCGTAC" appears periodically and repeatedly in the obtained base sequence information.
[0098] Next, the obtained base sequence information is converted into a numerical sequence by assigning A to 0, T to 1, C to 2, and G to 3 for each base in A, C, and G. The lower part of Figure 10A is a graph showing the numerical sequence.
[0099] FIG. 10B shows the transition of the calculated autocorrelation coefficient when the graph of the sequence is shifted to the right. In the graph, each element shown in a dark color indicates the shifted base position, and each element shown in a light color indicates the original base position. The horizontal axis also indicates the order of the arranged bases. As shown in the upper part of FIG. 10B, when the graph of the original several examples is shifted by one base at a time, the similarity increases when shifted by 11 bases, and as shown in the graph in the lower part of FIG. 10B, the value of the autocorrelation coefficient reaches a maximum value. Thereafter, as the graph of the original sequence is sequentially shifted to the right, a maximum value of the autocorrelation coefficient appears every 11 bases, which is the number of bases in the target circular nucleic acid. In other words, by analyzing the period in which the maximum value of the autocorrelation coefficient appears, the number of bases in the target circular nucleic acid can be identified.
[0100] Here, since the sequence derived from the base sequence of a circular nucleic acid has no continuity between adjacent values, the discontinuity of the autocorrelation sequence is extremely large. Therefore, in the absence of sequence errors, etc., as will be described later, the maximum values appearing as repeat units are easy to identify. In particular, the maximum point of the autocorrelation coefficient corresponding to the repeat unit is easy to identify because it is the largest among the candidate maximum values under ideal conditions where the sequence error is less than 0.1%.
[0101] Figure 11 shows a specific example of determining the number of bases in a circular nucleic acid from its base sequence information. In the example of Figure 11, in the graph showing the transition of the autocorrelation coefficient shown in the middle, it can be confirmed that a maximum value appears at a cycle of every 3607 bases. In other words, the number of bases in the circular nucleic acid of Figure 11 can be determined to be 3607 bases.
[0102] Since the autocorrelation sequence takes large values at or near 0, it is preferable to exclude, as noise, local maxima that clearly occur at positions shorter than the base length of the DNA fragment from candidates when analyzing the period of local maxima that occur in the autocorrelation sequence of the base sequence. Furthermore, among the local maxima candidates, those whose autocorrelation coefficient values are clearly smaller than those of nearby local maxima and therefore cannot be candidates for repeat units are also preferably excluded as noise by setting a threshold value.
[0103] (2-2) Weighting of Function As described above, this technology preferably identifies the number of bases in a target circular nucleic acid by analyzing the period in which the maximum value of the autocorrelation coefficient of the base sequence information obtained by decoding the base sequence of the circular nucleic acid appears. However, in the process of decoding the base sequence of a circular nucleic acid, sequence errors (insertion, deletion, or misreading of bases) may occur at a certain frequency. The higher the rate of sequence errors, the more difficult it becomes to distinguish between the maximum value of the autocorrelation coefficient of the base sequence information reflecting the number of bases in the target circular nucleic acid and the amplitude of the autocorrelation coefficient in other regions. In light of this, it is preferable to weight the autocorrelation sequence before identifying the maximum value of the autocorrelation coefficient (S115 in FIG. 8). A detailed explanation is provided below using the figures.
[0104] 12 shows an image of calculating the weight of an original signal and weighting the original signal using the calculated weight. The original signal, such as the autocorrelation sequence in <12A>, has a large amount of noise amplitude, making it difficult to distinguish the maximum value of the autocorrelation sequence in this state. Taking this into consideration, <12B> shows the transition of the average value within a certain range centered on the value of the target autocorrelation sequence as a smoothed signal. However, the amount of information contained in the smoothed signal is reduced.
[0105] Considering that the amount of information in the above smoothed signal is small, the weighted signal shown in <12C> can be calculated by taking the product of each element of the sequence of the original signal using the sequence of the smoothed signal shown in <12B> as a weight.
[0106] 13 shows an example in which the sequence of autocorrelation coefficients of base sequence information obtained by decoding the base sequence of a circular nucleic acid is smoothed and then used to weight the original sequence of autocorrelation coefficients. As shown in <13A>, the sequence of autocorrelation coefficients of the base sequence information obtained by decoding the base sequence of the circular nucleic acid, which is the original signal (corr in <13A>), has many amplitudes that become noise, making it difficult to distinguish the maximum value of the autocorrelation sequence in this state. In light of this, <13A> shows a smoothed sequence that is the transition of the average value of the autocorrelation sequence, with the width of a certain range centered on the value of the target autocorrelation sequence set to 50, 100, and 200.
[0107] By multiplying each element of the sequence of the original signal by the sequence of the smoothed signal as a weight, a weighted sequence is calculated as shown in <13B>. From this graph, it can be seen that by performing the above process, discontinuous amplitudes at the micro level are suppressed, and a smooth curve is obtained in which only macro peaks remain. Furthermore, in this weighted sequence of autocorrelation coefficients, peaks other than macro peaks are suppressed, which can reduce the probability of determining points outside the cycle in which the autocorrelation coefficient maximums actually occur as maximum values.
[0108] (2-3) Identification of repeat units using spectral analysis Figure 14 shows the flow of identifying repeat units using spectral analysis from the obtained base sequence information of multiple types of circular nucleic acids. The obtained base sequence information is converted into sequence information using the same method as in the case of identifying repeat units using the autocorrelation sequence described above (S121 in Figure 14).
[0109] Next, a spectral sequence is calculated from the obtained sequence (S122 in FIG. 14). Specifically, the sequence derived from the base sequence information of the circular nucleic acid is Fourier transformed to calculate the frequency component distribution, which is then converted into a period component distribution. As mentioned above, in the base sequence decoding step S101, the base sequence of the circular nucleic acid is decoded while rotating around its circular structure, resulting in repeated occurrences of identical base sequences. Therefore, the maximum value of the period component distribution of the obtained spectral sequence can be estimated to correspond to the number of bases in the circular nucleic acid. In other words, the number of bases in the circular nucleic acid can be determined by calculating the maximum value of the period component distribution of the spectral sequence obtained by spectral analysis of the decoded base sequence (S123 and S124 in FIG. 14). A detailed explanation is provided below using the figures.
[0110] 15 is an illustration of a case where a repeating unit is identified based on the component distribution of the period of a spectral sequence obtained by spectral analysis, using an example of a periodic function. <15A> shows an example of a periodic function as the original signal. <15B> shows the frequency component distribution calculated by Fourier transforming the periodic function. <15C> shows the component distribution of the period of the spectral sequence converted based on the frequency component distribution. In this case, as shown in <15C>, the maximum value of 1 in the component distribution of the period of the spectral sequence can be estimated as the repeating unit.
[0111] <15D> shows the transition of the autocorrelation coefficient of the periodic function, which is the original signal. From <15D>, it can be confirmed that a maximum value appears at period 1, which is calculated as the maximum value of the component distribution of the period of the spectral sequence, and therefore it can be estimated as a repeating unit.
[0112] More specifically, an example of identifying the number of bases in a circular nucleic acid based on the component distribution of the period of a spectral sequence obtained by spectral analysis from the base sequence information of the circular nucleic acid will be described with reference to FIG.
[0113] <16A> is a graph showing a sequence derived from the base sequence information of a circular nucleic acid. <16B> is a graph showing the sequence subjected to Fourier transform to calculate the frequency component distribution. <16C> is a graph showing the periodic component distribution of a spectral sequence converted based on the frequency component distribution.
[0114] In spectral analysis, local maxima other than those of repeat units, such as high-frequency components, are likely to be detected. However, since the high-frequency components correspond to periods that are much shorter than the base length of the DNA fragments contained in the circular nucleic acid and are not candidates for repeat units, they may be removed as noise using a low-pass filter. This allows the number of bases in the circular nucleic acid to be identified from the local maxima of the component distribution of the period of the spectral sequence, as shown in the red frame in <16C>.
[0115] In this case, by confirming that the maximum value of the component distribution of the period of the spectral sequence matches with the maximum value occurring in the autocorrelation sequence of the base sequence information of the circular nucleic acid, the accuracy in identifying the number of bases in the circular nucleic acid can be improved. Here, "match" in this case means that the number of bases in the maximum value of the component distribution of the period of the spectral sequence and the number of bases in the maximum value occurring in the autocorrelation sequence of the base sequence information of the circular nucleic acid may not strictly match, so an error of up to a certain number of bases may be allowed, and the number of bases identified in either one may be referenced and identified as the number of bases in the circular nucleic acid.
[0116] Even when the number of bases in a circular nucleic acid is determined by identifying the repeat unit through spectral analysis, sequence errors that occur with a certain frequency during the process of decoding the base sequence of the circular nucleic acid may make it difficult to distinguish the maximum value of the component distribution of the period of the spectral sequence. In light of this, it is preferable to weight the spectral sequence before identifying the maximum value of the component distribution of the period of the spectral sequence (S125 in FIG. 14). As a weighting method, the method described above for weighting the (2-2) function can be suitably used.
[0117] (2-4) Identification of Repeat Unit Using Autocorrelation Sequence and Spectral Sequence In light of the above, it is preferable that the identification of the number of bases of the circular nucleic acid in the repeat unit identification step S102 of the present technology is performed by combining identification of the repeat unit using an autocorrelation sequence and identification of the repeat unit using spectral analysis. Specifically, as shown in the example of the flow chart in Figure 17, it is preferable that the number of bases of the circular nucleic acid is identified by a maximum value occurring in the autocorrelation sequence of the decoded base sequence, which coincides with a maximum value of the component distribution of the period of the spectral sequence obtained by spectral analysis of the decoded base sequence.
[0118] In this case, taking into consideration sequence errors that occur with a certain frequency in the process of decoding the base sequence of a circular nucleic acid, it is preferable to perform the above-mentioned weighting in identifying the repeat unit using the autocorrelation sequence and in identifying the repeat unit using spectral analysis.
[0119] (2-5) Effect of the Number of RCAs on Identification of Repeat Units In the base sequence decoding step S101 of the present technology, the lower the rate of sequence errors, the more preferable it is. However, in the base sequence decoding step S101 of the present technology, since the base sequence is decoded multiple times while circling the circular structure of the circular nucleic acid using a circular nucleic acid, the accuracy of identifying the number of bases in the circular nucleic acid is high even if a sequence error occurs.
[0120] Figure 18 shows data on the effect of the number of rotations (RCA number) made around the circular structure of the circular nucleic acid to decode the base sequence on the sequencing error rate in the base sequence decoding step S101. The horizontal axis of the graph in Figure 18 represents the sequencing error rate, and the vertical axis represents the error rate in identifying the repeat unit.
[0121] The data shown in Figure 18 shows the measurement of the error rate for identifying repeat units at RCA times set between 2 and 5 times using a specific circular nucleic acid, with sequencing error rates of 0%, 0.1%, 1%, 10%, 20%, 30%, and 40%. The RCA times are set as 1.0 for one full revolution around the circular structure. For example, the RCA times for one and a half revolutions around the circular structure is 1.5. The measured values are the average of 10 trials, each with 200 reads per trial.
[0122] It can be seen from the data in Figure 18 that the greater the number of RCAs, the higher the accuracy of identifying the number of bases in a circular nucleic acid. For example, from the data in Figure 18, if the sequence error rate is 1% or less, setting the number of RCAs to 2.0 or more can achieve an accuracy of 99% or more in identifying the number of bases in a circular nucleic acid.
[0123] <DNA Fragment Identification Step S103> In the nucleic acid analysis method of the present technology, the DNA fragment identification step identifies regions of DNA fragments contained in each of the multiple types of circular nucleic acids identified in the repeat unit identification step S102.
[0124] In the nucleic acid analysis method of the present technology, multiple types of DNA fragments constituting a circular nucleic acid are randomly linked by linking moieties such as nucleic acid linkers, etc. Therefore, by identifying the sequence of the linking moieties that link the DNA fragments in the circular nucleic acid, it is possible to identify the regions of the DNA fragments in the circular nucleic acid in the base sequence information.
[0125] 19 is a conceptual diagram showing how the base sequence information of the circular nucleic acid identified in the repeat unit identification step S102 is divided into DNA fragment units by identifying the sequence of the linking portion. In this example, the base sequence information is divided into DNA fragment units of 780 bases, 804 bases, and 2023 bases.
[0126] <Duplicate elimination step S104> The nucleic acid analysis method of the present technology may include, after the DNA fragment identification step, a duplicate elimination step of comparing the arrangement orders of DNA fragments contained in the multiple types of circular nucleic acids to eliminate circular nucleic acids having the same order.
[0127] In this technology, multiple types of DNA fragments contained in the target of analysis are randomly linked. Therefore, the number of DNA fragments contained in the resulting circular nucleic acid is not constant. Furthermore, the number of bases in each DNA fragment is different, and the number of bases contained in the circular nucleic acid is not constant, so the total number of bases in the circular nucleic acid is also not constant. Furthermore, since the DNA fragments are randomly linked to form the circular nucleic acid, each circular nucleic acid has a unique DNA fragment arrangement pattern. Since the DNA fragment arrangement pattern can be used as an identifier for the circular nucleic acid, it is not necessary to assign a molecular barcode to the circular nucleic acid used in this technology.
[0128] Considering the above characteristics, as shown in Figure 20, the base sequence information of the circular nucleic acid identified in the repeat unit identification step S102 can be classified by the number of DNA fragments and the total number of bases in the circular nucleic acid. Furthermore, since each circular nucleic acid has a unique DNA fragment arrangement pattern, the base sequence information of the circular nucleic acid can be classified by the DNA fragment arrangement pattern. Since analyzing all of the obtained base sequence information of the circular nucleic acid at once would require an enormous amount of calculation, performing a classification process in advance narrows down the objects to be compared, reduces the calculation load, and enables efficient identification of circular nucleic acids.
[0129] Then, by comparing the arrangement order of the DNA fragments contained in the circular nucleic acid, the base sequence information of the duplicated circular nucleic acid can be eliminated from the base sequence information. As a result, duplicated DNA fragments derived from the nucleic acid contained in the target can be eliminated, so that, for example, if the target nucleic acid is mRNA, the expression level of each mRNA can be suitably determined. Using Figure 21, a method for eliminating duplicated circular nucleic acids by comparing the arrangement patterns of DNA fragments in the circular nucleic acid will be specifically described.
[0130] <21A> shows a population of base sequence information of circular nucleic acids classified using the procedure of FIG. 20. At this point, it can be confirmed that the population has been classified as base sequence information of two types of circular nucleic acids. Next, as shown in <21B>, the base sequence information of the classified circular nucleic acids is aligned with the same DNA fragment as the starting point, and the DNA fragment arrangement patterns in the base sequence information of the circular nucleic acids are compared. As a result of comparing the DNA fragment arrangement patterns, it can be confirmed that the base sequence information has been reclassified into three types of circular nucleic acids, as shown in <21C>. Based on the classification results, <21D> shows that duplicate base sequence information has been eliminated, and it has been determined that the original population of base sequence information of circular nucleic acids contains base sequence information of three types of circular nucleic acids.
[0131] As described above, the nucleic acid analysis method of the present technology can target all DNA fragments contained in the analysis target, and uses circular nucleic acids formed by randomly linking multiple types of DNA fragments into a ring, so it is not affected by amplification bias based on the expression frequency of nucleic acid fragments such as DNA, and can identify even nucleic acid fragments that appear relatively low in the entire multiple types of nucleic acid fragments contained in the target, and can preferably detect the expression level.Therefore, for example, it can also preferably analyze the expression level of mRNA in a single cell.
[0132] Furthermore, the nucleic acid analysis method of the present technology can be combined with a biological sample analyzer or the like to suitably analyze the mRNA expression levels of cells sorted by the biological sample analyzer. In this case, by taking advantage of the high throughput characteristics of the biological sample analyzer, which can sort single cells, it is possible to efficiently obtain the analysis target and quickly analyze the mRNA expression levels in single cells.
[0133] <Base sequence determination step S105> The nucleic acid analysis method of the present technology may include, after the DNA fragment identification step, a base sequence determination step of comparing base sequences of multiple DNA fragments identified as the same DNA fragment to determine the base sequence of the identical DNA fragment.
[0134] When the base sequence determination step is performed in the nucleic acid analysis method of the present technology, for each DNA fragment identified in the DNA fragment identification step as being located at the same position in the circular nucleic acid, the base sequences of the DNA fragments are aligned, the similarity of the bases at the same position in the base sequences is evaluated, and the base at that position is determined. In the example shown in Figure 19, the base sequences of the DNA fragments are aligned for each of the three types of divided DNA fragment units (780 bases, 804 bases, and 2023 bases).
[0135] FIG. 22 shows an image of determining bases in a base sequence. As shown in <22A>, in aligning the base sequences of DNA fragments, in the process of decoding the base sequence of a circular nucleic acid, the bases of the base sequences of the DNA fragments are aligned, taking into account sequence errors (insertion, deletion, or misreading of bases) that occur with a certain frequency. Next, as shown in <22B>, if the alignment results in bases at the same position being sufficiently similar to be considered identical, the bases are determined. The same process is performed for each base to obtain a consensus sequence. In this case, the similarity may be determined, for example, by majority vote of bases at the same position in the DNA fragments.
[0136] Figure 23 shows an image of repeating the base determination of the base sequence to determine the base sequence for each DNA fragment. <23A> shows the base sequence information of a circular nucleic acid with identified repeat units divided into DNA fragment units. In the example shown in <23A>, it can be confirmed that the target circular nucleic acid contains three types of DNA fragments. <23B> shows an image of aligning the bases of the base sequences of the DNA fragments, taking sequence errors into account, for each of the three divided DNA fragment units. In the figure, Xs are added to the positions of sequence errors contained in these DNA fragments. Because sequence errors occur randomly, it can be confirmed that sequence errors occur at different positions for each DNA fragment. The base at that position is determined by majority voting of bases at the same position, etc. <23C> shows a DNA fragment whose base sequence has been determined as a result of the alignment in <23B>.
[0137] [Nucleic Acid Analysis Program and Information Recording Medium] The present technology can also construct a nucleic acid analysis program that controls each of the above-mentioned steps.
[0138] That is, the nucleic acid analysis program constructed in the present technology executes a base sequence decoding step S101 for decoding the base sequences of multiple types of circular nucleic acids formed by randomly linking multiple types of DNA fragments, executes a repeat unit identifying step S102 for identifying the multiple types of circular nucleic acids by identifying the number of bases of the circular nucleic acid from the base sequence decoded in the base sequence decoding step S101, and executes a DNA fragment identifying step for identifying regions of DNA fragments contained in the multiple types of circular nucleic acids identified in the repeat unit identifying step S102. Here, after executing the DNA fragment identifying step, a duplicate elimination step may be executed for comparing the arrangement order of the DNA fragments contained in the multiple types of circular nucleic acids and excluding circular nucleic acids having the same order.
[0139] By executing the nucleic acid analysis program of this technology, it is possible to identify nucleic acid fragments that appear at relatively low frequencies among all the multiple types of nucleic acid fragments contained in a subject without being affected by amplification bias based on the expression frequency of nucleic acid fragments such as DNA, and it is possible to preferably detect the expression level. Note that the nucleic acid analysis program of this technology may be combined with any program as needed, depending on the purpose of the program, as long as the desired physical properties are not impaired.
[0140] The nucleic acid analysis program of the present technology can be implemented as a program on an information recording medium such as a personal computer, a control unit including a CPU, or hardware resources such as non-volatile memory (such as a USB memory), HDD, CD, etc., and can be operated by the personal computer or the control unit.
[0141] [Information Processing System] The nucleic acid analysis program of the present technology may be implemented in an information processing system. The system may be composed of a single device, multiple devices, or optional components that can be separated from the device. The nucleic acid analysis program of the present technology may be integrated into the entire system, or may be individually implemented in the devices or optional components that constitute the system.
[0142] Furthermore, by combining an information processing system that implements the nucleic acid analysis program of the present technology with any device such as a biological sample analyzer, it is possible to suitably analyze the nucleic acids contained in samples collected by the biological sample analyzer, etc. For example, when the nucleic acid analysis program of the present technology is combined with a biological sample analyzer, it is possible to suitably analyze the expression level of mRNA in single cells sorted by the biological sample analyzer.
[0143] [Nucleic Acid Analysis of Single Cells] An example of nucleic acid analysis of single cells using the present technology will be specifically described with reference to the drawings.
[0144] When performing single-cell analysis, an identifier is first assigned to each cell. The method for assigning an identifier to each cell is not particularly limited, but examples that can be used include a method in which one cell and one bar-coding bead (or gel bead) are encapsulated in a microspace such as a well or emulsion, or a method in which a bioparticle capture unit for a cell or the like, a molecule comprising a barcode sequence and a cleavable linker, is fixed to the surface of a chip or the like via the linker, and the cell or the like is captured via the bioparticle capture unit, and the linker is cleaved to release the cell from the surface, and the cell is then encapsulated in a microspace.
[0145] The barcode sequence on the surface of the barcoding bead or the chip or the like on which the bioparticle capture unit is fixed is provided with a unique identifier and a nucleic acid fragment involved in linking DNA fragments, such as a nucleic acid linking unit, etc. Thus, by mixing one cell with one barcoding bead or capturing a cell on the surface of the chip or the like on which the bioparticle capture unit is fixed, a unique identifier (cell identifier) and a nucleic acid fragment involved in linking DNA fragments are provided for each cell.
[0146] Figure 24 shows an image of a single-cell analysis in which cells are captured on the surface of a chip or the like to which the bioparticle capture unit is fixed, thereby assigning an identifier to each cell. The two cells in Figure 24 are assigned different cell identifiers. In this case, the cells are then sorted into individual microspaces using a biological sample analyzer or the like, and the nucleic acid analysis method of the present technology is performed using a sample obtained by lysing the cells, thereby enabling the mRNA expression level of each single cell to be suitably analyzed.
[0147] The method of creating a microspace for separating cells as single cells is not particularly limited, but a well method, emulsion method, or the like can be suitably used.
[0148] When an emulsion method is used as a microspace method for separating cells as single cells, emulsion particles can be generated using, for example, a microchannel. The device includes, for example, a channel through which a first liquid flows, which together form the dispersoid of the emulsion, and a channel through which a second liquid flows, which forms the dispersion medium. The first liquid may contain biological particles. The device further includes a region where these two liquids come into contact to form an emulsion. An example of a microchannel is described below with reference to FIG. 25.
[0149] The microchannel shown in FIG. 25 includes a channel 61 through which a first liquid containing bioparticles flows, and channels 62-1 and 62-2 through which a second liquid flows. The first liquid forms emulsion particles (dispersoids), and the second liquid forms the dispersion medium of the emulsion. Channel 61 and channels 62-1 and 62-2 converge, and emulsion particles are formed at this confluence. Bioparticles 63 are then isolated within the emulsion particles. For example, the size of the emulsion particles can be controlled by controlling the flow rates of these channels. To form an emulsion, the first liquid and the second liquid are immiscible with each other. For example, the first liquid may be a hydrophilic liquid and the second liquid a hydrophobic liquid, or vice versa.
[0150] 6 may also include a flow channel 64 for introducing a bioparticle-destroying substance into the emulsion particles. By configuring the microchannel so that flow channel 64 merges with flow channel 61 immediately before the junction, it is possible to prevent the bioparticles from being destroyed by the bioparticle-destroying substance before the emulsion particles are formed.
[0151] Next, an example of an apparatus for more efficiently forming emulsions containing emulsion particles containing one biological particle will be described with reference to Figures 26A and 26B. This emulsion-forming apparatus can isolate one biological particle within one emulsion particle with a very high probability, thereby reducing the number of empty emulsion particles. Furthermore, this emulsion-forming apparatus can also increase the probability of isolating one biological particle and one barcode sequence within one emulsion particle.
[0152] Fig. 26A shows an example of a microchip used to form emulsion particles in the device. Microchip 150 shown in Fig. 8A includes a main channel 155 through which bioparticles flow and a recovery channel 159 through which target particles are recovered. Microchip 150 is provided with a particle sorting section 157. An enlarged view of particle sorting section 157 is shown in Fig. 27. As shown in Fig. 27A, particle sorting section 157 includes a connection channel 170 that connects main channel 155 and recovery channel 159. A liquid supply channel 161 that can supply liquid to connection channel 170 is connected to connection channel 170. As described above, microchip 150 has a channel structure that includes main channel 155, recovery channel 159, connection channel 170, and liquid supply channel 161.
[0153] FIG. 26B is a schematic diagram illustrating the formation of emulsion particles in the microchip 150 shown in FIG. 26A and the segregation of bioparticles within the formed emulsion particles.
[0154] 26A , the microchip 150 constitutes a part of a bioparticle sorting device 200 that includes, in addition to the microchip, a light irradiation unit 191, a detection unit 192, and a control unit 193. The control unit 193 can include a signal processing unit, a determination unit, and a sorting control unit. The bioparticle sorting device 200 is used as the emulsion forming device described above.
[0155] Next, Figure 28 is a schematic diagram for explaining an example of decomposing cells sorted into an emulsion, which is a microspace, and modifying and extracting mRNA with an identifier. In this case, the method for decomposing cells is not particularly limited, and any method can be used for suitable decomposition.
[0156] By performing the nucleic acid analysis method of the present technology on samples obtained by disassembling the above cells as the analysis target, the mRNA expression level of each single cell can be suitably analyzed. Figure 29 is a schematic diagram illustrating an example of analyzing the mRNA expression level of each cell. In the example shown in Figure 29, when the mRNA expression levels of two cells are compared, it can be confirmed that the expression levels of each mRNA differ between the two cells. This makes it possible to verify, for example, the effects of various conditions on the living body for each cell.
[0157] [Application example of nucleic acid analysis of single cells] Furthermore, as an application example of nucleic acid analysis of single cells, it is also possible to combine mRNA expression analysis of single cells with any other analysis by performing processing on the sample to be analyzed in accordance with the other analysis.
[0158] For example, as shown in FIG. 30 , an antibody-nucleic acid molecule complex is used in which an antibody is bound to a nucleic acid molecule having a sequence serving as an antibody identifier sequence (antibody barcode) that identifies the origin of the antibody and a polyA sequence that is a sequence captured by a target nucleic acid capture unit or the like possessed by a nucleic acid linker. This allows for analysis of intracellular mRNA as well as analysis of molecules expressed within the cell, molecules expressed on the cell surface, or secreted molecules captured on the cell surface. This is described in detail below with reference to the diagram. Note that, as shown in the example of FIG. 30 , the nucleic acid molecule contained in the antibody-nucleic acid molecule complex may comprise a priming sequence (Amp. Primer) in addition to the antibody identifier sequence and polyA sequence. Furthermore, the complex used in this application example can be combined with a method of randomly linking multiple types of DNA fragments, which can be used in the nucleic acid analysis method of the present technology.
[0159] In the antibody-nucleic acid molecule complex shown in Figure 14C, the antibody contained in the complex is an antibody that specifically binds to a molecule expressed intracellularly, a molecule expressed on the cell surface, or a secreted molecule captured on the cell surface. Furthermore, the nucleic acid molecule contained in the complex has a unique antibody identifier sequence for each antibody. That is, if there are multiple secreted molecules on the cell surface to be analyzed, multiple types of antibody-nucleic acid molecule complexes are prepared according to the types. Furthermore, the antibody identifier sequence may have an error correction function. Although Figure 30 shows an example using an antibody-nucleic acid molecule complex, the entity that forms a complex with the nucleic acid molecule is not limited to an antibody, as long as it has the function of binding to a molecule expressed intracellularly, a molecule expressed on the cell surface, or a secreted molecule captured on the cell surface. For example, antibody fragments (e.g., Fab, scFv, VHH, Minobody), aptamers, molecularly imprinted polymers, etc., can be used. In this application example, a complex in which these are bound to a nucleic acid molecule may also be used.
[0160] The above-prepared antibody-nucleic acid molecule complex binds to cells on which the antigen molecule is expressed on the surface or on which the secreted molecule is captured on the surface through an antigen-antibody reaction. On the other hand, if the antigen molecule is not expressed intracellularly, expressed on the surface, or captured, the antibody-nucleic acid molecule complex does not bind to the cell.
[0161] By subjecting cells to an antigen-antibody reaction using the prepared antibody-nucleic acid molecule complex, which has been bound to a nucleic acid linking portion via the cell capture portion, cells can be obtained to which an antibody corresponding to a molecule expressed within the cell, a molecule expressed on the cell surface, or a secreted molecule captured on the cell surface is bound.
[0162] Next, the cells that have undergone the antigen-antibody reaction are isolated in a microspace by any method. Then, each cell is disrupted under conditions similar to those applicable to the nucleic acid analysis of single cells described above, and the nucleic acid analysis method of the present technology is performed on the resulting sample. In this application example, in the base sequence decoding step S101, when multiple types of DNA are randomly linked, an antibody identifier sequence is incorporated into the circular nucleic acid. Therefore, by performing the nucleic acid analysis method of the present technology, the presence or absence of the antibody identifier sequence contained in the circular nucleic acid can be suitably identified. This allows for analysis of mRNA expression in the cells isolated in the microspace, as well as simultaneous analysis of the expression and secretion status of molecules expressed and secreted on the cell surface.
[0163] It should be noted that, for example, by assigning antibody identifier sequences with different base sequence lengths for each antibody, it may be easier to identify the antibody identification sequence. In this case, the base sequence of the antibody identifier sequence may be adjusted to, for example, within the range of 5 to 25 bases.
[0164] As described above, by combining an information processing system that implements the nucleic acid analysis program of the present technology with a biological sample analyzer, it is possible to suitably analyze the mRNA expression level of each single cell. Below, the biological sample analyzer that can be combined with the information processing system of the present technology will be specifically described.
[0165] An example configuration of a biological sample analyzer according to the present disclosure is shown in Figure 31. The biological sample analyzer 6100 shown in Figure 31 includes a light irradiation unit 6101 that irradiates light onto a biological sample S flowing through a flow path C, a detection unit 6102 that detects light generated by irradiating the biological sample S with light, and an information processing unit 6103 that processes information related to the light detected by the detection unit. Examples of the biological sample analyzer 6100 include a flow cytometer and an imaging cytometer. The biological sample analyzer 6100 may also include a fractionation unit 6104 that separates specific biological particles P from within the biological sample. An example of a biological sample analyzer 6100 that includes the fractionation unit is a cell sorter.
[0166] (Biological Sample) The biological sample S may be a liquid sample containing biological particles. The biological particles may be, for example, cells or non-cellular biological particles. The cells may be living cells, and more specific examples include blood cells such as red blood cells and white blood cells, and reproductive cells such as sperm and fertilized eggs. The cells may be directly collected from a specimen such as whole blood, or may be cultured cells obtained after culturing. Examples of the non-cellular biological particles include extracellular vesicles, particularly exosomes and microvesicles. The biological particles may be labeled with one or more labeling substances (e.g., dyes (particularly fluorescent dyes) and fluorescent dye-labeled antibodies). Note that the biological sample analyzer of the present disclosure may also analyze particles other than biological particles, such as beads for calibration purposes.
[0167] (Flow Channel) The flow channel C is configured to allow the biological sample S to flow. In particular, the flow channel C can be configured to form a flow in which biological particles contained in the biological sample are aligned in a substantially straight line. The flow channel structure including the flow channel C may be designed to form a laminar flow. In particular, the flow channel structure is designed to form a laminar flow in which the flow of the biological sample (sample flow) is surrounded by the flow of sheath liquid. The design of the flow channel structure may be appropriately selected by those skilled in the art, and a known design may be adopted. The flow channel C may be formed in a flow channel structure such as a microchip (a chip having flow channels on the order of micrometers) or a flow cell. The width of the flow channel C may be 1 mm or less, particularly 10 μm or more and 1 mm or less. The flow channel C and the flow channel structure including it may be formed from a material such as plastic or glass.
[0168] The biological sample analyzer of the present disclosure is configured so that light from light irradiation unit 6101 is irradiated onto the biological sample flowing within flow path C, and particularly onto biological particles within the biological sample. The biological sample analyzer of the present disclosure may be configured so that the interrogation point of light on the biological sample is within the flow path structure in which flow path C is formed, or so that the interrogation point of light is outside the flow path structure. An example of the former is a configuration in which the light is irradiated onto flow path C within a microchip or flow cell. In the latter, the light may be irradiated onto biological particles after they have left the flow path structure (particularly its nozzle portion), and an example of this is a jet-in-air flow cytometer.
[0169] (Light Irradiation Unit) The light irradiation unit 6101 includes a light source unit that emits light and a light-guiding optical system that guides the light to an irradiation point. The light source unit includes one or more light sources. The type of light source is, for example, a laser light source or an LED. The wavelength of the light emitted from each light source may be any of ultraviolet light, visible light, and infrared light. The light-guiding optical system includes optical components such as a beam splitter group, a mirror group, or an optical fiber. The light-guiding optical system may also include a lens group for focusing light, such as an objective lens. There may be one or more irradiation points where the light intersects with the biological sample. The light irradiation unit 6101 may be configured to focus light irradiated from one or more different light sources onto one irradiation point.
[0170] (Detection Unit) The detection unit 6102 includes at least one photodetector that detects light generated by irradiating the bioparticles with light. The detected light is, for example, fluorescence or scattered light (e.g., one or more of forward scattered light, back scattered light, and side scattered light). Each photodetector includes one or more light-receiving elements, for example, a photodetector array. Each photodetector may include one or more PMTs (photomultiplier tubes) and / or photodiodes such as APDs and MPPCs as light-receiving elements. The photodetector includes, for example, a PMT array in which multiple PMTs are arranged in a one-dimensional direction. The detection unit 6102 may also include an imaging element such as a CCD or CMOS. The detection unit 6102 can acquire images of the bioparticles (e.g., bright-field images, dark-field images, and fluorescence images) using the imaging element.
[0171] The detection unit 6102 includes a detection optical system that allows light of a predetermined detection wavelength to reach a corresponding photodetector. The detection optical system includes a spectroscopic unit such as a prism or a diffraction grating, or a wavelength separation unit such as a dichroic mirror or an optical filter. The detection optical system is configured to, for example, disperse light generated by irradiating bioparticles with light, and detect the dispersed light using a plurality of photodetectors, the number of which is greater than the number of fluorescent dyes with which the bioparticles are labeled. A flow cytometer that includes such a detection optical system is called a spectral flow cytometer. The detection optical system is also configured to, for example, separate light corresponding to the fluorescent wavelength range of a specific fluorescent dye from the light generated by irradiating bioparticles with light, and detect the separated light using a corresponding photodetector.
[0172] The detection unit 6102 may also include a signal processing unit that converts the electrical signal obtained by the photodetector into a digital signal. The signal processing unit may include an A / D converter as a device that performs the conversion. The digital signal obtained by the conversion by the signal processing unit may be transmitted to the information processing unit 6103. The digital signal may be handled by the information processing unit 6103 as data related to light (hereinafter also referred to as "light data"). The light data may be light data including, for example, fluorescent light data. More specifically, the light data may be light intensity data, and the light intensity may be light intensity data of light including fluorescent light (which may include feature quantities such as area, height, and width).
[0173] (Information Processing Unit) The information processing unit 6103 includes, for example, a processing unit that processes various data (e.g., optical data) and a storage unit that stores various data. When the processing unit acquires optical data corresponding to a fluorescent dye from the detection unit 6102, the processing unit may perform fluorescence spillover correction (compensation processing) on the light intensity data. Furthermore, in the case of a spectral flow cytometer, the processing unit performs fluorescence separation processing on the optical data to acquire light intensity data corresponding to the fluorescent dye. The fluorescence separation processing may be performed, for example, according to the unmixing method described in Japanese Patent Application Laid-Open No. 2011-232259. When the detection unit 6102 includes an image sensor, the processing unit may acquire morphological information of bioparticles based on images acquired by the image sensor. The storage unit may be configured to store the acquired optical data. The storage unit may further be configured to store spectral reference data used in the unmixing processing.
[0174] If the biological sample analyzer 6100 includes a fractionating unit 6104 (described below), the information processing unit 6103 can determine whether to fractionate bioparticles based on the optical data and / or morphological information. The information processing unit 6103 can then control the fractionating unit 6104 based on the result of this determination, allowing the fractionating unit 6104 to fractionate the bioparticles.
[0175] The information processing unit 6103 may be configured to output various data (e.g., optical data and images). For example, the information processing unit 6103 may output various data (e.g., two-dimensional plots, spectral plots, etc.) generated based on the optical data. The information processing unit 6103 may also be configured to accept input of various data, such as accepting gating processing on a plot by a user. The information processing unit 6103 may include an output unit (e.g., a display, etc.) or an input unit (e.g., a keyboard, etc.) for executing the output or input.
[0176] The information processing unit 6103 may be configured as a general-purpose computer, for example, as an information processing device including a CPU, RAM, and ROM. The information processing unit 6103 may be included in a housing that includes the light irradiation unit 6101 and the detection unit 6102, or may be located outside the housing. Furthermore, various processes or functions performed by the information processing unit 6103 may be realized by a server computer or a cloud connected via a network.
[0177] (Sorting unit) The sorting unit 6104 sorts the bioparticles according to the determination result by the information processing unit 6103. The sorting method may be a method of generating droplets containing bioparticles by vibration, applying an electric charge to the droplets to be sorted, and controlling the direction of travel of the droplets using electrodes. The sorting method may also be a method of controlling the direction of travel of the bioparticles within the flow channel structure to perform sorting. The flow channel structure is provided with, for example, a control mechanism using pressure (spray or suction) or electric charge. An example of such a flow channel structure is a chip (for example, the chip described in JP 2020-76736 A) having a flow channel structure in which a flow channel C branches downstream into a recovery flow channel and a waste flow channel, and specific bioparticles are recovered into the recovery flow channel.
[0178] The present technology may be configured as follows: [1] A nucleic acid analysis method comprising: a base sequence decoding step of decoding the base sequences of multiple types of circular nucleic acids formed into a circle by randomly linking multiple types of DNA fragments; a repeat unit identifying step of identifying the multiple types of circular nucleic acids by determining the number of bases of the circular nucleic acid from the base sequence decoded in the base sequence decoding step; and a DNA fragment identifying step of identifying regions of DNA fragments contained in the multiple types of circular nucleic acids identified in the repeat unit identifying step. [2] The nucleic acid analysis method according to [1], wherein the multiple types of DNA fragments consist of all DNA fragments contained in the target of analysis. [3] The nucleic acid analysis method according to [1] or [2], wherein the multiple types of DNA fragments are cDNA fragments derived from mRNA. [4] The nucleic acid analysis method according to [3], wherein the mRNA is derived from a single cell. [5] The nucleic acid analysis method according to any one of [1] to [4], further comprising, after the DNA fragment identifying step, a duplicate elimination step of comparing the order of arrangement of DNA fragments contained in the multiple types of circular nucleic acids and excluding circular nucleic acids having the same order. [6] The nucleic acid analysis method according to any one of [1] to [5], wherein the number of bases in the repeat unit identifying step is determined by analyzing the period of a maximum value occurring in an autocorrelation sequence of the decoded base sequence. [7] The nucleic acid analysis method according to [6], wherein the autocorrelation sequence is weighted using a transition of an average value calculated from data within a specific range. [8] The nucleic acid analysis method according to any one of [1] to [7], wherein the number of bases in the repeat unit identifying step is determined by a maximum value in a component distribution of a period of a spectral sequence obtained by spectral analysis of the decoded base sequence. [9] The nucleic acid analysis method according to [8], wherein the spectral sequence is weighted using a transition of an average value calculated from data within a specific range.
[10] The nucleic acid analysis method according to any one of [1] to [9], wherein the number of bases in the repeat unit identifying step is determined by a maximum value occurring in the autocorrelation sequence of the decoded base sequence, which coincides with a maximum value in a component distribution of a period of a spectral sequence obtained by spectral analysis of the decoded base sequence.
[11] The nucleic acid analysis method according to any one of [1] to
[10] , comprising a base sequencing step of determining the base sequences of the identical DNA fragments by comparing the base sequences of multiple DNA fragments identified as the identical DNA fragments after the DNA fragment identification step.
[12] The nucleic acid analysis method according to any one of [1] to
[11] , wherein in the DNA fragment identification step, the region of the DNA fragment is identified by identifying the sequence of a linking portion that links the DNA fragments in the circular nucleic acid.
[13] A method for analyzing the expression level of mRNA in a single cell, using the nucleic acid analysis method according to any one of [4] to
[12] .
[14] The analysis method according to
[13] , wherein the single cell is a cell sorted by a biological sample analyzer.
[15] A nucleic acid analysis program comprising: a base sequence decoding step for decoding the base sequences of multiple types of circular nucleic acids formed into circles by randomly linking multiple types of DNA fragments; a repeat unit identifying step for identifying the multiple types of circular nucleic acids by specifying the number of bases of the circular nucleic acid from the base sequence decoded in the base sequence decoding step; and a DNA fragment identifying step for identifying regions of DNA fragments contained in the multiple types of circular nucleic acids identified in the repeat unit identifying step.
[16] The nucleic acid analysis program according to
[15] , further comprising, after the DNA fragment identifying step, a duplicate elimination step for comparing the order in which DNA fragments contained in the multiple types of circular nucleic acids are arranged, thereby eliminating circular nucleic acids having the same order.
[17] An information recording medium having the nucleic acid analysis program according to
[15] or
[16] implemented thereon.
[18] An information processing system having the nucleic acid analysis program according to
[15] or
[16] implemented thereon.
[0179] 10 Nucleic acid linking part 11 Target nucleic acid capturing part 12 Complementary strand capturing part 13 Double-stranded part 14 Sequence addition part
Claims
1. A nucleic acid analysis method comprising: a base sequence decoding step for decoding the base sequences of multiple types of circular nucleic acids formed into circles by randomly linking multiple types of DNA fragments; a repeat unit identification step for identifying the multiple types of circular nucleic acids by determining the number of bases in the circular nucleic acid from the base sequence decoded in the base sequence decoding step; and a DNA fragment identification step for identifying regions of DNA fragments contained in the multiple types of circular nucleic acids identified in the repeat unit identification step.
2. The nucleic acid analysis method according to claim 1, wherein the plurality of types of DNA fragments are composed of all DNA fragments contained in the subject of analysis.
3. The nucleic acid analysis method according to claim 1, wherein the plurality of types of DNA fragments are cDNA fragments derived from mRNA.
4. The nucleic acid analysis method according to claim 3, wherein the mRNA is derived from a single cell.
5. A nucleic acid analysis method as described in claim 1, which includes, after the DNA fragment identification step, a duplication elimination step of comparing the order of arrangement of DNA fragments contained in the multiple types of circular nucleic acids to eliminate circular nucleic acids having the same order.
6. A nucleic acid analysis method according to claim 1, wherein the number of bases is identified in the repeat unit identification step by analyzing the period of maximum values occurring in the autocorrelation sequence of the decoded base sequence.
7. The nucleic acid analysis method according to claim 6, wherein the autocorrelation sequence is weighted using a transition of the average value calculated from data within a specific range.
8. A nucleic acid analysis method according to claim 1, wherein the number of bases in the repeat unit identification step is determined by the maximum value of the component distribution of the period of the spectral sequence obtained by spectral analysis of the decoded base sequence.
9. The nucleic acid analysis method according to claim 8, wherein the spectral sequence is weighted using a transition of an average value calculated from data within a specific range.
10. A nucleic acid analysis method as described in claim 1, wherein the number of bases is identified in the repeat unit identification step by a maximum value occurring in the autocorrelation sequence of the decoded base sequence, which coincides with a maximum value of the component distribution of the period of the spectral sequence obtained by spectral analysis of the decoded base sequence.
11. A nucleic acid analysis method as described in claim 1, further comprising a base sequence determination step of determining the base sequence of the identical DNA fragment by comparing the base sequences of multiple DNA fragments identified as the same DNA fragment after the DNA fragment identification step.
12. A nucleic acid analysis method according to claim 1, wherein in the DNA fragment identification step, the region of the DNA fragment is identified by identifying the sequence of the linking portion that links the DNA fragments in the circular nucleic acid.
13. A method for analyzing the expression level of mRNA in a single cell, using the nucleic acid analysis method described in claim 4.
14. The analytical method according to claim 13, wherein the single cell is a cell sorted by a biological sample analyzer.
15. A nucleic acid analysis program that executes a base sequence decoding step for decoding the base sequences of multiple types of circular nucleic acids formed into circles by randomly linking multiple types of DNA fragments; executes a repeat unit identification step for identifying the multiple types of circular nucleic acids by determining the number of bases of the circular nucleic acid from the base sequence decoded by the base sequence decoding step; and executes a DNA fragment identification step for identifying regions of DNA fragments contained in the multiple types of circular nucleic acids identified by the repeat unit identification step.
16. A nucleic acid analysis program as described in claim 15, which, after performing the DNA fragment identification step, performs a duplication elimination step of comparing the arrangement order of DNA fragments contained in the multiple types of circular nucleic acids to eliminate circular nucleic acids having the same order.
17. An information recording medium on which the nucleic acid analysis program according to claim 15 is implemented.
18. An information processing system that implements the nucleic acid analysis program according to claim 15.
Citation Information
Patent Citations
Method for exclusive selection of circularized DNA from monomolecular DNA when circularizing DNA molecules
WO2013031700A1