Specification of nucleic acid molecule using identification molecule
The nucleic acid molecule complex with an identification molecule enables efficient and cost-effective identification of long nucleic acid sequences by sequencing the identification molecule, addressing the inefficiencies of existing methods.
Patent Information
- Application Number
- PCT/JP2025/013859
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-08
- Filing Date
- 2025-04-07
- Publication Date
- 2025-10-16
AI Technical Summary
Current methods for determining the full-length structure of long nucleic acid molecules are time-consuming and costly, and existing sequencers require multiple sequence analyses, which can introduce read errors and increase workload, especially when analyzing multiple candidates simultaneously.
A nucleic acid molecule complex is created containing a nucleic acid molecule with a known sequence and an identification molecule with a specific sequence, allowing the nucleic acid molecule's sequence to be identified by determining the identification molecule's sequence without sequencing the nucleic acid molecule itself, using a general-purpose device.
This approach simplifies and cost-effectively identifies long nucleic acid sequences by determining the identification molecule's sequence, reducing the need for repeated sequencing of the nucleic acid molecule, thereby decreasing time and financial effort.
Smart Images

Figure JPOXMLDOC01-APPB-T000001 
Figure JPOXMLDOC01-APPB-T000002 
Figure 00000032_0000
Abstract
Description
Identifying nucleic acid molecules using identifier molecules
[0001] The present invention relates to an invention of a method for identifying the structure of a nucleic acid molecule whose sequence has been identified using an identifier, an invention of a nucleic acid complex containing an identifier structure used for identifying the structure and a nucleic acid molecule whose sequence has been identified, or an invention of a library composed of such nucleic acid complexes.
[0002] In modern medical and biological research, genetic engineering has been established as a technology that supports basic research, and its related technology, gene sequence analysis (base sequence analysis), is a very important element. For this reason, various instruments based on various principles have been developed for deciphering base sequences, and the Sanger method, which deciphers fluorescently labeled nucleic acids one by one, has long been the mainstream.
[0003] In recent years, however, the advent of next-generation sequencers (NGS) has made it possible to simultaneously analyze millions of base sequences, which previously were analyzed individually, achieving an overwhelming increase in throughput. While the throughput differs between next-generation sequencers and electrophoresis-based or capillary-based instruments used in the Sanger method, the readable sequence length is not significantly different, at around 500 to 1,000 bases on general-purpose instruments.
[0004] In biomedical research, the selection of binding molecules that interact with specific molecules is a crucial topic. For example, the discovery of inhibitory or activating molecules that react with physiologically active molecules could potentially lead to therapeutic drugs for treating diseases caused by these molecules. Furthermore, in biochemical reactions, molecules that react with specific molecules have a wide range of applications, such as evaluation and production systems. Taking antibodies as an example, antibodies that bind to disease-causing molecules are recognized as effective therapeutic agents, and they are also highly valuable industrially, being used in diagnostic reagents and as components for purifying specific molecules. In addition to antibodies, research on binding molecules for partial proteins such as peptides is also actively being conducted, and the development of effective selection (screening) methods is desired.
[0005] Generally, screening for molecules that interact with a given molecule is effective when identifying the target binding molecule from a large number of candidates, known as a library. To date, libraries of small molecules, antibodies, and peptides have been frequently used. In this method, the given molecule is mixed with a group of molecules contained in the library, and reactive and non-reactive molecules are separated. By repeating this process, binding molecules can be identified. Subsequently, to identify the binding molecule, molecular information (e.g., nucleic acid sequence information) is analyzed, and the transcripts and translations are matched with the gene sequence.
[0006] The sequencer mentioned above is used as the genetic analysis device for this purpose, but as mentioned there, the analysis range is 500 to 1000 bases. Therefore, to analyze molecules longer than this, it is necessary to piece together the analysis information, which requires multiple sequence analyses, resulting in several times the effort required.
[0007] In other words, when determining the full-length structure of a long nucleic acid molecule using a conventional sequencer, the entire length cannot be determined in a single sequence analysis, and multiple sequences must be spliced together. This type of sequence analysis has the problem of the possibility of read errors and the high workload of analyzing multiple sequences.
[0008] In addition, it is often the case that multiple genes are combined and fused into a single protein molecule. In such cases, the workload of connecting multiple reads and decoding using reverse sequence analysis are required, which poses time and cost constraints in research and development where multiple candidates are analyzed simultaneously.
[0009] Using a long-read sequencer, it is possible to analyze gene sequences of tens of thousands of bases in length. However, currently, DNA sequence analysis using a long-read sequencer is extremely time-consuming and expensive, and due to time and financial constraints on research and development, it is not currently a method that can be easily used.
[0010] Patent No. 6338235 "Method for screening and producing minibodies" US20090088327A1 "Method for sequencing a polynucleotide template"
[0011] Sanger F, Nicklen S, Coulson AR. DNA sequencing with chain-terminating inhibitors. Proc Natl Acad Sci US A. 1977 Dec;74(12):5463-7.Sanger F. The early days of DNA sequences. Nat Med. 2001 Mar;7(3):267-8.Shendure J, Ji H. Next-generation DNA sequencing. Nat Biotechnol. 2008 Oct;26(10):1135-45.
[0012] An object of the present invention is to analyze and identify long gene sequences more simply, more cost-effectively, and more efficiently. Another object of the present invention is to efficiently perform such analysis by using a general-purpose device.
[0013] The inventors of the present invention have provided a nucleic acid molecule complex that contains, in a single molecule, a nucleic acid molecule whose sequence has been identified and an identification molecule that contains a specific sequence for identifying the nucleic acid molecule, and have revealed that the sequence of the nucleic acid molecule can be identified based on previously obtained information, without having to determine the sequence of the nucleic acid molecule each time, simply by determining the specific sequence of the identification molecule within the overall structure of the nucleic acid molecule complex.
[0014] More specifically, in order to solve the above-mentioned problems, the present application provides the following aspects: [1]: A nucleic acid molecule complex comprising, in a single molecule, a combination of (A) a nucleic acid molecule whose sequence has been identified, and (B) an identification molecule comprising, in a single molecule, a specific sequence for identifying the (A) nucleic acid molecule; [2]: The nucleic acid molecule complex according to [1], in which the (B) identification molecule is composed of nucleic acid; [3]: The nucleic acid molecule complex according to [1] or [2], which comprises an additional nucleic acid molecule between the (A) nucleic acid molecule and the (B) identification molecule; [4]: The nucleic acid molecule complex according to [1] or [2], which further comprises (C) another nucleic acid molecule between the (B) identification molecule and the (C) other nucleic acid molecule at a position such that the (A) nucleic acid molecule is not present; [5]: A nucleic acid molecule library comprising, in a single molecule, a combination of (A) a nucleic acid molecule whose sequence has been identified, and (B) an identification molecule comprising a specific sequence for identifying the (A) nucleic acid molecule, The nucleic acid molecule library as described above, in which the specific sequence of each (B) identification molecule does not exist identically for each nucleic acid molecule complex; [6]: The nucleic acid molecule library according to [5], in which the (B) identification molecule in each nucleic acid molecule complex contained in the library is composed of a nucleic acid; [7]: The nucleic acid molecule library according to [5] or [6], in which an additional nucleic acid molecule is contained between the (A) nucleic acid molecule and the (B) identification molecule in each nucleic acid molecule complex contained in the library; [8]: The nucleic acid molecule library according to [5] or [6], in which each nucleic acid molecule complex contained in the library further contains another (C) nucleic acid molecule at a position such that the (A) nucleic acid molecule is not present between each (B) identification molecule and the other (C) nucleic acid molecule; [9]: A method for identifying the sequence of a nucleic acid molecule complex containing, in a single molecule, a combination of (A) a nucleic acid molecule whose sequence has been identified and (B) an identification molecule containing a specific sequence for identifying the (A) nucleic acid molecule, the method comprising: determining the specific sequence of the (B) identification molecule contained in the nucleic acid molecule complex; a step of identifying the sequence of the (A) nucleic acid molecule based on the sequence of the (B) identification molecule without determining the sequence of the (A) nucleic acid molecule;
[10] : A method for determining the sequence of an (A) nucleic acid molecule according to [9], wherein the (B) identification molecule is composed of a nucleic acid;
[11] : A method for determining the sequence of an (A) nucleic acid molecule according to [9] or
[10] , wherein an additional nucleic acid molecule is included between the (A) nucleic acid molecule and the (B) identification molecule;
[12] : A method for determining the sequence of a nucleic acid molecule according to [9] or
[10] , wherein the (A) nucleic acid molecule is further included at a position where the (A) nucleic acid molecule is not present between the (B) identification molecule and the (C) other nucleic acid molecule;
[13] : A method for determining the sequence of a nucleic acid molecule complex comprising, in a single molecule, a combination of: (A) a first nucleic acid molecule whose sequence has been determined; (B) an identification molecule comprising a specific sequence for identifying the (A) nucleic acid molecule; and (C) a second nucleic acid molecule, wherein the (A) nucleic acid molecule is included at a position where the (A) nucleic acid molecule is not present between the (B) identification molecule and the (C) second nucleic acid molecule, A method for determining the sequence of a nucleic acid molecule complex, comprising: determining the sequence of the (C) second nucleic acid molecule and the sequence of the (B) identification molecule; and determining the sequence of the (A) first nucleic acid molecule from the sequence of the (B) identification molecule without determining the sequence of the (A) first nucleic acid molecule;
[14] : A method for determining the sequence of a nucleic acid molecule complex according to
[13] , wherein the (B) identification molecule is composed of nucleic acid;
[15] : A method for determining the sequence of a nucleic acid molecule complex according to
[13] or
[14] , which comprises an additional nucleic acid molecule between the (A) first nucleic acid molecule and the (B) identification molecule;
[16] : A method for determining the sequence of a nucleic acid molecule complex according to [9] or
[10] , which further comprises a (C') other nucleic acid molecule at a position where the (A) first nucleic acid molecule is not present between the (B) identification molecule or the (C) second nucleic acid molecule and the (C') other nucleic acid molecule;
[0015] The present invention provides a nucleic acid molecule complex that contains, in a single molecule, a combination of (A) a nucleic acid molecule whose sequence has been identified and (B) an identification molecule that contains a specific sequence for identifying the (A) nucleic acid molecule, and has revealed that the sequence of the (A) nucleic acid molecule can be identified by simply determining the specific sequence of the (B) identification molecule within the overall sequence structure of the nucleic acid molecule complex, without determining the sequence of the (A) nucleic acid molecule itself.
[0016] Figure 1 is a diagram showing an outline of the sequence of a nucleic acid molecule complex containing, in a single molecule, a combination of (A) a nucleic acid molecule whose sequence has been identified; and (B) an identification molecule containing a specific sequence for identifying the (A) nucleic acid molecule. Figure 2 is a diagram showing that the sequence of the (A) nucleic acid molecule can be identified by determining only the sequence of the (B) identification molecule in the nucleic acid molecule complex having the sequence shown in Figure 1, without determining the sequence structure of the (A) nucleic acid molecule whose sequence has been identified. Figure 3 is a diagram showing an outline of the sequence of a nucleic acid molecule complex having, from the 5' to 3' end, a structure in which (C) another nucleic acid molecule, (B) an identification molecule, and (A) a nucleic acid molecule whose sequence has been identified are arranged in tandem, or a structure in which (A) a nucleic acid molecule whose sequence has been identified, (B) an identification molecule, and (C) another nucleic acid molecule are arranged in tandem. Figure 4 shows that, among the nucleic acid molecule complexes shown in Figure 3, the sequence of the (A) nucleic acid molecule can be identified without determining the sequence structure of the (A) nucleic acid molecule whose sequence has already been determined by determining the sequences of the (B) identification molecule and (C) other nucleic acid molecules. Figure 5-1 shows an example of a nucleic acid molecule library containing multiple nucleic acid molecule complexes each containing, in a single molecule, (A) a nucleic acid molecule whose sequence has already been determined; and (B) an identification molecule comprising, on the 5' side of the (A) nucleic acid molecule, a specific sequence for identifying the (A) nucleic acid molecule. Figure 5-2 shows an example of a nucleic acid molecule library containing multiple nucleic acid molecule complexes each containing, in a single molecule, (A) a nucleic acid molecule whose sequence has already been determined; and (B) an identification molecule comprising, on the 3' side of the (A) nucleic acid molecule, a specific sequence for identifying the (A) nucleic acid molecule. Figure 6-1 is a diagram showing an overview of a nucleic acid molecule library containing, as an example, a plurality of nucleic acid molecule complexes each having a structure in which, from the 5' to the 3' side, (C) another nucleic acid molecule, (B) an identification molecule, and (A) a nucleic acid molecule whose sequence has been identified are arranged in tandem. Figure 6-2 is a diagram showing, as an example, a nucleic acid molecule library containing, as an example, a plurality of nucleic acid molecule complexes each having a structure in which, from the 5' to the 3' side, (B) an identification molecule, (C) another nucleic acid molecule, and (A) a nucleic acid molecule whose sequence has been identified are arranged in tandem.Figure 6-3 is a diagram showing an example of a nucleic acid molecule library containing a plurality of nucleic acid molecule complexes in which (A) a nucleic acid molecule with a specified sequence, (B) an identifier molecule, and (C) another nucleic acid molecule are arranged in tandem from the 5' to 3' end. Figure 6-4 is a diagram showing an example of a nucleic acid molecule library containing a plurality of nucleic acid molecule complexes in which (A) a nucleic acid molecule with a specified sequence, (C) another nucleic acid molecule, and (B) an identifier molecule are arranged in tandem from the 5' to 3' end. Figure 7 is a diagram showing an example of a nucleic acid molecule complex constituting an antibody library, showing the outline of the sequence of the nucleic acid molecule complex in which, from the 5' to 3' end, the nucleic acid molecule complex contains, (C) a VH gene as another nucleic acid molecule, (B) an identifier molecule, and (A) a nucleic acid molecule with a specified sequence (an antibody L chain gene consisting of a VL gene and a CL gene). Figure 8-1 shows the results of analyzing the full-length sequences of (C) the antibody VH gene as another nucleic acid molecule, (B) the identifier molecule, and the antibody L-chain gene of a nucleic acid molecule complex constituting an antibody library according to a conventional method, by separating them into sequences determined using a forward primer and sequences obtained using a reverse primer. Figure 8-2 shows the results of generating full-length sequences based on the superposition of sequences determined using a forward primer and sequences determined using a reverse primer, obtained according to a conventional method. Figure 9 shows the results of analyzing the nucleotide sequences of (C) the antibody VH gene as another nucleic acid molecule and (B) the identifier molecule of a nucleic acid molecule complex constituting an antibody library according to the method of the present invention, as well as the nucleotide sequence of (A) the antibody L-chain gene, the sequence of which has been previously assigned to the identifier molecule (B), and the full-length nucleotide sequence of the nucleic acid molecule complex identified by superposing these pieces of information. Figure 10 shows an alignment comparing the full-length base sequence of a nucleic acid molecule complex whose full-length base sequence was determined according to a conventional method (see Figure 8-2) with the full-length base sequence of a nucleic acid molecule complex obtained by combining the base sequence whose partial base sequence was determined according to the method of the present invention with information on the antibody L chain gene assigned to the identification molecule (see Figure 9).
[0017] As a result of extensive research conducted by the inventors of the present invention to solve the above-mentioned problems, they have demonstrated that the above-mentioned problems can be solved by producing a nucleic acid molecule complex having a new structure.
[0018] Specifically, in one aspect, the present invention can provide a nucleic acid molecule complex that contains, in a single molecule, a combination of (A) a nucleic acid molecule whose sequence has been identified and (B) an identification molecule that includes a specific sequence for identifying the (A) nucleic acid molecule. Here, when the (A) nucleic acid molecule and the (B) identification molecule are "combined in a single molecule," it means that the (A) nucleic acid molecule and the (B) identification molecule are combined in a single molecule without restrictions on the order or the distance between the two molecules, and that when the (B) identification molecule is sequenced, the (A) nucleic acid molecule is located somewhere in the same molecule.
[0019] When the (B) identification molecule is a substance composed of nucleic acid, this means that the (B) identification molecule may be on the 5' side or the 3' side of the (A) nucleic acid molecule, and that other components may be present between the (A) nucleic acid molecule and the (B) identification molecule. In this nucleic acid molecule complex, a (B) identification molecule containing a specific sequence for identifying the (A) nucleic acid molecule is combined in a one-to-one correspondence with a (A) nucleic acid molecule whose sequence has been identified, and the sequence structure is determined so that one type of (A) nucleic acid molecule whose sequence has been identified is always assigned to one (B) identification molecule (see Figure 1).
[0020] The (B) identification molecule used in the present invention can have any sequence as long as it is a nucleic acid molecule or a peptide molecule. In the present invention, the sequence structure of such a (B) identification molecule is also referred to as an "identifier structure." That is, a recognition molecule having a specific sequence is assigned to a (A) nucleic acid molecule whose sequence has already been identified, and the recognition molecule is combined with the nucleic acid molecule and arranged in a single molecule. The detailed sequence structure of this (B) identification molecule is a specific sequence for identifying the aforementioned (A) nucleic acid molecule. However, it is sufficient for this (B) identification molecule to be able to identify specific information based on its sequence; the (B) identification molecule itself may or may not have any physiological or biochemical function.
[0021] (B) When the identification molecule is a substance composed of nucleic acid, the identification molecule can be composed of a random sequence of four types of bases (adenine, thymine, guanine, and cytosine), and by arranging n bases, 4 n As mentioned above, such nucleic acid (B) recognition molecules do not need to have any physiological or biochemical function depending on the sequence, and the amino acids in the sequence may or may not be specified.
[0022] Examples of such nucleic acid molecules (B) for identification molecules include, but are not limited to, nucleic acid molecules encoding common linkers that serve to link proteins, such as a nucleic acid molecule encoding a 15-amino acid sequence called a GS linker consisting of glycine (G) and serine (S) (GGGGSGGGGSGGGGS, SEQ ID NO: 1), or a nucleic acid molecule encoding a molecule commonly called a tag, such as a nucleic acid molecule encoding a His-tag (sequence HHHHHH, SEQ ID NO: 2).
[0023] In addition, (B) when peptide molecules are used as the recognition molecules, 20 types of amino acid sequences that can be analyzed by a peptide sequencer can be randomly constructed, and by arranging n amino acids, 20 nIt is possible to create a variety of recognition molecules. The recognition molecules (B) using such peptides also do not need to have any physiological or biochemical function depending on their sequence.
[0024] Specifically, when a peptide molecule is used as the (B) recognition molecule, it is possible to identify the (A) molecule with a specified sequence from a complex consisting of the (B) recognition molecule and an arbitrary amino acid sequence with a specified sequence by simply determining the sequence by N-terminal analysis of the (B) protein, without analyzing the genetic information. This differs from systems in which genetic information and proteins are linked (such as phage display systems and yeast display systems), and can be used in secretion systems (systems that are not displayed by bacteria, etc.).
[0025] By analyzing the detailed sequence of the (B) identification molecule thus prepared, the sequence of the (A) nucleic acid molecule can be identified based solely on its relationship with the identified sequence of the (A) nucleic acid molecule, in which the sequences of the (B) identification molecules are specifically assigned and combined in a one-to-one correspondence and contained in a single molecule, without having to determine the sequence of the (A) nucleic acid molecule each time. In other words, the sequence of the (A) nucleic acid molecule can be identified by determining only the sequence of the (B) identification molecule in this nucleic acid molecule complex, without determining the sequence of the (A) nucleic acid molecule whose sequence has already been identified (see Figure 2).
[0026] By adopting the above-described configuration of the present invention, the sequence of the (A) nucleic acid molecule can be identified simply by determining the structure of the (B) identification molecule contained in the nucleic acid molecule complex, without having to perform sequence analysis of the (A) nucleic acid molecule whose sequence has already been identified each time. Therefore, the (A) nucleic acid molecule may have any number of bases, and may be so long that its structure cannot be determined in a single sequence determination when using a general-purpose structural analysis device, for example.
[0027] When identifying the sequence of the (A) nucleic acid molecule, the (A) nucleic acid molecule and the (B) identifying molecule are arranged so that the sequence of the nucleic acid molecule complex is structured such that sequencing of the (B) identifying molecule can be completed without sequencing the (A) nucleic acid molecule, and the start site of sequencing is determined. For example, the sequence of a nucleic acid molecule is generally determined by determining the bases from the 5' side to the 3' side of the nucleic acid molecule. For example, in the present invention, (A) a (B) identification molecule is placed on the 5' side of a nucleic acid molecule whose sequence has already been determined (i.e., a 5'-(B)-(A)-3' structure) and the sequence is determined from the 5' side of the (B) identification molecule, or (A) a (B) identification molecule is placed on the 3' side of a nucleic acid molecule whose sequence has already been determined (i.e., a 5'-(A)-(B)-3' structure) and the sequence is determined from the 5' side of the (B) identification molecule on the 3' side of the (A) nucleic acid molecule. This makes it possible to complete the sequence determination of the (B) identification molecule without having to sequence the (A) nucleic acid molecule each time.
[0028] As long as the (A) nucleic acid molecule and the (B) identification molecule are combined and contained in one molecule, the (A) nucleic acid molecule and the (B) identification molecule may be directly linked without using an additional nucleic acid molecule, or an additional nucleic acid molecule of any length may be included between them. By using an additional nucleic acid molecule, it becomes possible to add a sequence for a primer, to include an additional functional molecule, etc.
[0029] Using the nucleic acid molecule complex of this configuration, (A) a nucleic acid molecule with a specified sequence can be identified based on sequencing of the (B) identifier molecule, thereby selecting a nucleic acid molecule complex containing (A) a nucleic acid molecule with a specified sequence and (B) an identifier molecule. The nucleic acid molecule complex selected in this manner can be applied directly to a protein expression system to express a protein encoded by (A) a nucleic acid molecule with a specified sequence.
[0030] The nucleic acid molecule complex of the present invention may further contain another nucleic acid molecule (C) at a position where the (A) nucleic acid molecule is not present between the (B) identification molecule and the other nucleic acid molecule (C). This is intended to exclude sequencing of the (A) nucleic acid molecule when sequencing the (B) identification molecule and the other nucleic acid molecule (C) (in any order) in terms of positional relationship. That is, a nucleic acid molecule complex having this configuration has a structure in which, from the 5' side to the 3' side, (C) another nucleic acid molecule, (B) an identification molecule, and (A) a nucleic acid molecule whose sequence has been identified are arranged in series (i.e., a 5'-(C)-(B)-(A)-3' configuration), or a structure in which, from the 5' side to the 3' side, (B) an identification molecule, (C) another nucleic acid molecule, and (A) a nucleic acid molecule whose sequence has been identified are arranged in series (i.e., a 5'-(B)-(C)-(A)-3' configuration), or a structure in which, from the 5' side to the 3' side, (A) a nucleic acid molecule whose sequence has been identified, (B) an identification molecule, and (C) another nucleic acid molecule are arranged in series (i.e., a 5'-(A)-(B)-(C)-3' configuration), - Those having a structure in which (A) a nucleic acid molecule with a specified sequence, (C) another nucleic acid molecule, and (B) an identifier molecule are arranged in series from the 5' to the 3' end (i.e., a 5'-(A)-(C)-(B)-3' structure) are included (see Figure 3).
[0031] For example, in the case of a nucleic acid molecule complex having the first structure (i.e., the 5'-(C)-(B)-(A)-3' configuration), the sequence of another nucleic acid molecule (C) is determined from the 5' side, and then the sequence of the identifier molecule (B) is determined. At that point, the sequence of the nucleic acid molecule (A) whose sequence has been previously determined can be identified based on the structure of the identifier molecule (B). Alternatively, in the case of a nucleic acid molecule complex having the third structure (i.e., the 5'-(A)-(B)-(C)-3' configuration), the sequence of the identifier molecule (B) is determined from the 3' side of the nucleic acid molecule (A) and the 5' side of the identifier molecule (B). Then, the sequence of another nucleic acid molecule (C) is determined. At that point, the sequence of the nucleic acid molecule (A) whose sequence has been previously determined can be identified based on the structure of the identifier molecule (B) (see Figure 4). Due to these characteristics, in one embodiment of the present invention, by sequencing (B) the identification molecule and (C) other nucleic acid molecules, it is possible to identify the sequences of (C) other nucleic acid molecules and (A) nucleic acid molecules whose sequences have already been identified, without having to sequence (A) the nucleic acid molecule each time, thereby significantly reducing the time and financial effort required for sequence analysis.
[0032] Using a nucleic acid molecule complex further comprising (C) another nucleic acid molecule, the sequence of the (C) other nucleic acid molecule and the sequence of the (A) nucleic acid molecule whose sequence has been identified can be identified based on the sequencing of the (C) other nucleic acid molecule and the sequencing of the (B) identifier molecule, without having to sequence the (A) nucleic acid molecule each time, and since these are contained in a single molecule, it can be easily selected as a nucleic acid molecule complex containing a combination of the (C) other nucleic acid molecule and the (A) nucleic acid molecule whose sequence has been identified. The nucleic acid molecule complex selected in this manner can be applied directly to a protein expression system to express a protein encoded by the (C) other nucleic acid molecule and a protein encoded by the (A) nucleic acid molecule whose sequence has been identified.
[0033] In the present invention, a nucleic acid molecule library containing a plurality of (A) nucleic acid molecules whose sequences have been identified can be formed by preparing and combining a plurality of the above-mentioned nucleic acid molecule complexes. Specifically, in one aspect, the present invention can provide a nucleic acid molecule library containing a plurality of nucleic acid molecule complexes each containing, in a single molecule, a combination of (A) a nucleic acid molecule whose sequence has been identified and (B) an identification molecule containing a specific sequence for identifying the (A) nucleic acid molecule, wherein the specific sequence of each (B) identification molecule is not the same for any of the nucleic acid molecule complexes (see Figures 5-1 and 5-2).
[0034] In the present invention, when a plurality of nucleic acid molecule complexes having the above-described configuration are compiled into a library, the specific sequence of each (B) identification molecule is designed so that no identical sequences exist among the nucleic acid molecule complexes, and the (B) identification molecule and the (A) nucleic acid molecule whose sequence has been identified are always in one-to-one correspondence. By adopting such a structure, it is possible to select a specific (A) nucleic acid molecule whose sequence has been identified from the library simply by determining the sequence of the (B) identification molecule, without having to determine the sequence of the (A) nucleic acid molecule each time.
[0035] As described above, the nucleic acid molecule complex of the present invention may further include another (C) nucleic acid molecule at a position where the (A) nucleic acid molecule is not present between the (B) identification molecule and the other (C) nucleic acid molecule. When a library is created by combining multiple nucleic acid molecule complexes each including another (C) nucleic acid molecule, the library is constructed based on the structural characteristics of the nucleic acid molecule complex, with the (C) other nucleic acid molecule and the (A) nucleic acid molecule whose sequence has been identified being bound via the (B) identification molecule (see Figures 6-1 and 6-2). This structure allows nucleic acid molecule complexes containing another (C) nucleic acid molecule and an (A) nucleic acid molecule whose sequence has been identified to be selected from the library based on the sequencing of the other (C) nucleic acid molecule and the sequencing of the (B) identification molecule, without having to sequence the (A) nucleic acid molecule each time.
[0036] In the present invention, among nucleic acid molecule complexes containing a (B) identification molecule having the above-mentioned characteristics and a (A) nucleic acid molecule whose sequence has been identified in one molecule, by simply determining the sequence of the (B) identification molecule, which is a relatively small molecular region, the sequence of the (A) nucleic acid molecule whose sequence has been identified and bound to the (B) identification molecule in one molecule can be identified based only on information about the relationship between the (B) identification molecule and the (A) nucleic acid molecule whose sequence has been identified, without having to determine the sequence each time. That is, specifically, in one aspect, the present invention can provide a method for identifying the sequence of a nucleic acid molecule complex containing, in one molecule, a (A) nucleic acid molecule whose sequence has been identified and (B) an identification molecule containing a specific sequence for identifying the (A) nucleic acid molecule, the method comprising: determining the specific sequence of the (B) identification molecule contained in the nucleic acid molecule complex; and identifying the sequence of the (A) nucleic acid molecule based on the sequence of the (B) identification molecule, without determining the sequence of the (A) nucleic acid molecule (see Figure 2).
[0037] The present invention aims to more simply, more cost-effectively, and efficiently analyze and identify the sequence of the (A) nucleic acid molecule contained in a nucleic acid molecule complex, and to efficiently perform such analysis by using a general-purpose device. However, it has been shown that by using a nucleic acid molecule complex having the characteristic structure of the present invention, the sequence of the (A) nucleic acid molecule can be identified by determining the sequence of the (B) identification molecule linked to the sequence information of the (A) nucleic acid molecule obtained in advance, without having to analyze the sequence of the (A) nucleic acid molecule each time, thereby solving the above-mentioned problem.
[0038] As described above, the nucleic acid molecule complex of the present invention may further comprise another (C) nucleic acid molecule at a position such that the (A) nucleic acid molecule is not present between the (B) identification molecule and the other (C) nucleic acid molecule. The nucleic acid molecule complex of the present invention further comprising another (C) nucleic acid molecule can identify the sequence of the other (C) nucleic acid molecule and the sequence of the nucleic acid molecule whose sequence has been identified (A) based on the sequencing of the other (C) nucleic acid molecule and the sequencing of the (B) identification molecule, without sequencing the (A) nucleic acid molecule each time.
[0039] Therefore, in another aspect, the present invention can provide a method for identifying the sequence of a nucleic acid molecule complex containing a combination of (A) a first nucleic acid molecule whose sequence has been identified, (B) an identification molecule comprising a specific sequence for identifying the (A) nucleic acid molecule, and (C) a second nucleic acid molecule, the combination being located between the (B) identification molecule and the (C) second nucleic acid molecule at a position such that the (A) nucleic acid molecule is not present, the method comprising: determining the sequence of the (C) second nucleic acid molecule and the sequence of the (B) identification molecule; and identifying the sequence of the (A) first nucleic acid molecule from the sequence of the (B) identification molecule without determining the sequence of the (A) first nucleic acid molecule (see Figure 4).
[0040] In this aspect of the present invention, the nucleic acid molecule complex of the present invention, which is configured to include a (C) second nucleic acid molecule, can identify the sequence of the (C) second nucleic acid molecule and the sequence of the (A) first nucleic acid molecule, the sequence of which has been identified, based on sequencing of the (C) second nucleic acid molecule and sequencing of the (B) identification molecule.
[0041] The nucleic acid molecule complex of this aspect of the present invention may further comprise a (C') other nucleic acid molecule at a position such that the (A) nucleic acid molecule is not present between the (C) second nucleic acid molecule or (B) identifier molecule and the (C') other nucleic acid molecule. This is intended to prevent sequencing of the (A) nucleic acid molecule during sequencing of the (B) identifier molecule, the (C) second nucleic acid molecule, and the (C') other nucleic acid molecule (in any order). Examples of the above configuration include: (C') when another nucleic acid molecule is present between the second nucleic acid molecule (C) and the identifier molecule (B) (i.e., 5'-(C)-(C')-(B)-(A)-3', 5'-(B)-(C')-(C)-(A)-3', 5'-(A)-(C)-(C')-(B)-3', or 5'-(A)-(B)-(C')-(C)-3'); - (C') No other nucleic acid molecule is present between the (C) second nucleic acid molecule and the (B) identifier molecule (i.e., any of the following configurations may be used: 5'-(C')-(C)-(B)-(A)-3', 5'-(C')-(B)-(C)-(A)-3', 5'-(C)-(B)-(C')-(A)-3', 5'-(B)-(C)-(C')-(A)-3', 5'-(A)-(C')-(C)-(B)-3', 5'-(A)-(C')-(B)-(C)-3', 5'-(A)-(C')-(B)-(C')-3', 5'-(A)-(C)-(B)-(C')-3').
[0042] For example, a nucleic acid molecule complex of the present invention containing (C') another nucleic acid molecule in the configuration 5'-(C)-(C')-(B)-(A)-3' can identify the sequence of (C) second nucleic acid molecule-(C') other nucleic acid molecule-(A) first nucleic acid molecule whose sequence has been identified based on sequencing of (C) second nucleic acid molecule, sequencing of (C') other nucleic acid molecule, and sequencing of (B) identifying molecule.
[0043] For example, the nucleic acid molecule complex of the present invention, which contains (C') another nucleic acid molecule in the configuration 5'-(C')-(C)-(B)-(A)-3', can identify the sequence of (C') another nucleic acid molecule-(C) second nucleic acid molecule-(A) first nucleic acid molecule whose sequence has been identified based on sequencing of (C') another nucleic acid molecule, sequencing of (C) second nucleic acid molecule, and sequencing of (B) identifier molecule.
[0044] The present invention will be specifically illustrated by the following examples, which are not intended to limit the present invention in any way.
[0045] Example 1: Design of genetic identifier sequences as identification molecules The purpose of this example was to design (B) nucleic acid molecules (genetic identifiers) that function as identification molecules.
[0046] In this example, a 15-amino acid sequence called the GS linker (GGGGSGGGGSGGGGS, SEQ ID NO: 1), consisting of glycine (G) and serine (S), which is a common linker that plays a role in linking proteins, was selected for designing the genetic identifier as an identification molecule.
[0047] The base sequence encoding both the Gly and Ser amino acids that make up the linker site, which is this recognition molecule, is coded by four codons for Gly: ggt, ggc, gga, ggg; and six codons for Ser: agt, agc, tct, tcc, tca, tcg. Therefore, the number of combinations is (4 x 4 x 4 x 4 x 6). 3 = 3,623,878,656, or more than 3.6 billion combinations. Therefore, when creating a recognition molecule as a nucleic acid molecule, a 15-amino acid linker allows it to function as a recognition molecule with approximately 3.6 billion combinations using only silent mutations without any amino acid mutations.
[0048] As an example, 400 different sequences were designed so that each sequence would be unique across 45 bases, equivalent to these 15 amino acids. Specific combinations were selected from four codons per glycine amino acid and six codons per serine amino acid, as described above, so that all 400 sequences would consist of different base sequences. Some of the specific sequence designs are outlined in Table 1 below.
[0049]
[0050] Example 2: Construction of an antibody library In recent years, next-generation sequencing (NGS) analysis has become standard in antibody development using Phage Display technology. However, in antibody development technology using NGS analysis, the read length of general-purpose NGS machines is 500 bp or less. Therefore, when analyzing genes encoding antibody molecules, for example, NGS is used to analyze only the heavy chain variable region (VH) or the variable domain of heavy chain of heavy chain antibody (VHH antibody), but it is not possible to analyze the full-length scFv (single chain Fv).
[0051] Analysis of the combination of heavy chain variable region (VH chain) and light chain variable region (VL chain) is essential for antibody development. When using current general-purpose NGS machines, the full-length structure of the VH chain and VL chain combination is determined in several steps, with overlapping ends, and the information is then combined to identify the full-length structure.
[0052] When actually creating antibodies, the task of determining the sequence several times and combining the information to identify the full length combination of VH chain and VL chain that constitutes a single antibody is a heavy workload.
[0053] Under these circumstances, this Example was carried out with the aim of constructing a model of an antibody library containing a plurality of nucleic acid complexes that constitute antibodies, as an example of the nucleic acid complex of the present invention, for the purpose of serving as a model for verifying the present invention.
[0054] In this example, the phagemid was designed to ultimately encode a signal sequence and a heavy chain antibody gene sequence as (C) other nucleic acid molecules, the GS linker sequence prepared in Example 1 as (B) an identification molecule, a light chain antibody sequence as (A) nucleic acid molecule for structure identification, and the M13 phage cp3 protein, and was configured to further include an origin of replication for amplification in E. coli and an ampicillin resistance gene for drug selection.
[0055] (2-1) RNA extraction: A cell sample containing human B cells, including normal human bone marrow-derived mononuclear cells, normal human peripheral blood-derived mononuclear cells, and human umbilical cord blood, was prepared as a raw material for constructing a human antibody library. 5-10 × 10 cells were prepared. 8 After adding 1 mL of TRIzol Reagent (Thermo Fisher Scientific) to the cells to disrupt them, 50 μL of 4-bromoanisole was added. After incubation, the mixture was centrifuged and 600 μL of the upper aqueous layer containing RNA was collected. After adding 70% ethanol, the sample was transferred to a spin cartridge. The spin column was washed with 350 μL of Wash Buffer I and 500 μL of Wash Buffer II, and eluted with 100 μL of RNase-free water.
[0056] (2-2) cDNA Synthesis. Five microliters of the RNA prepared in (2-1) was ethanol-precipitated, and the resulting RNA precipitate was suspended in a premixed solution (1 μL of 50 μM Origo d(T)20 primer, 1 μL of 10 mM dNTP Mix, and 11 μL of DEPC-water). A premixed solution (4 μL of 5x SSIV RT Buffer, 1 μL of 100 mM DTT, 1 μL of RNaseOUT Recombinant Rnase Inhibitor, and 1 μL of SuperScript IV Reverse Transcriptase (200 U / μL) (Thermo Fisher Scientific) was then added, and the reaction was carried out at 50°C for 10 minutes and then at 80°C for 10 minutes. After the reaction, the cDNA was purified using a PCR Purification Kit (QIAGEN), and the resulting eluate was adjusted to 100 μL.
[0057] (2-3) Amplification of antibody heavy chain gene fragment Using 15 ng of the cDNA prepared above as a template, PCR was carried out using Q5 polymerase (NEW ENGLAND BioLabs). According to standard methods (see references below), seven forward primers (VH1 to VH7, SEQ ID NOs: 3 to 9) specifically amplifying the heavy chain variable region (VH) and four reverse primers (JH1-2, JH3, JH4-5, and JH6, SEQ ID NOs: 10 to 13) representing six families were used at a final concentration of 0.5 μM (Marks JD et al., J. Mol. Biol. (1991) 222, 581-597; Campbell M et al., Mol. Immunol. (1992) 29, 193-203; Zhu Z, Dimitrov DS. Methods Mol. Biol. 2009;525:129-42; Antibody Phage Display. Methods in Molecular Biology, vol. 178. Humana Press.). The PCR reaction was carried out for 30 cycles consisting of 30 seconds at 94°C, 30 seconds at 65°C, and 2 minutes at 74°C, and the amplified DNA fragment was collected.
[0058]
[0059] In this way, a series of procedures from RNA extraction to cDNA synthesis and amplification of the heavy chain fragment were repeated, and approximately 5 × 10 8 A large number of antibody heavy chain gene fragments were prepared from the cells.
[0060] (2-4) Construction of antibody heavy chain gene library An antibody heavy chain gene library was constructed by incorporating each amplified DNA fragment of each antibody heavy chain gene into a vector for phage display (phagemid vector).
[0061] To recombine the antibody heavy chain gene variable region DNA fragment into the vector, the PCR-amplified fragment was digested with Sfi-I and XhoI restriction enzymes, and the vector was also digested in the same way. The antibody heavy chain gene fragment and the vector fragment were then ligated, and the ligation product was introduced into E. coli as follows to obtain a transformant.
[0062] The ligation product was ethanol precipitated and dissolved in a QW flask. 2 μL of the precipitate was suspended in 20 μL of DH12S (Thermo Fisher Scientific) and electroporated using a Gene Pulser Xcell (Bio-Rad) with the bacterial program (25 μF, 200 Ω, 1.8 kV, cuvette width 0.1 cm). The entire ligation product was then transformed into transfected E. coli, which were then cultured in 2xYT medium containing ampicillin. After overnight incubation, the cells were harvested and the plasmid was purified.
[0063] (2-5) Design of antibody light chain gene The sequence of the antibody light chain gene was designed for use as the nucleic acid molecule (A) with a specified sequence. Specifically, the germline genes of human antibodies are published in "THE INTERNATIONAL IMMUNOGENETICS INFORMATION SYSTEM" (IMGT: https: / / www.Imgt.org / ), and the basic light chain germline gene was designed by referring to this sequence.
[0064] For kappa chains, functional clusters of IGKVs were 35-40 (IGKV1: 17 (+2) species, IGKV2: 9 (+1) species, IGKV3: 5 (+2) species, IGKV4: 1 species, IGKV5: 5 species, IGKV6: 2 species, IGKV7: - species), and for lambda chains, IGLVs were 29-30 (IGLV1: 5 species, IGLV2: 5 species, IGLV3: 8 species, IGLV4: 3 species, IGLV5: 3-4 species, IGLV6: 1 species, IGLV7: 1 species, IGLV8: 1 species, IGLV9: 1 species, IGLV10: 1 species, IGLV11: 1 species) (Williams, SC et al., J. Mol. Biol., 264, 220-232 (1996), Kawasaki, K. et al., Genome Res., 7, 250-261 (1997)).
[0065] To further diversify these basic sequences, artificial mutations were introduced into FR1, FR2, FR3, and FR4 in the framework region and CDR1, CDR2, and CDR3 in the complementarity-determining regions. The abundance of each subgroup was determined based on previous studies (e.g., Winter G, et al., Annu Rev Immunol. (1994) 12:433-55; Prabakaran P, et al., Immunogenetics. 2012 May; 64(5):337-50) to match the abundance of the kappa and lambda chain gene families in vivo. That is, for the kappa chains (IGKV1 to IGKV6), the proportion of each subgroup in the library was approximately IGKV1: 36%, IGKV2: 11%, IGKV3: 35%, IGKV4: 11%, IGKV5: 4%, and IGKV6: 3%, and for the lambda chains (IGLV1 to IGLV10), the proportion of each subgroup in the library was IGLV1: 40%, IGLV2: 16%, IGLV3: 21%, IGLV4: 3%, IGLV5: 6%, IGLV6: 6%, IGLV7: 2%, IGLV8: 2%, IGLV9: 2%, and IGLV10: 2%.
[0066] Each of the light chain genes thus prepared was designed to have 400 unique linker sequences prepared in Example 1 attached to its 5' end in a one-to-one correspondence, and 400 artificial sequences (5'-(B) recognition molecule-(A) light chain gene-3') were artificially synthesized in which 400 arbitrary recognition molecules were matched to the antibody light chain gene sequences arranged in series.
[0067] Specifically, (B) nucleic acid molecules (identifier molecules) specifying the 400 designed linkers were positioned at the 5' end of (A) a light chain antibody gene sequence, which is a nucleic acid molecule with a specified sequence, and 400 light chain gene sequences were artificially designed for the nucleic acid molecules specifying the 400 linkers. An example design is shown below.
[0068] Identification molecule (B) 001: ggtggcggagggagtgggggaggcggtagtggtggcggagggtcc (SEQ ID NO: 14)
[0069] Nucleic acid molecule (A) with a specified array 001 (light chain base sequence): atggcccagagcgtcctgacccaaccaccctccgtcagcgaggcaccgcgccagcgcgtgactatttcttgctctggatcctcgtctaacattggcaataacacggttaattggtaccagcagctcccggggaaagccccgaaactccttatttattcggatgatcttctgccgtctggcgtgtcggatcgcttcagcggtagcaaaagtggcacgtcagcctccttagccatttcccgcctgcagagcgaggatgaagccgattattactgcgcggcctgggatgattcattgagcggttgggtgtttggcggtggaactaaactgacggtgctgggccagccgaaagcggcaccctctgtgaccctttttccgccgagttcagaggaactgcaggcgaataaagcgaccctggtctgcctgatcagtgacttctatccaggtgcggtgacggtagcgtggaaagcagacagtagccctgtcaaagcaggcgtagagacgacgaccccgagtaagcagtctaataataaatacgccgcgtctagttacctgtcattgacgccggaacagtggaaaagtcatcggtcttacagctgtcaggttacccatgagggcagcacggtagaaaaaaccgttgcgccgactgagtgctcg (SEQ ID NO: 15)
[0070] Nucleic acid molecule with a specified sequence containing an identification molecule (the underlined part is the sequence of the identification molecule (B) 001): ggtggcggagggagtgggggaggcggtagtggtggcggagggtccatggcccagagcgtcctgacccaaccaccctccgtcagcgaggcaccgcgccagcgcgtgactatttcttgctctggatcctcgtctaacattggcaataacacggttaattggtaccagcagctcccggggaaagccccgaaactccttatttattcggatgatcttctgccgtctggcgtgtcggatcgcttcagcggtagcaaaagtggcacgtcagcctccttagccatttcccgcctgcagagcgaggatgaagccgattattactgcgcggcctgggatgattcattgagcggttgggtgtttggcggtggaactaaactgacggtgctgggccagccgaaagcggcaccctctgtgaccctttttccgccgagttcagaggaactgcaggcgaataaagcgaccctggtctgcctgatcagtgacttctatccaggtgcggtgacggtagcgtggaaagcagacagtagccctgtcaaagcaggcgtagagacgacgaccccgagtaagcagtctaataataaatacgccgcgtctagttacctgtcattgacgccggaacagtggaaaagtcatcggtcttacagctgtcaggttacccatgagggcagcacggtagaaaaaaccgttgcgccgactgagtgctcg (SEQ ID NO: 16)
[0071] The amino acid sequence (linker + light chain amino acid sequence) encoded by the nucleic acid molecule whose sequence includes the identifier molecule has been identified: GGGGSGGGGSGGGGSMAQSVLTQPPSVSEAPRQRVTISCSGSSSNIGNNTVNWYQQLPGKAPKLLIYSDDLLPSGVSDRFSGSKSGTSASLAISRLQSEDEADYYCAAWDDSLSGWVFGGGTKLTVLGQPKAAPSVTLFPPSSEELQANKATLVCLISDFYPGAVTVAWKADSSPVKAGVETTTPSKQSNNKYAASSYLSLTPEQWKSHRSYSCQVTHEGSTVEKTVAPTECS (SEQ ID NO: 17).
[0072] (2-6) Preparation of antibody gene library An antibody gene library was prepared by combining the antibody heavy chain gene library prepared in (2-4) above with the antibody light chain gene library combined with the recognition molecule prepared in (2-5) above.
[0073] The sequence containing the antibody light chain gene and the recognition molecule designed and artificially synthesized in (2-5) above was inserted into an antibody heavy chain gene library. Specifically, the DNA fragment containing the antibody light chain gene and the recognition molecule was treated with XhoI and AscI restriction enzymes and purified using a PCR Purification Kit (QIAGEN). Similarly, the antibody heavy chain gene library was treated with XhoI and AscI restriction enzymes, and vector-sized DNA fragments were excised by agarose electrophoresis. These were recovered using a GEL Extraction Kit (QIAGEN), purified by ethanol precipitation, and then dissolved in QW.
[0074] The DNA fragments containing the antibody light chain genes and the identification molecules were then ligated with the antibody heavy chain gene library vector under the following conditions: 150 μg of the DNA fragments containing the antibody light chain genes and the identification molecules and 100 μg of the heavy chain library were reacted overnight at 15°C with 10 μL of T4 Ligase (Takara). The DNA was then purified by ethanol precipitation, and dissolved in a QW flask. This process resulted in the completion of an antibody gene library consisting of nucleic acid molecule complexes composed of combinations of (A) nucleic acid molecules with identified sequences, (B) identification molecules, and (C) other nucleic acid molecules. The structures of each nucleic acid molecule complex constituting the resulting antibody library are shown schematically in Figure 7.
[0075] A 2 μL aliquot of the ligation product was suspended in 20 μL of DH12S (Thermo Fisher Scientific) and electroporated using the Gene Pulser Xcell (Bio-Rad) bacterial program (25 μF, 200 Ω, 1.8 kV, cuvette width 0.1 cm). The entire ligation product was then transformed, and the transfected E. coli was cultured in 2xYT medium containing ampicillin. The number of transformants was calculated by diluting the electroporated E. coli and plating it on agar medium, and the number was calculated to be 2 x 10 10 Transformants were obtained.
[0076] (2-7) Preparation of phage antibody library The transformed E. coli was diluted into 2 L of 2xYT medium (2xYTGA) supplemented with 0.05% glucose and 100 μg / mL ampicillin and cultured overnight at 30°C. Then, 625 mL of E. coli cultured overnight in 5 L of 2xYTGA was added, and cultured at 37°C. The absorbance at 600 nm was measured and grown until it reached 1.0.
[0077] Next, M13KO7 helper phage solution (NEW ENGLAND BioLabs) was added to the culture medium at an MOI of 10 per flask to infect the E. coli, and the cells were cultured at 37°C for 2 hours. 2xYTGA was then added to the culture medium to bring the total volume to 12 L, and the cells were cultured again overnight at 30°C.
[0078] The next day, the culture medium was centrifuged at 8000 rpm for 10 minutes, and 1 / 6 the volume of 20% polyethylene glycol / 2.5 M NaCl was added to the collected culture supernatant and stirred. After centrifugation at 8000 rpm for 20 minutes, the precipitate was suspended in PBS and used as a phage solution.
[0079] Example 3: Structural identification using a recognition molecule In this example, an experiment was conducted to create a nucleic acid complex in which the light chain (L chain) of an antibody whose structure had already been determined was used as a model for the (A) nucleic acid molecule for structural identification, to which a (B) recognition molecule composed of a base sequence was bound, and to which another (C) nucleic acid molecule encoding the heavy chain variable region (VH chain) of the antibody was further bound on the 5' side of the recognition molecule, and to which the nucleic acid complex was then used to identify the (A) nucleic acid molecule.
[0080] (3-1) Verification of structural identification using a recognition molecule Since the present invention utilizes a (B) recognition molecule to identify a nucleic acid molecule whose (A) sequence has already been identified, in this example, one of the 400 types of nucleic acid complexes prepared in Example 2 was used as a model to demonstrate that (A) sequence identification using a (B) recognition molecule and (A) sequence identification without a (B) recognition molecule (conventional method) result in the same sequence. To demonstrate this, the sequence of the model nucleic acid complex was generated by reading the (A) sequence to be identified using multiple primers according to the conventional method, and also by (A) sequence identification using a (B) recognition molecule derived from the present invention, and the sequences obtained from both were compared.
[0081] Verification was performed using the nucleotide sequence of another model antibody gene (SEQ ID NO: 18) contained in the antibody gene library containing 400 constructs prepared in Example 2 above. Specifically, plasmid DNA of one of the model antibody genes (SEQ ID NO: 19) from the nucleic acid molecule complexes prepared in Example 2 was purified from E. coli, and the nucleotide sequence was determined using a forward primer (T7 primer: taatacgactcacta taggg, SEQ ID NO: 19). The same plasmid was also ligated to a sequence decoded from the opposite strand using a reverse primer (gccagcattg acaggaggtg, SEQ ID NO: 20) (see Figure 8). The genes inserted into the antibody library were arranged in the following order from the 5' end: antibody heavy chain gene variable region, GS linker (identifier molecule), antibody light chain gene variable region, and constant region. The forward primer could decode the antibody heavy chain genes and GS linkers (identifier molecules) of the antibody library, while the reverse primer could analyze the antibody light chain genes.
[0082] On the other hand, the base sequence determined from the forward primer side above contained the sequence of the GS linker (B) as an identification molecule (denoted as identification molecule (B)002) (see Figure 8-1). Based on the sequence of this identification molecule, the base sequence of the nucleic acid molecule complex was identified by applying information on the antibody light chain gene (denoted as nucleic acid molecule (A)002), the sequence of which had been specified in advance (Figure 9).
[0083] Furthermore, the nucleotide sequence of the desired nucleic acid molecule complex, generated as a single sequence by reading with multiple primers according to the conventional method (see Figure 8-2), was compared with the nucleotide sequence of the nucleic acid molecule complex obtained by identifying the sequence of (A) nucleic acid molecule using (B) identifier molecule of the present invention (Figure 9). The results are shown in Figure 10. As a result of this sequence comparison, the nucleotide sequence of the obtained antibody light chain gene completely matched the nucleotide sequence determined by the conventional method and the nucleotide sequence of the antibody light chain assigned to the nucleotide sequence of the identifier molecule (SEQ ID NO: 18).
[0084] Based on these results, in the case of the antibody gene in this Example, the sequence of the (C) antibody heavy chain gene obtained from the forward primer was determined, and it was shown that it was possible to simultaneously determine the (B) identification molecule, and the structure of the nucleic acid molecule of the (A) antibody light chain gene, whose sequence had been previously linked from the sequence of the (B) identification molecule, could be determined, and the full-length sequence of the scFv antibody comprising the heavy chain (VH chain) and light chain (VL chain + CL chain) could be identified.
[0085] In addition, the conventional method using a reverse primer allowed us to identify the full-length sequence of the scFv antibody by overlapping the sequences obtained from the forward and reverse primers. Finally, by comparing the full-length sequences obtained from both methods, we confirmed that the identified sequences were identical (Figure 10).
[0086] The present invention provides a nucleic acid molecule complex that contains, in a single molecule, a combination of (A) a nucleic acid molecule whose sequence has been identified and (B) an identification molecule that contains a specific sequence for identifying the (A) nucleic acid molecule, and has revealed that the structure of the (A) nucleic acid molecule can be identified by simply determining the specific sequence of the (B) identification molecule within the overall structure of the nucleic acid molecule complex, without determining the structure of the (A) nucleic acid molecule.
Claims
1. A nucleic acid molecule complex comprising, in a single molecule, a combination of (A) a nucleic acid molecule whose sequence has been identified and (B) an identification molecule containing a specific sequence for identifying the (A) nucleic acid molecule.
2. The nucleic acid molecule complex according to claim 1, wherein the (B) identifying molecule is composed of a nucleic acid.
3. A nucleic acid molecule complex according to claim 1 or 2, which comprises an additional nucleic acid molecule between the (A) nucleic acid molecule and the (B) identifying molecule.
4. A nucleic acid molecule complex according to claim 1 or 2, further comprising (C) another nucleic acid molecule in a position such that the (A) nucleic acid molecule is not present between the (B) identifier molecule and the (C) other nucleic acid molecule.
5. A nucleic acid molecule library comprising a plurality of nucleic acid molecule complexes each comprising a combination of (A) a nucleic acid molecule whose sequence has been identified and (B) an identification molecule containing a specific sequence for identifying the (A) nucleic acid molecule, wherein the specific sequence of each (B) identification molecule is not the same for any other nucleic acid molecule complex.
6. The nucleic acid molecule library according to claim 5, wherein the (B) identifying molecules in each nucleic acid molecule complex contained in the library are each composed of a nucleic acid.
7. A nucleic acid molecule library according to claim 5 or 6, which comprises an additional nucleic acid molecule between the (A) nucleic acid molecule and the (B) identifying molecule in each nucleic acid molecule complex contained in the library.
8. A nucleic acid molecule library according to claim 5 or 6, wherein each nucleic acid molecule complex contained in the library further comprises (C) another nucleic acid molecule at a position such that the (A) nucleic acid molecule is not present between each (B) identifier molecule and the (C) other nucleic acid molecule.
9. A method for identifying the sequence of a nucleic acid molecule complex that contains a combination of (A) a nucleic acid molecule whose sequence has already been identified and (B) an identification molecule that contains a specific sequence for identifying the (A) nucleic acid molecule, the method comprising: a step of determining the specific sequence of the (B) identification molecule contained in the nucleic acid molecule complex; and a step of identifying the sequence of the (A) nucleic acid molecule based on the sequence of the (B) identification molecule, without determining the sequence of the (A) nucleic acid molecule.
10. A method for identifying the sequence of a (A) nucleic acid molecule according to claim 9, wherein the (B) identifying molecule is composed of a nucleic acid.
11. A method for identifying the sequence of the (A) nucleic acid molecule according to claim 9 or 10, comprising an additional nucleic acid molecule between the (A) nucleic acid molecule and the (B) identifying molecule.
12. A method for identifying the sequence of the (A) nucleic acid molecule described in claim 9 or 10, further comprising (C) another nucleic acid molecule in a position such that the (A) nucleic acid molecule is not present between the (B) identifier molecule and the (C) other nucleic acid molecule.
13. A method for identifying the sequence of a nucleic acid molecule complex containing, in a single molecule, a combination of: (A) a first nucleic acid molecule whose sequence has been identified; (B) an identification molecule comprising a specific sequence for identifying the (A) nucleic acid molecule; and (C) a second nucleic acid molecule, the second nucleic acid molecule being located at a position such that the (A) nucleic acid molecule is not present between the (B) identification molecule and the (C) second nucleic acid molecule, the method comprising: determining the sequence of the (C) second nucleic acid molecule and the sequence of the (B) identification molecule; and identifying the sequence of the (A) first nucleic acid molecule from the sequence of the (B) identification molecule without determining the sequence of the (A) first nucleic acid molecule.
14. A method for identifying the sequence of a nucleic acid molecule complex according to claim 13, wherein the (B) identifying molecule is composed of a nucleic acid.
15. A method for identifying the sequence of a nucleic acid molecule complex described in claim 13 or 14, which comprises an additional nucleic acid molecule between the (A) first nucleic acid molecule and the (B) identifier molecule.
16. A method for identifying the sequence of a nucleic acid molecule complex described in claim 9 or 10, further comprising (C') another nucleic acid molecule in a position such that the (A) nucleic acid molecule is not present between the (B) identifier molecule or (C) second nucleic acid molecule and the (C') other nucleic acid molecule.
Citation Information
Patent Citations
Methods and reagents for detecting and assessing genotoxicity - Patent Application 20070122999
JP2021513364A
Polynucleotides, reagents, and methods for nucleic acid hybridization
JP2021526366A
Method and system for characterizing tumors and identifying tumor heterogeneity - Patent Application 20070122997
JP2022522221A