DNA synthesis method

The use of template DNA with universal and stop base sequences in an aqueous environment addresses DNA synthesis limitations, enhancing speed, reducing errors, and lowering costs for efficient long-length DNA production.

WO2026038781A1PCT designated stage Publication Date: 2026-02-19KOREA UNIV RES & BUSINESS FOUND
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/011724
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-16
Filing Date
2025-08-05
Publication Date
2026-02-19

AI Technical Summary

Technical Problem

Existing DNA synthesis technologies face challenges in terms of speed, cost, environmental pollution, error rate, and synthesis of long-length DNA, particularly using microarray methods, which are limited by density and solvent use.

Method used

A method involving template DNA with universal and stop base sequences, using DNA polymerase in an aqueous environment, allows for selective and stepwise DNA base insertion, reducing errors and synthesis costs, and enabling long-length DNA synthesis.

Benefits of technology

The method significantly improves DNA synthesis rate and reduces error rates, eliminates environmental pollution, and lowers synthesis costs, facilitating efficient and accurate long-length DNA production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025011724_19022026_PF_FP_ABST
    Figure KR2025011724_19022026_PF_FP_ABST
Patent Text Reader

Abstract

A DNA synthesis method is provided. The DNA synthesis method according to some embodiments is a method for synthesizing a target DNA to be used as a data storage medium, and may comprise the steps of: preparing a template DNA; and synthesizing the target DNA from the template DNA on the basis of a DNA polymerase. Here, the sequence of the template DNA includes a universal base sequence, which can complementarily bind to any one of two or more types of data DNA bases, and a stop base sequence, wherein a starting base of the stop base sequence does not complementarily bind to the data DNA bases and can complementarily bind to a pre-designated non-data DNA base. According to this method, a target DNA having a desired sequence can be accurately and quickly synthesized in an eco-friendly manner.
Need to check novelty before this filing date? Find Prior Art

Description

DNA synthesis method

[0001] This application claims priority to Korean Patent Application No. 10-2024-0109768, filed August 16, 2024, the entire disclosure of which is incorporated herein by reference.

[0002] The present disclosure relates to a DNA (deoxyribonucleic acid) synthesis technology and a DNA-based data storage technology.

[0003] Due to the increasing use of smartphones, social networking services (SNS), and Internet of Things (IoT) devices, data production is exploding every year. Consequently, active efforts are being made to develop new storage media that can overcome the limitations of existing data storage media.

[0004] Recently, deoxyribonucleic acid (DNA) has attracted considerable attention as a new data storage medium due to its advantages, such as high data density, ease of long-term storage, and easy data replication. As illustrated in Figure 1, actually storing data (11) in DNA requires several steps, including data encoding (12), DNA synthesis (13), and DNA storage (14). However, among these, the DNA synthesis step (13) can act as a bottleneck in terms of synthesis cost and speed.

[0005] Most DNA synthesis technologies proposed to date are based on phosphoramidite chemistry and utilize microarrays to accelerate DNA synthesis. Microarray DNA synthesis technologies also include inkjet printers, electrochemistry, and photochemistry.

[0006] However, all of the above microarray DNA synthesis technologies fail to guarantee sufficient DNA synthesis rates, and technological limitations have already been reached in increasing the density of microarrays, which is directly related to DNA synthesis speed. Furthermore, all of these microarray DNA synthesis technologies require high synthesis costs and suffer from frequent errors during the DNA synthesis process. Furthermore, all of these microarray DNA synthesis technologies also suffer from the problem of environmental pollution due to the use of organic solvents.

[0007] The technical problems to be solved through some embodiments of the present disclosure relate to a method for synthesizing DNA (deoxyribonucleic acid) used as a data storage medium.

[0008] Specifically, a technical problem to be solved through some embodiments of the present disclosure is to provide a method capable of improving the speed of DNA synthesis.

[0009] In addition, another technical problem to be solved through some embodiments of the present disclosure is to provide a method for reducing the cost required for DNA synthesis.

[0010] In addition, another technical challenge to be solved through some embodiments of the present disclosure is to provide a method for easily synthesizing long-length DNA (e.g., more than 1 kb).

[0011] In addition, another technical challenge to be solved through some embodiments of the present disclosure is to provide a method for synthesizing DNA in an environmentally friendly manner.

[0012] In addition, another technical problem to be solved through some embodiments of the present disclosure is to provide a method capable of improving the DNA synthesis error rate and DNA synthesis success rate.

[0013] In addition, another technical problem to be solved through some embodiments of the present disclosure is to provide structure and sequence information of template DNA that can be used for DNA synthesis.

[0014] In addition, another technical challenge to be solved through some embodiments of the present disclosure is to provide a method for storing various data in DNA.

[0015] The technical problems of the present disclosure are not limited to the technical problems mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art of the present disclosure from the description below.

[0016] According to some embodiments of the present disclosure for solving the above-described technical problem, a DNA synthesis method is provided, which is a method for synthesizing a target DNA used as a data storage medium, comprising: a step of preparing a template DNA; and a step of synthesizing the target DNA from the template DNA based on a DNA polymerase. In this case, the sequence of the template DNA includes: a universal base sequence capable of complementarily binding to any one of two or more data DNA bases; and a stop base sequence, wherein the start base of the stop base sequence does not complementarily bind to the data DNA bases but can complementarily bind to a pre-designated non-data DNA base.

[0017] In some embodiments, the termination base sequence within the template DNA may be located adjacent to the universal base sequence in the 5' end direction of the template DNA.

[0018] In some embodiments, the universal base sequence and the termination base sequence may alternately appear within the template DNA.

[0019] In some embodiments, the data DNA bases may have a mapping relationship with values ​​representing data stored in the target DNA, and the non-data DNA bases may not have a mapping relationship with the values ​​representing the data.

[0020] In some embodiments, the universal base sequence may include inosine (I), and the starting base of the termination base sequence may be cytosine (C).

[0021] In some embodiments, the universal base sequence may include uracil (U) and the starting base of the termination base sequence may be guanine (G).

[0022] In some embodiments, the sequence of the template DNA may further include a coding sequence for binding a primer, the sequence being located at the 3'-terminal segment of the template DNA.

[0023] In some embodiments, the sequence of the template DNA may further include an address code sequence used to identify the template DNA.

[0024] In some embodiments, the sequence of the template DNA includes a first universal base sequence and a first termination base sequence, and the step of synthesizing the target DNA may include: encoding specific data into a DNA base sequence including a first data DNA base; introducing a first solution including the first data DNA base to form a complementary bond between the first universal base sequence and the first data DNA base; and removing the first solution and introducing a second solution including the non-data DNA base to form a complementary bond between the start base of the first termination base sequence and the non-data DNA base.

[0025] In some embodiments, the sequence of the template DNA further comprises a second universal base sequence, the first termination base sequence is located between the first universal base sequence and the second universal base sequence, and the DNA base sequence further comprises a second data DNA base, wherein the step of synthesizing the target DNA may further comprise a step of removing the second solution and introducing a third solution containing the second data DNA base to form a complementary bond between the second universal base sequence and the second data DNA base.

[0026] In some embodiments, the types of the data DNA bases are three, and the encoding step may include: converting the specific data into ternary data; and converting the ternary data into the DNA base sequence based on a mapping relationship between the ternary number and the data DNA bases.

[0027] In some embodiments, the types of the data DNA bases are two, the specific data is binary data, and the encoding step may include a step of converting the specific data into the DNA base sequence based on a mapping relationship between the binary number and the data DNA bases.

[0028] In some embodiments, the step of synthesizing the target DNA may include: preparing the template DNA in a circular structure; hybridizing an auxiliary DNA having a base sequence complementary to at least a portion of the template DNA in the circular structure with the template DNA in the circular structure; performing a first round of synthesis starting from an end of the auxiliary DNA and referring to the template DNA in the circular structure, thereby synthesizing a first DNA fragment of the target DNA; and performing a second round of synthesis again referring to the template DNA in the circular structure, thereby synthesizing a second DNA fragment of the target DNA continuously to the first DNA fragment.

[0029] In some embodiments, the sequence of the template DNA further comprises: an error accumulation prevention code sequence used to prevent accumulation of synthesis errors related to the target DNA, the error accumulation prevention code sequence starting with a DNA base that does not complementarily bind to an error accumulation prevention DNA base that is at least one of the data DNA bases and does not complementarily bind to the non-data DNA base, and the step of synthesizing the first DNA fragment of the target DNA may include a step of introducing a solution containing the error accumulation prevention DNA base and the non-data DNA base when the synthesis process for the central segment of the template DNA where the universal base sequence and the termination base sequence are located is completed.

[0030] In some embodiments, when the universal base sequence comprises inosine (I) and the termination base sequence comprises cytosine (C): the error accumulation prevention DNA base comprises adenine (A), the non-data DNA base is guanine (G), and the starting base of the error accumulation prevention code sequence may be adenine (A).

[0031] In some embodiments, when the universal base sequence comprises uracil (U) and the termination base sequence comprises guanine (G): the error accumulation prevention DNA base comprises adenine (A), the non-data DNA base is cytosine (C), and the starting base of the error accumulation prevention code sequence may be adenine (A).

[0032] According to some embodiments of the present disclosure for solving the above-described technical problem, a template DNA sequence for synthesizing a target DNA used as a data storage medium based on DNA polymerase, wherein the template DNA sequence may include: a universal base sequence capable of complementarily binding to any one of two or more data DNA bases; and a stop base sequence. In this case, the start base of the stop base sequence may not complementarily bind to the data DNA bases but may complementarily bind to a pre-designated non-data DNA base.

[0033] In some embodiments, the data DNA bases are: (1) when the data DNA bases are adenine (A) and thymine (T), the starting base of the termination base sequence is guanine (G) or cytosine (C), (2) when the data DNA bases are adenine (A) and guanine (G), the starting base of the termination base sequence is adenine (A) or guanine (G), (3) when the data DNA bases are adenine (A) and cytosine (C), the starting base of the termination base sequence is adenine (A) or cytosine (C), (4) when the data DNA bases are thymine (T) and guanine (G), the starting base of the termination base sequence is thymine (T) or guanine (G), (5) when the data DNA bases are thymine (T) and cytosine (C), the starting base of the termination base sequence is thymine (T) or cytosine (C), and (6) when the data DNA bases are guanine (G) and cytosine (C). In this case, the starting base of the termination base sequence may be adenine (A) or thymine (T).

[0034] In some embodiments, when the data DNA bases are: (1) adenine (A), cytosine (C), and thymine (T), the starting base of the termination base sequence is cytosine (C), (2) adenine (A), guanine (G), and thymine (T), the starting base of the termination base sequence is guanine (G), (3) cytosine (C), guanine (G), and thymine (T), the starting base of the termination base sequence is thymine (T), and (4) adenine (A), cytosine (C), and guanine (G), the starting base of the termination base sequence may be adenine (A).

[0035] According to some embodiments of the present disclosure, target DNA (deoxyribonucleic acid) can be synthesized from template DNA using DNA polymerase in an aqueous environment. This not only eliminates environmental pollution issues, but also significantly improves the DNA synthesis rate. Furthermore, the error rate in DNA synthesis can be significantly reduced.

[0036] Additionally, template DNA can support the selective and stepwise (sequential) insertion of DNA bases via a universal base sequence and a stop base sequence. In this case, target DNA with the desired sequence can be easily and accurately synthesized, allowing the target DNA to function as a data storage medium.

[0037] Additionally, target DNA can be synthesized through a repetitive synthesis process over multiple rounds based on a circular DNA template. In this case, long target DNA can be easily synthesized, significantly reducing the overall cost of DNA synthesis.

[0038] Furthermore, by utilizing the coding sequence located at the 5' end of the template DNA, synthesis errors (e.g., deletion errors) occurring in a specific round can be reset (i.e., the accumulation of synthesis errors is prevented). In this case, synthesis errors in a specific round can be prevented from influencing other rounds (i.e., the problem of having to discard the entire synthesized DNA is resolved), thereby further reducing the overall cost of DNA synthesis. Furthermore, both the error rate and the success rate of DNA synthesis can be significantly improved.

[0039] The effects according to the technical idea of ​​the present disclosure are not limited to the effects mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art from the description below.

[0040] Figure 1 is an exemplary drawing to explain the problems of existing DNA (deoxyribonucleic acid) synthesis technology.

[0041] FIG. 2 is an exemplary drawing illustrating the structure and sequence of template DNA according to some embodiments of the present disclosure.

[0042] FIG. 3 is an exemplary drawing to further explain the structure and sequence of template DNA according to some embodiments of the present disclosure.

[0043] FIG. 4 is an exemplary flowchart illustrating a DNA synthesis method according to some embodiments of the present disclosure.

[0044] Figure 5 is an exemplary drawing for explaining the detailed process of the target DNA synthesis step illustrated in Figure 4.

[0045] FIG. 6 is an exemplary flowchart illustrating a DNA synthesis method according to some other embodiments of the present disclosure.

[0046] Figure 7 is an exemplary drawing to further explain the detailed process of the DNA production step of the circular structure illustrated in Figure 6.

[0047] Figure 8 is an exemplary drawing to further explain the detailed process of the target DNA synthesis step illustrated in Figure 6.

[0048] FIG. 9 is an exemplary diagram illustrating a code sequence of template DNA that can be used in a synthetic error reset method according to some embodiments of the present disclosure.

[0049] FIG. 10 is an exemplary diagram illustrating a synthetic error reset method according to some embodiments of the present disclosure.

[0050] FIG. 11 is an exemplary diagram illustrating the effect of a synthetic error reset method according to some embodiments of the present disclosure.

[0051] FIG. 12 is an exemplary flowchart illustrating a DNA-based data storage method according to some embodiments of the present disclosure.

[0052] Figure 13 is an exemplary drawing to further explain the detailed process of the encoding step illustrated in Figure 12.

[0053] Figure 14 shows the results of an experiment conducted by the inventors on inosine (I).

[0054] Figure 15 shows the results of an experiment conducted to demonstrate the effectiveness of a synthetic error (e.g., deletion error) reset method.

[0055] Hereinafter, various embodiments of the present disclosure will be described in detail with reference to the attached drawings. The advantages and features of the present disclosure, and methods for achieving them, will become clear with reference to the embodiments described in detail below together with the attached drawings. However, the technical idea of ​​the present disclosure is not limited to the following embodiments and may be implemented in various different forms. The following embodiments are provided only to complete the technical idea of ​​the present disclosure and to fully inform those skilled in the art of the present disclosure of the scope of the present disclosure, and the technical idea of ​​the present disclosure is defined only by the scope of the claims.

[0056] In describing various embodiments of the present disclosure, if it is determined that a detailed description of a related known configuration or function may obscure the gist of the present disclosure, the detailed description will be omitted.

[0057] Unless otherwise defined, the terms (including technical and scientific terms) used in the following examples may be used with meanings commonly understood by those of ordinary skill in the art to which this disclosure pertains; however, this may vary depending on the intentions of engineers working in the relevant field, precedents, the emergence of new technologies, etc. The terminology used in this disclosure is for the purpose of describing the embodiments and is not intended to limit the scope of this disclosure.

[0058] In the following examples, singular expressions include plural concepts unless the context clearly specifies that they are singular. Furthermore, plural expressions include singular concepts unless the context clearly specifies that they are plural.

[0059] In addition, terms such as first, second, A, B, (a), (b), etc. used in the following embodiments are only used to distinguish certain components from other components, and the nature, order, or sequence of the components are not limited by the terms.

[0060] Hereinafter, various embodiments of the present disclosure will be described in detail.

[0061] According to various embodiments of the present disclosure, a method for synthesizing DNA (deoxyribonucleic acid) that supports selective insertion of DNA bases is provided. Specifically, a method is provided for synthesizing target DNA having a desired sequence (i.e., DNA base sequence) by selectively inserting DNA bases using template DNA designed / manufactured by the inventors of the present invention. This method can be used, for example, to store various data in the target DNA (i.e., the target DNA can be used as a data storage medium), and examples of this will be described later with reference to FIGS. 12 and 13. In the following description, the term "DNA base sequence" may also be abbreviated as "base sequence" or "sequence." Furthermore, the term "DNA" may be used interchangeably with terms such as "DNA molecule" or "DNA strand."

[0062] First, for ease of understanding, we will describe template DNA designed / manufactured to support selective insertion of DNA bases.

[0063] FIG. 2 is an exemplary drawing for explaining the structure and sequence of template DNA (21) according to some embodiments of the present disclosure.

[0064] As shown in Fig. 2, the template DNA (21) can be composed of, for example, a 3' terminal segment (22), a 5' terminal segment (24), and a central segment (23).

[0065] The 3' terminal segment (22) is a segment located at the 3' end of the template DNA (21) and includes one or more code sequences (see 'CODE 1'). Here, the code sequence refers to a base sequence (or base sequence pattern) having a specific function. For example, as illustrated in FIG. 3, the 3' terminal segment (22) may include a code sequence for binding to a forward primer (31). Those skilled in the art will already be familiar with the function of the primer (31) indicating the starting point of DNA synthesis, and therefore, a description thereof will be omitted. FIGS. 2 and 3 illustrate exemplary code sequences (25) of the 3' terminal segment (22). In some cases, the 3' terminal segment (22) may further include a unique address code sequence that can be used to identify the template DNA (21). The address code sequence can be used, for example, to identify and select a desired template DNA (e.g., 21) from among a plurality of template DNAs.

[0066] Next, the 5' terminal segment (24) is a segment located at the 5' end of the template DNA (21) and may include one or more code sequences (see 'CODE 2'). For example, the 5' terminal segment (24) may include a code sequence for resetting an error (e.g., deletion error) that occurs during DNA synthesis, which will be described in detail later with reference to FIGS. 9 to 11. FIGS. 2 and 3 illustrate exemplary code sequences (26) of the 5' terminal segment (24). In some cases, the 5' terminal segment (24) may further include a unique address sequence that can be used to identify the template DNA (21).

[0067] Next, the central segment (23) is a segment located between the 3'-terminal segment (22) and the 5'-terminal segment (24) and supports selective and stepwise (sequential) insertion of DNA bases. For this purpose, the central segment (23) includes a universal base sequence (27) and a stop base sequence (28). In the drawings below Fig. 2, the 'universal base sequence' and the 'stop base sequence' are abbreviated as 'UBS' and 'SBS', respectively.

[0068] The universal base sequence (27) is a sequence that supports selective insertion into DNA bases. The universal base sequence (27) is composed of one or more universal bases that can complementarily bind to any one of two or more DNA bases, thereby supporting selective insertion into DNA bases. Examples of such universal bases include inosine (I) and uracil (U), and the universal base sequence (27) can be composed of, for example, one or more inosine (I) or uracil (U). However, the scope of the present disclosure is not limited thereto.

[0069] For reference, since inosine (I) can complementarily bind to adenine (A), guanine (G), and thymine (T), inosine (I) can be used to selectively insert them into target DNA (see 32 in Fig. 3). Fig. 3 illustrates an example where the universal base sequence (27) consists of one inosine (I).

[0070] Next, the terminator sequence (28) is a sequence for supporting stepwise (sequential) insertion of DNA bases without errors. Specifically, a DNA base (so-called 'terminator') that does not complementarily bind with DNA bases that complementarily bind with the universal base sequence (27) but complementarily binds with a specific DNA base designated in advance can be used (designated) as the starting base and the last base of the terminator sequence (28) (e.g., a DNA base that has a low binding efficiency (or binding force) with the complementary bases of the universal base sequence (27) but a high binding efficiency with a specific DNA base is used as the starting base and the last base of the terminator sequence (28)). In addition, the terminator sequence (28) is configured to include one or more of these terminator bases and is positioned adjacent to the universal base sequence (27) in the 5'-end direction of the template DNA (21), thereby supporting stepwise insertion of DNA bases without errors. Since the terminator has the characteristic of not binding to the complementary bases of the universal base sequence (27), the specific form of the terminator sequence (28) may vary depending on the universal base sequence (27) (or its complementary bases). For various examples of the complementary bases of the universal base sequence (27) and the start / last base of the terminator sequence (28), please refer to Tables 1 and 2 below. For better understanding, the function of the terminator sequence (28) will be further explained with reference to FIG. 3.

[0071] FIG. 3 illustrates an example in which the universal base sequence (27) is composed of one inosine (I) and the termination base sequence (28) is composed of one cytosine (C) (i.e., the starting base and the last base are cytosine (C)). However, the scope of the present disclosure is not limited thereto. For example, unlike as illustrated in FIG. 3, the termination base sequence (28) may be composed of a plurality of cytosines (C), a plurality of guanines (G), etc., and a general DNA base other than the termination base may be located in the middle of the termination base sequence (28) (e.g., the termination base sequence (28) may be composed of CC, CCC, CCCC, CAC, CAAC, CCAC, CCAAC, etc.).

[0072] Referring to FIG. 3, the termination base sequence (28, cytosine (C)) does not complementarily bind to the complementary bases of the universal base sequence (27) (i.e., adenine (A), guanine (G), and thymine (T)). Therefore, when any one of adenine (A), guanine (G), and thymine (T) is inserted, the inserted DNA base (32) forms an exact complementary bond only with the universal base sequence (27, inosine (I)) (i.e., only the desired DNA base (32) is inserted into the target DNA).

[0073] Next, let us assume that guanine (33, G), which is the complementary base of the terminating base sequence (28, cytosine (C)), is inserted. In this case, guanine (33, G) does not complementarily bond with the complementary bases (i.e., adenine (A), guanine (G), and thymine (T)) of the next universal base sequence (i.e., inosine (I)), and therefore forms an exact complementary bond only with the terminating base sequence (28, cytosine (C)). Therefore, guanine (33, G) can function as an indicator base indicating that the insertion step of the preceding DNA base (32) has been terminated (completed).

[0074] For reference, in FIG. 3, etc., the fact that the DNA base (eg, 32) combined with the universal base sequence (eg, 27) is indicated as 'data' can be understood to mean that in the DNA-based data storage method, the DNA base (32) that is selectively inserted represents the 'data' itself that is stored in the target DNA (i.e., the selective insertion of the DNA base means writing data to the target DNA). In addition, it can be understood that the reason why the DNA base (eg, 33) combined with the start / last base (i.e., the termination base) of the termination base sequence (eg, 28) is marked as 'non-data' is because, in the DNA-based data storage method, the complementary base (eg, guanine (G)) of the termination base (eg, 28) does not represent the data itself stored in the target DNA (i.e., the complementary base of the termination base (eg, 28) or the complementary sequence of the termination base sequence (eg, 28) plays a role in distinguishing between data writing in one step and data writing in the next step).

[0075] In the following description, the complementary base of a universal base (or universal base sequence) may be named as a 'data base' or a 'data DNA base', and the complementary base of a terminator base (or terminator base sequence) may be named as a 'non-data base' or a 'non-data DNA base'.

[0076] This is explained again with reference to Figure 2.

[0077] The universal base sequence (27) and the terminator base sequence (28) may appear (position) alternately within the central segment (23) of the template DNA (21). In other words, pairs of the universal base sequence (27) and the terminator base sequence (28) may appear repeatedly within the central segment (23) so that a target DNA of a certain length or longer can be easily synthesized. However, in some cases, two or more terminator base sequences (e.g., 28) may appear consecutively adjacent to one universal base sequence (27). The length of the central segment (23) and the template DNA (21) (e.g., the number of pairs of the universal base sequence (27) and the terminator base sequence (28)) may be designed in various ways in consideration of the length of the target DNA, the synthesis cost, etc.

[0078] Tables 1 and 2 below show various examples of complementary bases of the universal base sequence (27) (or universal base) and the start / last base (i.e., the terminal base) of the terminal base sequence (28). Table 1 shows an example where there are two complementary bases of the universal base sequence (27), and Table 2 shows an example where there are three complementary bases of the universal base sequence (27).

[0079] Complementary base of universal base sequenceStart / last base of terminal base sequenceAdenine (A), Thymine (T)Guanine (G) or Cytosine (C)Adenine (A), Guanine (G)Adenine (A) or Guanine (G)Adenine (A), Cytosine (C)Adenine (A) or Cytosine (C)Thymine (T), Guanine (G)Thymine (T) or Guanine (G)Thymine (T), Cytosine (C)Thymine (T) or Cytosine (C)Guanine (G), Cytosine (C)Adenine (A) or Thymine (T)

[0080] Complementary bases of universal base sequenceStart / last base of terminal base sequenceAdenine (A), cytosine (C), thymine (T)Cytosine (C)Adenine (A), guanine (G), thymine (T)Guanine (G)Cytosine (C), guanine (G), thymine (T)Thymine (T)Adenine (A), cytosine (C), guanine (G)Adenine (A)

[0081] As shown in Table 2, when the complementary bases of the universal base sequence (27) are adenine (A), cytosine (C), and thymine (T) (e.g., when the universal base sequence (27) is composed of one inosine (I)), cytosine (C) can be used as the start / last base of the termination base sequence (28). This is because cytosine (C) does not complementarily bind with adenine (A), cytosine (C), and thymine (T), but only complementarily binds with guanine (G), which is a specific DNA base designated in advance (in this case, it can be understood that adenine (A), cytosine (C), and thymine (T) are used as data DNA bases, and guanine (C) is used as a non-data DNA base). The structure and sequence of the template DNA (21) according to some embodiments of the present disclosure have been described with reference to FIGS. 2 and 3 so far. Below, a method for synthesizing target DNA using the template DNA (e.g., 21) described above will be described.

[0082] FIG. 4 is an exemplary diagram illustrating a DNA synthesis method according to some embodiments of the present disclosure. However, this is merely an exemplary embodiment for achieving the purpose of the present disclosure, and it is understood that some steps may be added or deleted as needed.

[0083] As illustrated in FIG. 4, the present embodiments may begin with step S41 of preparing template DNA. For example, the template DNA (21) illustrated in FIG. 2 may be prepared (manufactured). At this time, the universal base sequence (27) of the template DNA (21) may be composed of one inosine (I) as illustrated in FIG. 3, and the termination base sequence (28) may be composed of one cytosine (C), but the scope of the present disclosure is not limited thereto. For example, the universal base sequence (27) may be composed of multiple DNA bases including one inosine (I), and the termination base sequence (28) may be composed of multiple DNA bases and have cytosine (C) as the starting base and the last base. Alternatively, the universal base sequence (27) may be composed of uracil (U) (i.e., the universal base is uracil (U)), and the termination base sequence (28) may have guanine (G) as the start / last base (i.e., the termination base is guanine (G)).

[0084] In step S42, target DNA is synthesized from template DNA using DNA polymerase. That is, similar to DNA synthesis within cells, target DNA is synthesized using an enzyme (i.e., DNA polymerase) in an aqueous environment (i.e., target DNA is synthesized using DNA polymerase after a primer is bound to the template DNA). Specifically, target DNA can be synthesized by repeatedly performing the steps of introducing a DNA base selected from among complementary bases of a universal base sequence in the form of a solution and the steps of introducing a complementary base of a terminal base sequence in the form of a solution. In this case, target DNA having a desired sequence can be synthesized quickly and accurately without environmental pollution issues (because no organic solvent is used). Those skilled in the art are likely already familiar with the advantages of enzyme-based DNA synthesis (e.g., fast synthesis speed, low error rate, etc.) and the reasons for them, so a description thereof will be omitted.

[0085] For reference, the synthesis process in step S42 can be performed in an aqueous environment with the template DNA immobilized on beads. However, the scope of the present disclosure is not limited to this. For example, the synthesis process can also be performed with the template DNA immobilized on a flat glass slide (e.g., a glass sectioned into a microarray format).

[0086] To provide better understanding, the detailed process of step S42 will be described with reference to Fig. 5.

[0087] Fig. 5 is an exemplary diagram showing the detailed process of step S42. Fig. 5 assumes that the universal base sequence (e.g., 52) is composed of one inosine (I) (i.e., the universal base is inosine (I)) and the terminator base sequence (e.g., 53) is composed of one cytosine (C) (i.e., the terminator used as the start / last base is cytosine (C)). In addition, Fig. 5 illustrates two universal base sequences (52 and 54, hereinafter referred to as the 'first universal base sequence' and the 'second universal base sequence') and terminator base sequences (53 and 55, hereinafter referred to as the 'first terminator base sequence' and the 'second terminator base sequence') among the central segments of the template DNA (51).

[0088] As illustrated in FIG. 5, when a solution containing cytosine (57, C) (hereinafter referred to as the 'first solution') is introduced, a complementary bond is formed between the first universal base sequence (52, inosine (I)) of the template DNA (51) and cytosine (57, C), and as a result, cytosine (57, C) is inserted into the target DNA (56). At this time, the first termination base sequence (53, cytosine (C)) of the template DNA (51) (more precisely, the starting base of the first termination base sequence (53)) does not complementarily bind to cytosine (57, C) (see 'X' mark), so only cytosine (57, C) is accurately inserted into the target DNA (56) as intended.

[0089] Next, when the first solution is removed and a solution containing guanine (58, G) (hereinafter referred to as the 'second solution') is introduced, a complementary bond is formed between the first terminal base sequence (53, cytosine (C)) of the template DNA (51) and guanine (C), and as a result, guanine (58, C) is inserted into the target DNA (56). At this time, since the second universal base sequence (54, inosine (I)) of the template DNA (51) does not complementarily bind to guanine (58, G) (see 'X' mark), only guanine (58, G), which indicates the end of the insertion step of cytosine (57, C), is accurately inserted into the target DNA (56).

[0090] By repeatedly performing the above steps, a target DNA (56) having a desired sequence can be synthesized. For example, as illustrated, when the second solution is removed and a solution containing thymine (59, T) is further added, a complementary bond is formed between the second universal base sequence (54, inosine (I)) of the template DNA (51) and thymine (59, T), and as a result, thymine (59, T) is inserted into the target DNA (56). At this time, the second termination base sequence (55, cytosine (C)) of the template DNA (51) (more precisely, the starting base of the second termination base sequence (55)) does not complementarily bind to thymine (59, T) (see 'X' mark), so that only thymine (59, T) is accurately inserted into the target DNA (56) as intended.

[0091] Hereinafter, DNA synthesis methods according to some embodiments of the present disclosure have been described with reference to FIGS. 4 and 5. As described above, target DNA can be synthesized from template DNA using DNA polymerase in an aqueous environment. This not only eliminates environmental pollution issues, but also significantly improves the DNA synthesis rate. Furthermore, the error rate of DNA synthesis can be significantly reduced.

[0092] Additionally, template DNA can support the selective and stepwise (sequential) insertion of DNA bases via a universal base sequence and a termination base sequence. In this case, target DNA with the desired sequence can be easily and accurately synthesized, allowing the target DNA to function as a data storage medium.

[0093] Hereinafter, a DNA synthesis method according to several other embodiments of the present disclosure will be described with reference to FIGS. 6 to 11. However, for the clarity of the present disclosure, descriptions of content overlapping with the previous embodiments will be omitted.

[0094] Unlike the previous examples that utilized linear template DNA, these examples relate to a method for synthesizing target DNA using circular template DNA. Using circular template DNA facilitates the synthesis of long target DNAs, and offers various advantages, such as reduced synthesis costs and reduced error rates. These will be described in detail later.

[0095] Figure 6 is an exemplary flowchart illustrating a DNA synthesis method according to several other embodiments of the present disclosure. However, this is merely an exemplary embodiment for achieving the objectives of the present disclosure, and it is to be understood that some steps may be added or deleted as needed.

[0096] As illustrated in FIG. 6, the present embodiments may begin with step S61 of preparing (manufacturing) template DNA and auxiliary DNA. Here, the auxiliary DNA is used to form the template DNA into a circular structure, and refers to DNA having a base sequence complementary to at least a portion of the template DNA.

[0097] For example, as illustrated in FIG. 7, let us assume that the template DNA (71) is configured to include a 3'-terminal segment (72), a central segment (73), and a 5'-terminal segment (74) (for reference, the template DNA (71) of FIG. 7 may correspond to the template DNA (21) of FIG. 2). In this case, the auxiliary DNA (75) may be configured to include a first segment (76, i.e., a segment having a base sequence complementary to the 3'-terminal segment (72)) that can complementarily bind to the 3'-terminal segment (72) of the template DNA (71), and a second segment (77) that can complementarily bind to the 5'-terminal segment (74) of the template DNA (71). In some cases, the auxiliary DNA (75) may further include a third segment and may further include biotin (78) (or another substance) that serves to immobilize the DNA on the bead. In the drawings below FIG. 7, a segment (eg, 76) of auxiliary DNA (75) and a segment (eg, 72) of complementary template DNA (71) are indicated with the same hatching.

[0098] This is explained again with reference to Figure 6.

[0099] In step S62, template DNA and auxiliary DNA are hybridized to produce a circular DNA structure (hereinafter referred to as "combined DNA"). For better understanding, this description is again provided with reference to Figure 7.

[0100] Figure 7 is an exemplary drawing for explaining the detailed process of step S62.

[0101] As illustrated in Fig. 7, first, template DNA (71) and auxiliary DNA (75) are combined. That is, the 3'-terminal segment (72) of template DNA (71) is combined with the second segment (77) of complementary auxiliary DNA (75), and the 5'-terminal segment (74) of template DNA (71) is combined with the first segment (76) of complementary auxiliary DNA (75). For the convenience of understanding, Fig. 7 also illustrates exemplary base sequences of related segments (72, 74, 76, 77).

[0102] Next, the 3'-terminal segment (72) and the 5'-terminal segment (74) of the template DNA (71) are linked to each other through a ligation process. As a result, the template DNA (71) and the combined DNA (79) form a circular structure.

[0103] For reference, the part where synthesis begins in the combined DNA (79) may be a double-stranded segment (76), and it may be understood that the segment (76) or segments (76, 77) serve as a primer.

[0104] Meanwhile, in some other embodiments of the present disclosure, circular binding DNA may be generated in a manner different from that illustrated in FIG. 7. For example, rather than generating circular binding DNA through a ligation process, the circular binding DNA may be generated by preparing template DNA in a circular form and hybridizing auxiliary DNA (i.e., DNA having a base sequence complementary to at least a portion of the template DNA) with the circular template DNA. Thus, the specific method for generating circular binding DNA may vary in many ways.

[0105] This is explained again with reference to Figure 6.

[0106] In step S63, the target DNA is synthesized by repeatedly performing the synthesis process with reference to the binding DNA. For example, as illustrated in FIG. 8, while the binding DNA (79) is immobilized on a magnetic bead (81) (or other type of bead) using biotin (78), the circular template DNA (71) is referenced (reused) and multiple rounds of the synthesis process can be repeatedly performed. For example, starting from the end of the auxiliary DNA (75), the first round of the synthesis process is performed with reference to the circular template DNA (71), thereby synthesizing the first DNA fragment of the target DNA, and then the second round (e.g., the next round of the first round) of the synthesis process is performed with reference to the template DNA (71) again, thereby synthesizing the second DNA fragment of the target DNA (e.g., the next DNA fragment of the target DNA that is continuous to the first DNA fragment). In this case, even a long target DNA can be easily synthesized. In addition, since there is no need to prepare (manufacture) long-length template DNA to synthesize long-length target DNA, the overall cost of DNA synthesis can be greatly reduced.

[0107] Meanwhile, according to some embodiments of the present disclosure, when repeatedly performing a synthesis process using a circularly structured binding DNA (e.g., 79), a process of periodically or aperiodically resetting a synthesis error can be performed using the code sequence of the template DNA (e.g., 71). By doing so, a synthesis error (e.g., deletion error) occurring in a specific round can be prevented from affecting other rounds (i.e., the accumulation of synthesis errors can be prevented). These embodiments will be described in detail below with reference to FIGS. 9 to 11.

[0108] FIG. 9 is an exemplary diagram illustrating a code sequence (96) that can be used in a synthetic error reset method according to some embodiments of the present disclosure. FIG. 9 illustrates a case where the universal base sequence (94) is composed of one inosine (I) and the termination base sequence (95) is composed of one cytosine (C). The template DNA (91) of FIG. 9 may correspond to the template DNA (21) of FIG. 2.

[0109] As illustrated in FIG. 9, the template DNA (91) is configured to include a 3' terminal segment (91), a central segment (92), and a 5' terminal segment (93), and a code sequence (96) for resetting (i.e., preventing accumulation) a synthesis error may be included in the 5' terminal segment (93). For example, the code sequence (96) may be positioned at the beginning of the 5' terminal segment (93) to initiate a synthesis error resetting process simultaneously with the completion of a DNA synthesis process based on the central segment (92), but the scope of the present disclosure is not limited thereto.

[0110] The code sequence (96) is a code sequence that resets (i.e., prevents accumulation) a synthetic error, and therefore may be named as an 'error reset code sequence', an 'error reset sequence', an 'error accumulation prevention code sequence', or an 'error accumulation prevention sequence' depending on the case.

[0111] The code sequence (96) may be a sequence starting with a DNA base that does not complementarily bind to at least one DNA base (so-called 'error reset DNA base' or 'error accumulation prevention DNA base') among the complementary bases (i.e., data DNA bases) of the universal base sequence (94) and does not complementarily bind to the complementary base (i.e., non-data DNA base) of the start / last base (i.e., the termination base) of the termination base sequence (95). This is because if the code sequence (96) is designed in this way, errors (e.g., deletion errors) that occur during the synthesis of the target DNA can be accurately reset.

[0112] For example, as illustrated, if the universal base sequence (94) and the terminator base sequence (95) are each composed of one inosine (I) and one cytosine (C), the code sequence (96) may be a sequence starting with adenine (A). This is because adenine (A) does not complementarily bind to at least one complementary base (e.g., adenine (A), which is a data DNA base and an error accumulation prevention DNA base) of the universal base sequence (94, inosine (I)) and does not complementarily bind to the complementary base (i.e., guanine (G), which is a non-data DNA base) of the terminator base sequence (95, cytosine (C)). For the reason why the code sequence (96) is designed in this way, please refer to the description of FIG. 10. FIG. 9 illustrates a case where the code sequence (96) is composed of adenine (A) and guanine (G) as an example.

[0113] As another example, unlike the illustrated one, suppose that the universal base sequence (e.g., 94) consists of a single uracil (U) and the terminal base sequence (e.g., 95) consists of a single guanine (G). In this case, the code sequence (96) may be a sequence starting with adenine (A). This is because adenine (A) does not complementarily bind to at least one complementary base of uracil (U) (e.g., adenine (A), which is a data DNA base and an error-prevention DNA base) and also does not complementarily bind to cytosine (C), which is a complementary base of guanine (G) (i.e., a non-data DNA base).

[0114] FIG. 10 is an exemplary diagram illustrating a synthesis error reset method according to some embodiments of the present disclosure. FIG. 10 assumes that target DNA is synthesized from the template DNA (91) of FIG. 9 according to the DNA synthesis method illustrated in FIG. 6. In the following description, for ease of understanding, the terms "first" and "second" are used to distinguish between two base sequences of the same type (e.g., 102-1 and 102-2, 103-1 and 103-2).

[0115] As shown in the upper part of Fig. 10, let us assume that a specific round of synthesis process is initiated based on a template DNA (91) to synthesize a DNA fragment (101) of a target DNA, and that the synthesis of a central segment (92) of the template DNA (91) is completed (see the first universal base sequence (102-1), the first terminator base sequence (103-1), and the DNA bases (105, 106) combined therewith). In addition, let us assume that a deletion error has occurred in the second universal base sequence (102-2, inosine (I)) and the second terminator base sequence (103-2, cytosine (C)) (in this case, if the deletion error is not reset in the round, the DNA bases of the next round are combined with the second universal base sequence (102-2) and the second terminator base sequence (103-2), which causes a problem that the entire target DNA must be discarded).

[0116] In the above case, in order to reset the deletion error (or prevent the deletion error from accumulating), a solution containing a DNA base that prevents error accumulation (e.g., adenine (A)) and a complementary base (i.e., guanine (G)) of a termination base sequence (e.g., 103-1, cytosine (C)) may be injected. This is because when this solution is injected, all deletion errors related to the universal base sequence (e.g., 102-2) and the termination base sequence (e.g., 103-2) can be accurately reset. In other words, deletion errors occurring in the current round can be prevented from affecting the next round in advance. At this time, the injected error accumulation prevention DNA base (e.g., adenine (A)) may be designated as a DNA base that does not complementarily bind to the starting base (104, adenine (A)) of the code sequence (96) for error reset.

[0117] For example, as shown in the middle part of FIG. 10, when a solution containing adenine (A) and guanine (G) is injected, a complementary bond is formed between adenine (A) and the second universal base sequence (102-2, inosine (I)), and a complementary bond is formed between guanine (G) and the second termination base sequence (103-2, cytosine (C)). At this time, since neither the injected adenine (A) nor the guanine (G) complementarily binds to the starting base (104, adenine (A)) of the code sequence (96), only the deletion error associated with the middle segment (92) of the template DNA (91) is accurately reset. Consequently, through this process, the deletion error occurring in the current round can be prevented from affecting other rounds in advance (i.e., if the deletion error does not occur again in other rounds, the DNA fragments generated in other rounds can be used normally).

[0118] Meanwhile, the DNA fragment (101) in which the deletion error is reset is not a DNA fragment having the desired sequence, and thus can be discarded (e.g., if formation of a complementary bond is detected after inputting adenine (A) and guanine (G), it is determined that a deletion error has occurred, and the DNA fragment (101) of the current round can be discarded). Then, as illustrated in the lower part of Fig. 10, by inputting a solution containing a complementary base (109, thymine (T)) of the starting base (104, adenine (A)), the synthesis process of a new DNA fragment (i.e., the synthesis process of the next round) can be started (initiated).

[0119] Figure 11 is an exemplary drawing for explaining the effect according to the synthetic error reset method described above.

[0120] As shown on the left side of Fig. 11, in order to synthesize a 1 kb (kilobase) target DNA (112), at least 1 kb of template DNA (111) is required, and if an error (e.g., deletion error) occurs even once during the synthesis process, the entire target DNA (112) must be discarded (because a single synthesis error affects the entire sequence of the target DNA (112)). Furthermore, as the length of the target DNA (112) increases, the frequency of occurrence of synthesis errors inevitably increases. Therefore, as the length of the target DNA (112) increases, the cost required for preparing (manufacturing) the template DNA (111) increases and the synthesis success rate decreases significantly.

[0121] On the other hand, as illustrated on the right side of Fig. 11, by using a circular structure of binding DNA (113) (or template DNA), a 1 kb target DNA (114) can be synthesized from a much shorter template DNA. For example, a 1 kb target DNA (114) can be synthesized from a 100 bp (base pair) template DNA through approximately 10 rounds of repeated synthesis. In addition, even if an error (e.g., deletion error) occurs during the synthesis of a specific round, only the DNA fragment (115) of that round is affected by the error, and the DNA fragments (e.g., 116) of other rounds are not affected by the error. Therefore, the cost of preparing (manufacturing) template DNA can be significantly reduced, and a low synthesis error rate and a high synthesis success rate can be guaranteed regardless of the length of the target DNA (114).

[0122] Hereinafter, DNA synthesis methods according to several other embodiments of the present disclosure have been described with reference to FIGS. 6 through 11. As described above, target DNA can be synthesized through a repeated synthesis process over multiple rounds based on a circular template DNA (or conjugated DNA). In this case, long target DNA can be easily synthesized, and the overall cost of DNA synthesis can be significantly reduced.

[0123] Furthermore, by utilizing the coding sequence located at the 5' end of the template DNA, synthesis errors (e.g., deletion errors) occurring in a specific round can be reset (i.e., the accumulation of synthesis errors can be prevented). In this case, synthesis errors in a specific round can be prevented from influencing other rounds (i.e., the problem of having to discard the entire synthesized DNA is resolved), thereby further reducing the overall cost of DNA synthesis. Furthermore, both the error rate and the success rate of DNA synthesis can be significantly improved.

[0124] The DNA synthesis methods described with reference to FIGS. 4 through 11 can be used to store various data in target DNA. Below, these DNA-based data storage methods will be described with reference to FIGS. 12 and 13.

[0125] Figure 12 is an exemplary flowchart illustrating a DNA-based data storage method according to some embodiments of the present disclosure. However, this is merely an exemplary embodiment for achieving the objectives of the present disclosure, and it is understood that some steps may be added or deleted as needed.

[0126] As illustrated in FIG. 12, first, in step S121, data (i.e., digital data) to be stored in the target DNA is acquired. Here, the acquired data may be, for example, binary data, and the type of data may be diverse, such as text, images, audio, video, etc.

[0127] In step S122, the acquired data is encoded into a DNA base sequence. For example, the acquired data may be converted into a DNA base sequence according to a predefined encoding rule. Here, the encoding rule may be defined to convert a specific value of the data into any one of the data DNA bases based on a mapping relationship between a value representing the data (e.g., binary, ternary, etc.) and data DNA bases (i.e., complementary bases of a universal base sequence) (provided that non-data DNA bases (i.e., complementary bases of the start / last base of the termination base sequence) do not have a mapping relationship with a value representing the data). However, the scope of the present disclosure is not limited thereto. For better understanding, step S122 will be described in more detail with reference to FIG. 13.

[0128] Figure 13 is an exemplary diagram to further explain the detailed process of step S122. Figure 13 assumes that the universal base sequence is composed of one inosine (I) and the complementary bases of the universal base sequence (i.e., the data DNA bases adenine (A), cytosine (C), and thymine (T)) have a mapping relationship with the ternary number.

[0129] As shown in Fig. 13, let us assume that binary data (131) has been obtained. In this case, the binary data (131) can be converted into ternary data (132), and the ternary data (132) can be converted into a DNA sequence (133) based on the mapping relationship between data DNA bases and ternary numbers.

[0130] For example, let us assume that the mapping relationship is defined so that the ternary numbers '0', '1', and '2' are mapped to cytosine (C), adenine (A), and thymine (T), respectively. Then, let us assume that the binary number '1013' (see 'subscript') included in the binary data (131) is converted to the ternary number '102'. In this case, the ternary number '102' can be converted to the DNA base sequence 'ACT' based on the predefined mapping relationship.

[0131] For reference, if there are two types of data DNA bases, binary data (131) can be converted into a DNA base sequence based on the mapping relationship between data DNA bases and binary numbers.

[0132] This is explained again with reference to Figure 12.

[0133] In step S123, target DNA is synthesized based on the DNA base sequence. For example, the target DNA may be synthesized according to the DNA synthesis method illustrated in FIG. 4, or according to the DNA synthesis method illustrated in FIG. 6.

[0134] For reference, the process of reading (retrieving) data stored in DNA can be performed through steps such as DNA sequencing and data decoding.

[0135] Heretofore, DNA-based data storage methods according to some embodiments of the present disclosure have been described with reference to FIGS. 12 and 13 . As described above, by utilizing DNA as a data storage medium, limitations of existing data storage media (e.g., low data density, difficulty in long-term storage, etc.) can be easily overcome.

[0136] Below, we will briefly introduce the results of the experiments conducted by the inventors.

[0137] [Experiment preparation]

[0138] For the experiment, first, a 60-mer-long template DNA (100 pmole / μl, 1 μl), splint DNA to be used for ligation (100 pmole / μl, 3 μl), and primer DNA with biotinylated 5' ends (100 pmole / μl, 1 μl) were prepared. The template DNA and splint DNA were mixed and hybridized, and then T4 DNA ligase and its buffer were added to perform a ligation reaction to circularize the template DNA. Subsequently, T4 DNA ligase was inactivated, and the splint DNA was isolated. Afterwards, linear DNA residues were removed by treatment with Exonuclease I, and Exonuclease I was also inactivated.

[0139] Next, the circular template DNA and the 5' biotin-incorporated primer DNA were mixed and hybridized. The hybridized DNA was bound to Invitrogen Dynabeads MyOne Streptavidin magnetic beads, allowing the 5' biotin of the primer to bind to the streptavidin beads. Afterwards, Bst 2.0 DNA polymerase was bound to the DNA-bead complex, completing the preparation for DNA synthesis.

[0140] During the synthesis process, the DNA solution was added to the monomer solution and reacted at 65°C for 1 minute. After synthesis, the solution was washed three times with 1x isothermal amplification buffer to maintain enzyme activity. The synthesis step was repeated as many times as necessary.

[0141] The final synthesized DNA was denatured at 95°C for 5 minutes and then separated from magnetic beads. The concentration of the separated DNA was measured using a Nanodrop device. To confirm DNA quality, bands were observed using polyacrylamide gel electrophoresis.

[0142] [Experimental Example 1]

[0143] The present inventors prepared a template DNA (e.g., 21) of the structure illustrated in Fig. 3 and conducted an experiment in which cytosine (C), adenine (A), thymine (T), and guanine (G) were each added to a universal base sequence consisting of one inosine (I). The experiment was conducted in an aqueous environment using DNA polymerase. The experimental results are illustrated in Fig. 14.

[0144] Figure 14 shows the experimental results analyzed through gel electrophoresis.

[0145] As shown in Figure 14, cytosine (C), adenine (A), and thymine (T) bind to inosine (I) within 1 minute (note the extension to a '21-mer'), but guanine (G) does not bind to inosine (I) even after a long time (note the persistence of a '20-mer'). This shows that inosine (I) can function as a universal base (or universal base sequence) that supports selective insertion into three DNA bases (i.e., cytosine (C), adenine (A), and thymine (T)), and that cytosine (C) can function as the start / last base of a termination base sequence.

[0146] [Experimental Example 2]

[0147] The present inventors conducted an experiment to demonstrate the effectiveness of a method for resetting synthetic errors (e.g., deletion errors) according to FIG. 10. To this end, a universal base sequence was created as in Experimental Example 1 above, but adenine (A) was added to the middle of the sequence as a code sequence for resetting errors that occur during DNA synthesis.

[0148] As shown in Figure 15, the experimental results confirmed that when cytosine (C) and guanine (G) were added, complementary DNA was synthesized only up to A of the template DNA. This confirmed that it is possible to prevent deletion errors from accumulating during long DNA synthesis.

[0149] Various embodiments of the present disclosure and effects according to the embodiments have been described with reference to FIGS. 1 through 14 so far. The effects according to the technical concept of the present disclosure are not limited to the effects mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art from the description below.

[0150] Furthermore, even though the above embodiments have described multiple components as being combined or operating in combination, the technical concept of the present disclosure is not necessarily limited to these embodiments. That is, within the scope of the technical concept of the present disclosure, all of the components may be selectively combined and operated one or more times.

[0151] Although various embodiments of the present disclosure have been described with reference to the attached drawings, those skilled in the art will appreciate that the technical concepts of the present disclosure can be implemented in other specific forms without changing the technical concepts or essential features thereof. Therefore, it should be understood that the embodiments described above are exemplary in all respects and not restrictive. The scope of protection of the present disclosure should be interpreted by the claims below, and all technical concepts within a scope equivalent thereto should be interpreted as being included within the scope of the technical concepts defined by the present disclosure.

Claims

1. A method for synthesizing target DNA used as a data storage medium, Step of preparing template DNA; and Comprising a step of synthesizing the target DNA from the template DNA based on DNA polymerase, The sequence of the above template DNA is: A universal base sequence capable of complementarily binding to any one of two or more data DNA bases; and Contains a stop base sequence, The starting base of the above termination base sequence is capable of complementarily binding with a pre-designated non-data DNA base without complementarily binding with the data DNA bases. DNA synthesis method.

2. In paragraph 1, In the template DNA, the termination base sequence is located adjacent to the universal base sequence in the 5' end direction of the template DNA. DNA synthesis method.

3. In paragraph 1, In the above template DNA, the universal base sequence and the termination base sequence appear alternately. DNA synthesis method.

4. In paragraph 1, The above data DNA bases have a mapping relationship with values ​​representing data stored in the target DNA, The above non-data DNA bases do not have a mapping relationship with the values ​​representing the data, DNA synthesis method.

5. In paragraph 1, The above universal base sequence includes inosine (I), The starting base of the above termination base sequence is cytosine (C). DNA synthesis method.

6. In paragraph 1, The above universal base sequence includes uracil (U), The starting base of the above termination base sequence is guanine (G). DNA synthesis method.

7. In paragraph 1, The sequence of the above template DNA is: Located at the 3' terminal segment of the above template DNA and further comprising a code sequence for binding a primer, DNA synthesis method.

8. In paragraph 1, The sequence of the above template DNA is: Further comprising an address code sequence used to identify the above template DNA, DNA synthesis method.

9. In paragraph 1, The sequence of the above template DNA includes a first universal base sequence and a first termination base sequence, The step of synthesizing the above target DNA is: A step of encoding specific data into a DNA base sequence including a first data DNA base; A step of forming a complementary bond between the first universal base sequence and the first data DNA base by introducing a first solution containing the first data DNA base; and A step of removing the first solution and introducing a second solution containing the non-data DNA base to form a complementary bond between the starting base of the first termination base sequence and the non-data DNA base, DNA synthesis method.

10. In paragraph 9, The sequence of the above template DNA further comprises a second universal base sequence, The first termination base sequence is located between the first universal base sequence and the second universal base sequence, The above DNA base sequence further comprises a second data DNA base, The step of synthesizing the above target DNA is: Further comprising a step of removing the second solution and introducing a third solution containing the second data DNA base to form a complementary bond between the second universal base sequence and the second data DNA base. DNA synthesis method.

11. In paragraph 9, There are three types of DNA bases in the above data: The above encoding step is, A step of converting the above specific data into ternary data; and A step of converting the ternary data into the DNA base sequence based on the mapping relationship between the ternary number and the data DNA bases, DNA synthesis method.

12. In paragraph 9, There are two types of DNA bases in the above data, The above specific data is binary data, The above encoding step is, A step of converting the specific data into the DNA base sequence based on the mapping relationship between the binary number and the data DNA bases, DNA synthesis method.

13. In paragraph 1, The step of synthesizing the above target DNA is: A step of preparing the above template DNA into a circular structure; A step of hybridizing an auxiliary DNA having a base sequence complementary to at least a part of the template DNA of the circular structure with the template DNA of the circular structure; A step of synthesizing a first DNA fragment of the target DNA by performing a first round of synthesis process with reference to the template DNA of the circular structure, starting from the end of the auxiliary DNA; and A step of synthesizing a second DNA fragment of the target DNA consecutively to the first DNA fragment by performing a second round of synthesis process again referring to the template DNA of the circular structure. DNA synthesis method.

14. In paragraph 13, The sequence of the above template DNA is: Further comprising an error accumulation prevention code sequence used to prevent the accumulation of synthetic errors related to the target DNA, The above error accumulation prevention code sequence begins with a DNA base that does not complementarily bind to at least one error accumulation prevention DNA base among the data DNA bases and does not complementarily bind to the non-data DNA base, The step of synthesizing the first DNA fragment of the target DNA is: When the synthesis process for the central segment of the template DNA in which the universal base sequence and the termination base sequence are located is completed, a step of introducing a solution containing the error accumulation prevention DNA base and the non-data DNA base is included. DNA synthesis method.

15. In paragraph 14, When the above universal base sequence includes inosine (I) and the above termination base sequence includes cytosine (C): The above error accumulation prevention DNA base includes adenine (A), The above non-data DNA base is guanine (G), The starting base of the above error accumulation prevention code sequence is adenine (A). DNA synthesis method.

16. In paragraph 14, When the universal base sequence comprises uracil (U) and the termination base sequence comprises guanine (G): The above error accumulation prevention DNA base includes adenine (A), The above non-data DNA base is cytosine (C), The starting base of the above error accumulation prevention code sequence is adenine (A). DNA synthesis method.

17. In the sequence of template DNA for synthesizing target DNA used as a data storage medium based on DNA polymerase, The sequence of the above template DNA is: A universal base sequence capable of complementarily binding to any one of two or more data DNA bases; and Contains a stop base sequence, The starting base of the above termination base sequence is capable of complementarily binding with a pre-designated non-data DNA base without complementarily binding with the data DNA bases. Template DNA sequence.

18. In paragraph 17, The above data DNA bases are: (1) In the case of adenine (A) and thymine (T), the starting base of the termination base sequence is guanine (G) or cytosine (C), (2) In the case of adenine (A) and guanine (G), the starting base of the termination base sequence is adenine (A) or guanine (G). (3) In the case of adenine (A) and cytosine (C), the starting base of the termination base sequence is adenine (A) or cytosine (C), (4) In the case of thymine (T) and guanine (G), the starting base of the termination base sequence is thymine (T) or guanine (G). (5) In the case of thymine (T) and cytosine (C), the starting base of the termination base sequence is thymine (T) or cytosine (C), (6) In the case of guanine (G) and cytosine (C), the starting base of the termination base sequence is adenine (A) or thymine (T). Template DNA sequence.

19. In paragraph 17, The above data DNA bases are: (1) In the case of adenine (A), cytosine (C) and thymine (T), the starting base of the termination base sequence is cytosine (C), (2) In the case of adenine (A), guanine (G) and thymine (T), the starting base of the termination base sequence is guanine (G). (3) In the case of cytosine (C), guanine (G) and thymine (T), the starting base of the termination base sequence is thymine (T), (4) In the case of adenine (A), cytosine (C) and guanine (G), the starting base of the termination base sequence is adenine (A). Template DNA sequence.

20. In paragraph 17, In the above template DNA, the universal base sequence and the termination base sequence appear alternately. Template DNA sequence.

21. In paragraph 17, The sequence of the above template DNA is: Located at the 3' terminal segment of the above template DNA and further comprising a code sequence for binding a primer, Template DNA sequence.

22. In paragraph 17, The sequence of the above template DNA is: Further comprising an address code sequence used to identify the above template DNA, Template DNA sequence.

23. In paragraph 17, The sequence of the above template DNA is: Further comprising an error accumulation prevention code sequence located at the 5' terminal segment of the central segment of the template DNA and used to prevent the accumulation of synthesis errors related to the target DNA, The above error accumulation prevention code sequence starts with a DNA base that does not complementarily bind to at least one of the data DNA bases and does not complementarily bind to the non-data DNA base, Template DNA sequence.

Citation Information

Patent Citations

  • Methods of storing information using nucleic acids

    KR1020150037824A

  • Information processing apparatus, storage medium, and analysis method

    KR1020240001039A

  • KR20190116297A

  • KR20200025430A

  • KR20220052995A