Method for encoding digital data onto nucleic acids using biological processes
The nucleic acid-based data storage method converts digital data into nucleotide bioblocks for efficient and sustainable storage, overcoming the limitations of current digital data storage technologies.
Patent Information
- Application Number
- JP2024568588
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-05-19
- Publication Date
- 2025-06-17
AI Technical Summary
Current digital data storage methods are inefficient and unsustainable due to the fragility, energy consumption, and limited durability of existing media, which cannot keep pace with exponentially increasing data generation.
A nucleic acid-based data storage method that converts digital data into bioblocks consisting of nucleotides, which are then assembled into larger components for storage, utilizing libraries of data storage nucleic acid molecules with cleavage sites for assembly and retrieval.
This method provides a durable, energy-efficient, and scalable means of data storage that can maintain large amounts of data over time, addressing the limitations of current technologies.
Smart Images

Figure 2025518553000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a nucleic acid-based data storage method for storing digital information.
Background Art
[0002] Storing and archiving digital data is a major challenge in our modern society. Current digital media stored in data centers are fragile, bulky, and consume a large amount of energy. Optical media, magnetic tapes, hard drives, and flash memories have been developed, but their durability generally does not exceed 10 years on average. These data must be regularly copied onto new, reliable media and maintained under controlled temperature and humidity, resulting in astronomical energy costs and requiring large amounts of raw materials. The amount of energy consumed by data centers corresponds to 2% of the world's total electricity consumption (Masanet et al., 2020). The carbon footprint of data centers exceeds that of commercial aviation worldwide. Despite their increasing energy costs, carbon dioxide emissions, and the need for large areas, data centers can only store 30% of the data we generate while our data generation volume is increasing exponentially. "Even if we can store about 30% of the information we generate today, within just 10 or 12 years, we will only be able to store about 3%" (Dr. Karin Strauss, Microsoft Research, 2018). These general considerations indicate that the data revolution, the big data market, and the development of artificial intelligence cannot progress further without finding an innovative solution to the problem of data storage.
[0003] U.S. Patent Application Publication No. 2018 / 0137418 describes the use of chemically generated DNA bricks, some of which (3 to 6) are assembled to create larger molecules (hundreds of base pairs) encoding information bits (0 or 1). However, these processes require a great deal of time and cost.
[0004] As a result, there is still a need for new means to maintain the encoding of large amounts of data and, furthermore, for biocompatibility, i.e., for storing digital data that can be copied, edited, written, and / or read using organisms.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Summary of the Invention
[0006] The present invention is a nucleic acid-based data storage method for storing information, comprising: a) recovering data in the form of a digital sequence formed from a plurality of bits, each bit having a value of 0 or 1; b) subdividing the digital sequence into n digital subsequences, each comprising m bits, where m is in the range from 2 to 16; c) converting each of the n digital subsequences into a bioblock consisting of a sequence of m nucleotides, where the digital subsequence resides in m bits assigned to positions from 0 to m - 1, and the conversion of the digital subsequence to the bioblock - converting bits at even positions to a first nucleotide N1 when the bit has a value of 0 and to a second distinct nucleotide N2 when the bit has a value of 1, and - converting bits at odd positions to a third nucleotide N3 when the bit has a value of 0 and to a fourth distinct nucleotide N4 when the bit has a value of 1, where - N1, N2, N3, and N4 are distinct nucleotides, and d) constructing a plurality of x components, each individual component of the plurality of x components including at least one bioblock, the x components together including n bioblocks, e) assembling the plurality of x components together in one or more steps in a fixed order, related to a nucleic acid-based data storage method for storing information including
[0007] In some embodiments, the nucleotide is selected from the group of natural nucleotides consisting of adenine, guanine, cytosine, uracil, and thymine or non-natural nucleotides.
[0008] In some embodiments, the x components are x DNA molecules, preferably x double-stranded DNA molecules.
[0009] In some embodiments, in step (d), the construction of the plurality of x components, each including at least one bioblock, - selectively capturing x data storage nucleic acid molecules from at least one library of data storage nucleic acid molecules, each data storage nucleic acid molecule including at least one bioblock surrounded by a region including a cleavage site, - includes cleaving each of the x data storage nucleic acid molecules, thereby releasing at least one bioblock.
[0010] In some embodiments, in step (d), the construction of a plurality of x components, each comprising at least one bioblock, comprises: - selectively capturing n data storage nucleic acid molecules from at least two libraries of data storage nucleic acid molecules, each data storage nucleic acid molecule of each library comprising one bioblock surrounded by a region comprising a cleavage site, each library comprising all possible bioblocks of m nucleotides; - cleaving each of the n data storage nucleic acid molecules, thereby releasing n bioblocks.
[0011] In some embodiments, the region comprising the cleavage site comprises from 2 to 25 nucleotides.
[0012] In some embodiments, the region surrounding each bioblock comprises a site for a restriction enzyme, and step (d) comprises digesting each of the x data storage nucleic acid molecules with one or two restriction enzymes.
[0013] In some embodiments, step (e) comprises one or more assembly steps using overlap extension polymerase chain reaction (PCR), polymerase cycling assembly, sticky end ligation, biobrick assembly, golden gate assembly, Gibson assembly, recombinase assembly, ligase cycling reaction, template-directed ligation, in vivo assembly, or any other DNA assembly protocol.
[0014] The present invention further relates to a data storage nucleic acid molecule comprising at least one bioblock, the bioblock consisting of a nucleic acid sequence of m nucleotides assigned positions from 0 to m-1, - the bioblock being formed from at least two and at most four distinct nucleotides; - The nucleotides at even positions can be selected from the first and second nucleotides, and the nucleotides at odd positions can be selected from the third and fourth nucleotides, and the first, second, third, and fourth nucleotides are distinctly different.
[0015] In some embodiments, the data storage nucleic acid molecule is a double-stranded molecule, preferably a DNA molecule.
[0016] In some embodiments, the data storage nucleic acid molecule is a plasmid, cosmid, fosmid, prokaryotic chromosome, or eukaryotic chromosome.
[0017] In some embodiments, each bioblock is surrounded by a region containing a cleavage site, preferably surrounded by two sites for one restriction enzyme.
[0018] In some embodiments, the data storage nucleic acid molecule is replicable.
[0019] The present invention further relates to a library comprising a plurality of data storage nucleic acid molecules according to the present invention, each of the data storage nucleic acid molecules of the library contains one bioblock, each data storage nucleic acid molecule of the library contains the same peripheral region containing a cleavage site, and the library contains all possible bioblocks of m nucleotides.
[0020] The present invention further relates to a nucleic acid-based data storage system comprising at least two libraries according to the present invention.
[0021] Definitions In the present invention, the following terms have the following meanings.
[0022] The term "digital data" refers to data that can be managed by a computerized machine. As used herein, the expression "digital data" is meant to refer to data represented in binary. As used herein, "binary" refers to a language composed of the bits "0" and "1". Non-limiting examples of digital data include program files, text files, music files, image files, video files, and combinations thereof.
[0023] The phrase "store" or "storing" refers to the act of placing an item in a particular location for future use or safekeeping. More specifically, the expression "storage of digital data" is intended to mean the act of securely storing digital information for future use.
[0024] The term "replicable" refers to the ability to be replicated in vivo by a polymerase such as, for example, DNA polymerase, i.e., the ability to be replicated accurately within the error range of the organism's replication machinery. As used herein, "replicable nucleic acid molecule" is intended to refer to a nucleic acid molecule that can be replicated at least once in vivo. In one embodiment, the nucleic acid molecule according to the invention is selected from the group consisting of plasmids, cosmids, and chromosomes. In practice, a replicable nucleic acid molecule contains one or more origins of replication (also called ORIs), or one or more centromeres (for chromosomes).
[0025] Within the scope of the present invention, the terms "nucleotide" and "nucleobase" have meanings as alternative terms to each other and are intended to refer to the nucleic acid building blocks of DNA or RNA molecules. Nucleotides include both natural nucleotides and unnatural nucleotides. As used herein, natural nucleotides refer to adenine (A) or guanine (G) of purines, or cytosine (C), thymine (T), or uracil (U) of pyrimidines. For DNA nucleic acids, A refers to dAMP deoxyribonucleotide, G refers to dGMP deoxyribonucleotide, C refers to dCMP deoxyribonucleotide, and T refers to dTMP deoxyribonucleotide. For RNA nucleic acids, A refers to AMP ribonucleotide, G refers to GMP ribonucleotide, C refers to CMP ribonucleotide, and U refers to UMP ribonucleotide. As used herein, the term "unnatural nucleotide" refers to chemically modified A, T, U, C, or G nucleotides. Non-limiting examples of unnatural nucleotides are 2-amino-ATP, 8-aza-ATP, 2'-fluoro-dATP, 2'-fluoro-dCTP, 2'-fluoro-dGTP, 2'-fluoro-dUTP, 5-iodo-CTP, 5-iodo-UTP, N6-methyl-ATP, 5-methyl-CTP, 2'-O-methyl-ATP, 2'-O-methyl-CTP, 2'-O-methyl-GTP, 2'-O-methyl-UTP, pseudo-UTP, ITP, 2'-O-methyl-ITP, puromycin-TP, xanthosine-TP, 5-methyl-UTP, 4-thio-UTP, 2'-amino-dCTP, 2'-amino-dUTP, 2'-azido-dCTP, 2'-azido-dUTP, 06-methyl-GTP, 2-thio-UTP, Ara-CTP, Ara-UTP, 5,6-dihydro-UTP, 2-thio-CTP, 6-aza-CTP, 6-aza-UTP, N1-methyl-GTP, 2'-O-methyl-2-amino-ATP, 2'-O-methylpseudo-UTP, N1-methyl-ATP, 2'-O-methyl-5-methyl-UTP, 7-deaza-GTP, 2'-azido-dATP, 2'-amino-dATP, Ara-ATP, 8-azido-ATP, 5-bromo-CTP, 5-bromo-UTP, 2'-fluoro-dTTP,3'-O-methyl-ATP, 3'-O-methyl-CTP, 3'-O-methyl-GTP, 3'-O-methyl-UTP, 7-deaza-ATP, 5-AA-UTP, 2'-azido-dGTP, 2'-amino-dGTP, 5-AA-CTP, 8-oxo-GTP, pseudo-iso-CTP, N4-methyl-CTP, N1-methylpseudo-UTP, 5,6-dihydro-5-methyl-UTP, N6-methylamino-ATP, 5-carboxy-CTP, 5-formyl-CTP, 5-hydroxymethyl-UTP, 5-hydroxymethyl-CTP, thieno-GTP, 5-hydroxy-CTP, 5-formyl-UTP, thieno-UTP, 2-amino-dATP, 5-bromo-dCTP, 5-bromo-dUTP, 7-deaza-dATP, 7-deaza-dGTP, dITP, 5-propynyl-dCTP, 5-propynyl-dUTP, 2'-dUTP, 5-fluoro-dUTP, 5-iodo-dCTP, 5-iodo-dUTP, N6-methyl-dATP, 5-methyl-dCTP, O6-methyl-dGTP, N2-methyl-dGTP, 8-oxo-dATP, 8-oxo-dGTP, 2-thio-dTTP, 2'-dPTP, 5-hydroxy-dCTP, 4-thio-dTTP, 2-thio-dCTP, 6-aza-dUTP, 6-thio-dGTP, 8-chloro-dATP, 5-AA-dCTP, 5-AA-dUTP, N4-methyl-dCTP, 2'-deoxyzebularine-TP, 5-hydroxymethyl-dUTP, 5-hydroxymethyl-dCTP, 5-propynylamino-dCTP, 5-propynylamino-dUTP, 5-carboxy-dCTP, 5-formyl-dCTP, 5-indolyl-AA-dUTP, 5-carboxy-dUTP, 5-formyl-dUTP, 3'-dATP, 3'-dGTP, 3'-dCTP, 5-methyl-3'-dUTP, 3'-dUTP, ddATP, ddGTP, ddUTP, ddTTP, ddCTP, 3'-azido-ddATP, 3'-azido-ddGTP, 3'-azido-ddTTP, 3'-amino-ddATP, 3'-amino-ddCTP, 3'-amino-ddGTP, 3'-amino-ddTTP, 3'-azido-ddCTP, 3'-azido-ddUTP, 5-bromo-ddUTP, ddITP, (1-thio)-dATP, (1-thio)-dCTP, (1-thio)-dGTP, (1-thio)-dTTP(1-thio)-ATP, (1-thio)-CTP, (1-thio)-GTP, (1-thio)-UTP, (1-thio)-ddATP, (1-thio)-ddCTP, (1-thio)-ddGTP, (1-thio)-ddTTP, (1-thio)-3'-azido-ddTTP, (1-thio)-ddUTP, (1-borano)-dATP, (1-borano)-dCTP, (1-borano)-dGTP, (1-borano)-dTTP, ganciclovir-TP, cidofovir-DP, 3-methyl-6-amino-5-(1'-b-D-2'-deoxyribofuranosyl)-pyrimidin-2-one, 6-amino-9[(1'-b-D-2'-deoxyribofuranosyl)-4-hydroxy-5-(hydroxymethyl)-oxolan-2-yl]-1H-purin-2-one, 6-amino-3-(1'-b-D-2'-deoxyribofuranosyl)-5-nitro-1H-pyridin-2-one and 2-amino-8-(1'-b-D-2'-deoxyribofuranosyl)-imidazo[1,2a]-1,3,5-triazine-[8H]-4-one are included.,
[0026] Detailed description The present invention is a nucleic acid-based data storage method for storing information, (a) restoring data in the form of a digital sequence formed from a plurality of bits each having a value of 0 or 1, (b) subdividing the digital sequence into n digital subsequences each containing m bits, where m is in the range from 2 to 16, (c) converting each of the n digital subsequences into a bioblock consisting of a sequence of m nucleotides, the digital subsequence resides in m bits assigned to positions from 0 to m-1, the conversion of the digital subsequence into a bioblock is - converting the bits at even positions into a first nucleotide N1 when the bit has a value of 0 and into a second distinct nucleotide N2 when the bit has a value of 1, and - converting bits in odd positions to a third nucleotide N3 when the bit has a value of 0 and to a fourth distinct nucleotide N4 when the bit has a value of 1, consists in, - N1, N2, N3, and N4 are distinct nucleotides, converting, (d) constructing a plurality of x components, each individual component of the plurality of x components including at least one bioblock, and the x components together including n bioblocks, constructing, (e) assembling the plurality of x components together, in one or more steps, in a fixed order, a nucleic acid-based data storage method for storing information including.
[0027] As used herein, the term "bit" (binary digit) refers to the smallest basic unit of digital information. In practice, a bit depends on binary notation and can have a value of either 0 or 1. Methods of storing bits involve the use of electronic devices and are well known in the art.
[0028] Within the scope of the present invention, the term "byte", which is interchangeable with the terms "bit string" or "bit chain", refers to a contiguous sequence of bits and is also referred to herein as a "digital subsequence". Within the scope of the present invention, the number of bits per byte corresponds to the value of m.
[0029] In one embodiment, the value of m is in the range from 2 to 16. As used herein, the phrase "from 2 to 16" means 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, and 16. In one embodiment, the value of m is selected from the group consisting of or including 2, 4, 6, 8, 10, 12, 14, and 16. In one embodiment, the value of m is selected from the group consisting of or including 2, 4, 8, and 16.
[0030] In one embodiment, the value of m is 8. In fact, a byte consisting of 8 bits is referred to as an octet in this specification, and a bio-block resulting from the conversion of an octet is referred to as a bio-octet in this specification.
[0031] In one embodiment, the value of m is 16. In one embodiment, the value of m is 4. In one embodiment, the value of m is 2.
[0032] In one embodiment, the digital sequence may be included in or consist of any digital file stored on a computer. In one embodiment, the file may be a .3dm (Rhino 3D Model), .3ds (3D Studio Scene), .3g2 (3GPP2 multimedia file), .3gp (3GPP multimedia file), .accdb (Access 2007 database file), .ai (Adobe Illustrator file), .aif (AIF / Audio Interchange audio file), .apk (Android package file), .asp and .aspx (Active Server Page files), .avi (Audio Video Interleave file), .bak (backup file), .bat (batch file), .bin (binary file), .bmp (bitmap image file), .cab (Windows cabinet file), .cda (CD audio track file), .cer (Internet security certificate), .cfg (configuration file), .cfm (ColdFusion markup file), .cgi (Common Gateway Interface script), .cgi or .pl (Perl script file), .com (MS-DOS command file), .cpl (Windows control panel file), .css (Cascading Style Sheet file), .csv (comma-separated values file), .cur (Windows cursor file), .dat (data file), .db or .dbf (database file), .dll (DLL file), .dmp (dump file), .doc and .docx (Microsoft Word files), .drv (device driver file), .exe (executable file), .flv (Adobe Flash Video file), .gif (GIF / Graphical Interchange Format image), .h264 (H.264 video file), .htm and .html (HTML / Hypertext Markup Language files), .icns (macOS X icon resource file),.ico (icon file),.ico (icon file),.iff (Interchange File Format),.ini (initialization file),.jar (Java Archive file),.jpeg or.jpg (JPEG image),.js (JavaScript file),.jsp (Java Server Page file),.key (Keynote presentation),.lnk (Windows shortcut file),.log (log file),.m4v (Apple MP4 video file),.max (3ds Max Scene file),.mdb (Microsoft Access database file),.mid or.midi (MIDI audio file),.mkv (Matroska Multimedia Container),.mov (Apple QuickTime movie file),.mp3 (MP3 audio file),.mp4 (MPEG-4 video file),.mpa (MPEG-2 audio file),.mpg or.mpeg (MPEG video file),.msg (Outlook mail message),.msi (Windows installer package),.obj (Wavefront 3D Object file),.odp (OpenOffice Impress presentation file),.ods (OpenOffice Calc spreadsheet file),.odt (OpenOffice Writer document file),.part (partially downloaded file),.pdb (program database),.pdf (PDF file),.php (PHP source code file),.png (PNG / Portable Network Graphic image),.pps (PowerPoint slide show),.ppt (PowerPoint presentation),.pptx (PowerPoint Open XML presentation),.ps (PostScript file),.psd (PSD / Adobe Photoshop Document image),.py (Python file),.rm (Real Media file),.A file type selected from the group including rss (RSS / Rich Site Summary file),.rtf (Rich Text Format file),.sav (Saved file),.sql (SQL / Structured Query Language database file),.svg (Scalable Vector Graphics file),.swf (Small Web Format file, formerly known as ShockWave Flash file),.sys (Windows system file),.tar (Linux / Unix tarball file archive),.tex (TeX document file),.tif or.tiff (TIFF image),.tmp (Temporary file),.txt (Plain text file),.vob (DVD video object file),.wav (WAVE file),.wks and.wps (Microsoft Works word processor document file),.wma (Windows Media audio file),.wmv (Windows Media video file),.wpd (WordPerfect document),.wpl (Windows Media Player playlist),.wsf (Windows script file),.xhtml (XHTML / Extensible Hypertext Markup Language file),.xlr (Microsoft Works spreadsheet file),.xls (Microsoft Excel file),.xlsx (Microsoft Excel Open XML spreadsheet file).
[0033] In one embodiment, the digital sequence may be selected from the group including program files, text files, table files, audio files, image files, video files, and combinations thereof.
[0034] In one embodiment, the digital sequence may be included in or consist of a program file. Non-limiting examples of program files include .accdb (Access 2007 database file), .apk (Android package file), .bak (backup file), .bat (batch file), .bin (binary file), .cab (Windows cabinet file), .cfg (configuration file), .cgi (Common Gateway Interface script), .com (MS-DOS command file), .cpl (Windows control panel file), .csv (comma-separated values file), .cur (Windows cursor file), .dat (data file), .db or .dbf (database file), .dll (DLL file), .dmp (dump file), .drv (device driver file), .exe (executable file), .icns (macOS X icon resource file), .ico (icon file), .ini (initial settings file), .jar (Java archive file), .lnk (Windows shortcut file), .log (log file), .mdb (Microsoft Access database file), .msi (Windows installer package), .pdb (program database), .py (Python file), .sav (saved file), .sql (SQL / Structured Query Language database file), .sys (Windows system file), .tar (Linux / Unix tarball file archive), .tmp (temporary file), and .wsf (Windows script file).
[0035] In one embodiment, the digital sequence may be included in or consist of a text file. Non-limiting examples of text files include.doc and.docx (Microsoft Word files),.odt (OpenOffice Writer document files),.msg (Outlook mail messages),.pdf (PDF files),.rtf (rich text format files),.tex (TeX document files),.txt (plain text files),.wks and.wps (Microsoft Works word processor document files), and.wpd (WordPerfect documents).
[0036] In one embodiment, the digital sequence may be included in or consist of a table file, such as a spreadsheet. Non-limiting examples of table files include.ods (OpenOffice Calc spreadsheet files),.xlr (Microsoft Works spreadsheet files),.xls (Microsoft Excel files), and.xlsx (Microsoft Excel Open XML spreadsheet files).
[0037] In one embodiment, the digital sequence may be included in or consist of an audio file, such as a music file. Non-limiting examples of audio files include.aif (AIF / Audio Interchange audio files),.cda (CD audio track files),.iff (Interchange File Format),.mid or.midi (MIDI audio files),.mp3 (MP3 audio files),.mpa (MPEG-2 audio files),.wav (WAVE files),.wma (Windows Media audio files), and.wpl (Windows Media Player playlists).
[0038] In one embodiment, the digital sequence may be included in or consist of an image file. Non-limiting examples of image files include .ai (Adobe Illustrator file), .bmp (bitmap image file), .gif (GIF / Graphical Interchange Format image), .ico (icon file), .jpeg or .jpg (JPEG image), .max (3ds Max Scene file), .obj (Wavefront 3D Object file), .png (PNG / Portable Network Graphic image), .ps (PostScript file), .eps (Encapsulated PostScript file), .psd (PSD / Adobe Photoshop Document image), .svg (Scalable Vector Graphics file), .tif or .tiff (TIFF image), .3ds (3D Studio Scene), and .3dm (Rhino 3D Model).
[0039] In one embodiment, the digital sequence may be included in or consist of a video file. Non-limiting examples of video files include .avi (Audio Video Interleave file), .flv (Adobe Flash Video file), .h264 (H.264 video file), .m4v (Apple MP4 video file), .mkv (Matroska Multimedia Container), .mov (Apple QuickTime movie file), .mp4 (MPEG-4 video file), .mpg or .mpeg (MPEG video file), .rm (Real Media file), .swf (Shockwave flash file), .vob (DVD video object file), .wmv (Windows Media video file), .3g2 (3GPP2 multimedia file), and .3gp (3GPP multimedia file).
[0040] In one embodiment, the total number of bytes in a digital sequence, i.e., the digital subsequence containing m bits, is referred to as n, and the value of n is at least 1. As used herein, the phrase "at least one" means 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 88, 90, 92, 94, 96, 98, 100, 128, 256, 500, 512, 1000, 1024, 2048, 4096, 8192, 10 4 、 10 5 、 10 6 、 10 7 、 10 8 、 10 9 、 10 10 、 10 11 、 10 12 、 10 13 、 10 14 、 10 15 、 10 16 、 10 17 、 10 18 、 10 19 、 10 20 、 10 21 bytes, or more. Thus, in practice, the number of bits included in the digital sequence is equal to the number obtained by multiplying m (i.e., the number of bits per byte) by n (i.e., the number of bytes, or digital subsequences, included in the digital sequence).
[0041] Each bit has a defined position within a digital subsequence (or byte) containing m bits, where the first position is position 0 and the last position is equal to m - 1. Thus, the position of each bit in the digital sequence can be even or odd, with even positions including 0, 2, 4, 6, 8, 10, 12, and 14, and odd positions including 1, 3, 5, 7, 9, 11, 13, and 15. In one embodiment, the digital subsequence is an octet, with even positions including 0, 2, 4, and 6, and odd positions including 1, 3, 5, and 7.
[0042] In one embodiment, the invention includes the step of converting a byte stored on an electronic device to a byte stored on a nucleic acid molecule, where the byte stored on the nucleic acid molecule is referred to herein as a bioblock, and the bioblock includes m nucleotides. In one embodiment, the byte is an octet, i.e., m = 8, and the bioblock is referred to herein as a biooctet.
[0043] In one embodiment, the bioblock includes 2, 3, or 4 distinct nucleotides, which are referred to herein as N1, N2, N3, and N4. In one embodiment, the biooctet includes exactly 4 distinct nucleotides.
[0044] In one embodiment, both the value and position of each bit included in the byte are encoded in the corresponding bioblock, - A bit having a value of 0 and located at an even position corresponds to the first nucleotide N1, - A bit having a value of 1 and located at an even position corresponds to the second nucleotide N2, - A bit having a value of 0 and located at an odd position corresponds to the third nucleotide N3, - A bit having a value of 1 and located at an odd position corresponds to the fourth nucleotide N4, N1, N2, N3, and N4 are distinct nucleotides.
[0045] The method according to the present invention includes constructing at least one component, preferably a plurality of components, each component including or consisting of at least one bioblock (e.g., at least one biooctet), and the total number of components is x. In one embodiment, the number of bioblocks (e.g., biooctets) per component is y, and the value of y is at least 1. In one embodiment, the value of x is the value obtained by dividing n by y
[0046]
Number
[0047] 。
[0048] As used herein, the phrase "a plurality of" means two, three, four, five, six, seven, eight, nine, ten, twenty, thirty, forty, fifty, one hundred, one thousand, or more. As used herein, the phrase "at least one" means one, two, three, four, five, six, seven, eight, nine, ten, twelve, fourteen, sixteen, eighteen, twenty, twenty-two, twenty-four, twenty-six, twenty-eight, thirty, thirty-two, thirty-four, thirty-six, thirty-eight, forty, forty-two, forty-four, forty-six, forty-eight, fifty, fifty-two, fifty-four, fifty-six, fifty-eight, sixty, sixty-two, sixty-four, sixty-six, sixty-eight, seventy, seventy-two, seventy-four, seventy-six, seventy-eight, eighty, eighty-two, eighty-four, eighty-six, eighty-eight, ninety, ninety-two, ninety-four, ninety-six, ninety-eight, one hundred, one hundred and twenty-eight, two hundred and fifty-six, five hundred, five hundred and twelve, one thousand, ten 4 、10 5 、10 6 or more.
[0049] In one embodiment, each component includes the same number of bioblocks. In one embodiment, x = n, i.e., y = 1.
[0050] In another embodiment, x and n are clearly different, i.e., y ≠ 1, which means that each component includes from 2 to n bioblocks (e.g., from 2 to n biooctets).
[0051] In some embodiments, the value of x is not the value obtained by dividing n by y.
[0052] In one embodiment, y does not have a fixed value, i.e., at least two, three, four, five, or more components contain a distinct number of bio - blocks. In some embodiments, each component contains a distinct number of bio - blocks.
[0053] In some embodiments, each component contains the same number of bio - blocks (y), except for one component that contains from 1 to y - 1 bio - blocks, and the value of y is at least 2.
[0054] In one embodiment, the x components are assembled together in a fixed order, and the fixed order used to assemble the x components is the same as the order of the n digital sub - sequences within the digital sequence.
[0055] In one embodiment, the assembly of the x components is performed in one or more steps. In one embodiment, the assembly of the x components is performed in one step. In one embodiment, the assembly of the x components is performed in multiple steps. In one embodiment, the assembly of the x components is performed sequentially, separately, simultaneously, or in combination thereof.
[0056] In one embodiment, the nucleotide is selected from the group consisting of natural nucleotides and non - natural nucleotides.
[0057] Natural nucleotides include adenine, guanine, cytosine, uracil, and thymine.
[0058] Non-limiting examples of unnatural nucleotides are 2-amino-ATP, 8-aza-ATP, 2'-fluoro-dATP, 2'-fluoro-dCTP, 2'-fluoro-dGTP, 2'-fluoro-dUTP, 5-iodo-CTP, 5-iodo-UTP, N6-methyl-ATP, 5-methyl-CTP, 2'-O-methyl-ATP, 2'-O-methyl-CTP, 2'-O-methyl-GTP, 2'-O-methyl-UTP, pseudo-UTP, ITP, 2'-O-methyl-ITP, puromycin-TP, xanthosine-TP, 5-methyl-UTP, 4-thio-UTP, 2'-amino-dCTP, 2'-amino-dUTP, 2'-azido-dCTP, 2'-azido-dUTP, O6-methyl-GTP, 2-thio-UTP, Ara-CTP, Ara-UTP, 5,6-dihydro-UTP, 2-thio-CTP, 6-aza-CTP, 6-aza-UTP, N1-methyl-GTP, 2'-O-methyl-2-amino-ATP, 2'-O-methylpseudo-UTP, N1-methyl-ATP, 2'-O-methyl-5-methyl-UTP, 7-deaza-GTP, 2'-azido-dATP, 2'-amino-dATP, Ara-ATP, 8-azido-ATP, 5-bromo-CTP, 5-bromo-UTP, 2'-fluoro-dTTP, 3'-O-methyl-ATP, 3'-O-methyl-CTP, 3'-O-methyl-GTP, 3'-O-methyl-UTP, 7-deaza-ATP, 5-AA-UTP, 2'-azido-dGTP, 2'-amino-dGTP, 5-AA-CTP, 8-oxo-GTP, pseudoiso-CTP, N4-methyl-CTP, N1-methylpseudo-UTP, 5,6-dihydro-5-methyl-UTP, N6-methylamino-ATP, 5-carboxy-CTP, 5-formyl-CTP, 5-hydroxymethyl-UTP, 5-hydroxymethyl-CTP, thieno-GTP, 5-hydroxy-CTP, 5-formyl-UTP, thieno-UTP, 2-amino-dATP, 5-bromo-dCTP, 5-bromo-dUTP, 7-deaza-dATP, 7-deaza-dGTP, dITP, 5-propynyl-dCTP, 5-propynyl-dUTP, 2'-dUTP, 5-fluoro-dUTP, 5-iodo-dCTP, 5-iodo-dUTP, N6-methyl-dATP, 5-methyl-dCTP, O6-methyl-dGTP, N2-methyl-dGTP, 8-oxo-dATP,8-oxo-dGTP, 2-thio-dTTP, 2'-dPTP, 5-hydroxy-dCTP, 4-thio-dTTP, 2-thio-dCTP, 6-aza-dUTP, 6-thio-dGTP, 8-chloro-dATP, 5-AA-dCTP, 5-AA-dUTP, N4-methyl-dCTP, 2'-deoxyze-braline-TP, 5-hydroxymethyl-dUTP, 5-hydroxymethyl-dCTP, 5-propynylamino-dCTP, 5-propynylamino-dUTP, 5-carboxy-dCTP, 5-formyl-dCTP, 5-indolyl-AA-dUTP, 5-carboxy-dUTP, 5-formyl-dUTP, 3'-dATP, 3'-dGTP, 3'-dCTP, 5-methyl-3'-dUTP, 3'-dUTP, ddATP, ddGTP, ddUTP, ddTTP, ddCTP, 3'-azido-ddATP, 3'-azido-ddGTP, 3'-azido-ddTTP, 3'-amino-ddATP, 3'-amino-ddCTP, 3'-amino-ddGTP, 3'-amino-ddTTP, 3'-azido-ddCTP, 3'-azido-ddUTP, 5-bromo-ddUTP, ddITP, (1-thio)-dATP, (1-thio)-dCTP, (1-thio)-dGTP, (1-thio)-dTTP, (1-thio)-ATP, (1-thio)-CTP, (1-thio)-GTP, (1-thio)-UTP, (1-thio)-ddATP, (1-thio)-ddCTP, (1-thio)-ddGTP, (1-thio)-ddTTP, (1-thio)-3'-azido-ddTTP, (1-thio)-ddUTP, (1-borano)-dATP, (1-borano)-dCTP, (1-borano)-dGTP, (1-borano)-dTTP, ganciclovir-TP, cidofovir-DP, 3-methyl-6-amino-5-(1'-b-D-2'-deoxyribofuranosyl)-pyrimidin-2-one, 6-amino-9[(1'-b-D-2'-deoxyribofuranosyl)-4-hydroxy-5-(hydroxymethyl)-oxolan-2-yl]-1H-purin-2-one, 6-amino-3-(1'-b-D-2'-deoxyribofuranosyl)-5-nitro-1H-pyridin-2-one and 2-amino-8-(1'-b-D-2'-deoxyribofuranosyl)-imidazo[1,2a]-1,3,5-triazin-[8H]-4-one are included.,
[0059] In one embodiment, N1, N2, N3, and N4 are selected from the group consisting of or including adenine, guanine, cytosine, uracil, thymine, and non-natural nucleotides. In one embodiment, N1, N2, N3, and N4 are selected from the group consisting of or including adenine, guanine, cytosine, uracil, and thymine.
[0060] In one embodiment, N1, N2, N3, and N4 are selected from the group consisting of or comprising adenine, guanine, cytosine, and thymine. In one embodiment, N1 is adenine, N2 is guanine, N3 is cytosine, and N4 is thymine. In another embodiment, N1 is adenine, N2 is guanine, N3 is thymine, and N4 is cytosine. In another embodiment, N1 is adenine, N2 is cytosine, N3 is thymine, and N4 is guanine. In another embodiment, N1 is adenine, N2 is cytosine, N3 is guanine, and N4 is thymine. In another embodiment, N1 is adenine, N2 is thymine, N3 is cytosine, and N4 is guanine. In another embodiment, N1 is adenine, N2 is thymine, N3 is guanine, and N4 is cytosine. In another embodiment, N1 is guanine, N2 is adenine, N3 is cytosine, and N4 is thymine. In another embodiment, N1 is guanine, N2 is adenine, N3 is thymine, and N4 is cytosine. In another embodiment, N1 is guanine, N2 is cytosine, N3 is adenine, and N4 is thymine. In another embodiment, N1 is guanine, N2 is cytosine, N3 is thymine, and N4 is adenine. In another embodiment, N1 is guanine, N2 is thymine, N3 is adenine, and N4 is cytosine. In another embodiment, N1 is guanine, N2 is thymine, N3 is cytosine, and N4 is adenine. In another embodiment, N1 is cytosine, N2 is adenine, N3 is guanine, and N4 is thymine. In another embodiment, N1 is cytosine, N2 is adenine, N3 is thymine, and N4 is guanine. In another embodiment, N1 is cytosine, N2 is guanine, N3 is adenine, and N4 is thymine. In another embodiment, N1 is cytosine, N2 is guanine, N3 is thymine, and N4 is adenine. In another embodiment, N1 is cytosine, N2 is thymine, N3 is adenine, and N4 is guanine.In another embodiment, N1 is cytosine, N2 is thymine, N3 is guanine, and N4 is adenine. In another embodiment, N1 is thymine, N2 is adenine, N3 is guanine, and N4 is cytosine. In another embodiment, N1 is thymine, N2 is adenine, N3 is cytosine, and N4 is guanine. In another embodiment, N1 is thymine, N2 is guanine, N3 is adenine, and N4 is cytosine. In another embodiment, N1 is thymine, N2 is guanine, N3 is cytosine, and N4 is adenine. In another embodiment, N1 is thymine, N2 is cytosine, N3 is adenine, and N4 is guanine. In another embodiment, N1 is thymine, N2 is cytosine, N3 is guanine, and N4 is adenine.
[0061] In one embodiment, N1, N2, N3, and N4 are selected from the group consisting of or comprising adenine, guanine, cytosine, and uracil. In one embodiment, N1 is adenine, N2 is guanine, N3 is cytosine, and N4 is uracil. In another embodiment, N1 is adenine, N2 is guanine, N3 is uracil, and N4 is cytosine. In another embodiment, N1 is adenine, N2 is cytosine, N3 is uracil, and N4 is guanine. In another embodiment, N1 is adenine, N2 is cytosine, N3 is guanine, and N4 is uracil. In another embodiment, N1 is adenine, N2 is uracil, N3 is cytosine, and N4 is guanine. In another embodiment, N1 is adenine, N2 is uracil, N3 is guanine, and N4 is cytosine. In another embodiment, N1 is guanine, N2 is adenine, N3 is cytosine, and N4 is uracil. In another embodiment, N1 is guanine, N2 is adenine, N3 is uracil, and N4 is cytosine. In another embodiment, N1 is guanine, N2 is cytosine, N3 is adenine, and N4 is uracil. In another embodiment, N1 is guanine, N2 is cytosine, N3 is uracil, and N4 is adenine. In another embodiment, N1 is guanine, N2 is uracil, N3 is adenine, and N4 is cytosine. In another embodiment, N1 is guanine, N2 is uracil, N3 is cytosine, and N4 is adenine. In another embodiment, N1 is cytosine, N2 is adenine, N3 is guanine, and N4 is uracil. In another embodiment, N1 is cytosine, N2 is adenine, N3 is uracil, and N4 is guanine. In another embodiment, N1 is cytosine, N2 is guanine, N3 is adenine, and N4 is uracil. In another embodiment, N1 is cytosine, N2 is guanine, N3 is uracil, and N4 is adenine. In another embodiment, N1 is cytosine, N2 is uracil, N3 is adenine, and N4 is guanine.In another embodiment, N1 is cytosine, N2 is uracil, N3 is guanine, and N4 is adenine. In another embodiment, N1 is uracil, N2 is adenine, N3 is guanine, and N4 is cytosine. In another embodiment, N1 is uracil, N2 is adenine, N3 is cytosine, and N4 is guanine. In another embodiment, N1 is uracil, N2 is guanine, N3 is adenine, and N4 is cytosine. In another embodiment, N1 is uracil, N2 is guanine, N3 is cytosine, and N4 is adenine. In another embodiment, N1 is uracil, N2 is cytosine, N3 is adenine, and N4 is guanine. In another embodiment, N1 is uracil, N2 is cytosine, N3 is guanine, and N4 is adenine.
[0062] In one embodiment, N1, N2, N3, and N4 are selected from the group consisting of or comprising unnatural nucleotides.
[0063] In one embodiment, the x components are nucleic acid molecules selected from the group consisting of or comprising double-stranded DNA molecules, single-stranded DNA molecules, double-stranded RNA molecules, single-stranded RNA molecules, and nucleic acid molecules comprising at least one unnatural nucleotide.
[0064] In one embodiment, the x components are x DNA molecules, preferably x double-stranded DNA molecules.
[0065] In one embodiment, the x components are double-stranded DNA molecules. In one embodiment, the x components are single-stranded DNA molecules.
[0066] In another embodiment, the x components are double-stranded RNA molecules or single-stranded RNA molecules. In another embodiment, the x components are nucleic acid molecules comprising at least one unnatural nucleotide.
[0067] In one embodiment, the construction of a plurality of x components, each containing at least one bioblock, - selectively capturing x data storage nucleic acid molecules from at least one library of data storage nucleic acid molecules, each data storage nucleic acid molecule containing at least one bioblock surrounded by a region containing a cleavage site; - cleaving each of the x data storage nucleic acid molecules, thereby releasing at least one bioblock.
[0068] Within the scope of the present invention, a "data storage nucleic acid molecule" is a molecule, typically a plasmid, containing at least one bioblock (e.g., at least one biooctet) or a component according to the present invention, and each bioblock (e.g., biooctet) or component is flanked by a region containing a cleavage site. In one embodiment, the data storage nucleic acid molecule contains or consists of nucleotides selected from the group consisting of natural and unnatural nucleotides.
[0069] Within the scope of the present invention, the term "library of data storage nucleic acid molecules" refers to a plurality of clearly defined data storage nucleic acid molecules as defined herein, and each data storage nucleic acid molecule in the library contains distinct bioblocks (e.g., biooctets) or components.
[0070] As used herein, the term "cleavage site" refers to a nucleotide sequence targeted by an enzyme selected from the group consisting of or including restriction enzymes (also called restriction endonucleases), endonucleases, exonucleases, deoxyribonucleases, ribonucleases, nickases, transposases, integrases. In a preferred embodiment, the enzyme is a site-specific enzyme, i.e., an enzyme that recognizes a specific nucleic acid sequence.
[0071] In one embodiment, the cleavage site is a target for a restriction enzyme. In one embodiment, the cleavage site is a restriction enzyme recognition site. As used herein, the term "restriction enzyme recognition site" refers to a nucleotide sequence that is a target for a specific restriction enzyme. Non-limiting examples of restriction enzymes include EcoRI, BamHI, HindIII, KpnI, NotI, PstI, SmaI, and XhoI. Restriction enzymes and corresponding restriction enzyme recognition sites are well known in the art.
[0072] In another embodiment, the cleavage site is a target for an enzyme selected from the group consisting of or including endonucleases, exonucleases, deoxyribonucleases, ribonucleases, nickases, integrases, and transposases.
[0073] In one embodiment, the region containing the cleavage site includes a first nucleotide sequence recognized by an enzyme, typically a restriction enzyme, and a second nucleotide sequence that is digested or cleaved by the enzyme. In one embodiment, the first nucleotide sequence and the second nucleotide sequence are distinct. In some embodiments, the first nucleotide sequence and the second nucleotide sequence are separated by at least one nucleotide. In one embodiment, digestion of the cleavage site separates the first nucleotide sequence from the second nucleotide sequence.
[0074] In one embodiment, digestion of the cleavage site produces overhangs or blunt ends, preferably overhangs. Within the scope of the present invention, these overhangs are referred to herein as "fusion sites". In one embodiment, the overhangs are 3' overhangs or 5' overhangs. In one embodiment, the nucleotide sequences of the 3' overhangs and 5' overhangs are complementary.
[0075] In one embodiment, the construction of a plurality of x components, each containing at least one bio-block (e.g., a bio-octet), is - Selectively capturing n data storage nucleic acid molecules from at least two libraries of data storage nucleic acid molecules, wherein each data storage nucleic acid molecule in each library comprises one bio-block (e.g., a bio-octet) surrounded by a region containing a cleavage site, and each library comprises all possible bio-blocks of m nucleotides (e.g., all possible bio-octets of 8 nucleotides), - Cleaving each of the n data storage nucleic acid molecules, thereby releasing n bio-blocks (e.g., bio-octets).
[0076] In one embodiment, the region containing the cleavage site comprises from 2 to 25 nucleotides.
[0077] As used herein, the expression "from 2 to 25 nucleotides" includes 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, and 25 nucleotides.
[0078] In one embodiment, the region containing the cleavage site comprises from 2 to 20 nucleotides. In one embodiment, the region containing the cleavage site comprises from 2 to 15 nucleotides. In one embodiment, the region containing the cleavage site comprises from 2 to 10 nucleotides.
[0079] In one embodiment, the cleavage site is located both upstream and downstream of the bio-block (e.g., bio-octet) or component.
[0080] As used herein, the term "upstream" refers to the following position. - When the data storage nucleic acid molecule is a single-stranded nucleic acid molecule, it is adjacent to the 5'-side of the most 5'-terminal of the sequence of a bioblock (e.g., biooctet) or a component, and the adjacency means being continuous or separated by a spacer (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides), or - On the 5'-side of the most 5'-terminal of the sequence of a bioblock (e.g., biooctet) or a component on the plus strand, and when the data storage nucleic acid molecule is a double-stranded nucleic acid molecule, on the 3'-side of the most 3'-terminal of the sequence of a bioblock (e.g., biooctet) or a component on the minus strand, and the adjacency means being continuous or separated by a spacer (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides).
[0081] As used herein, the term "downstream" refers to the following positions. - When the data storage nucleic acid molecule is a single-stranded nucleic acid molecule, it is adjacent to the 3'-side of the most 3'-terminal of the sequence of a bioblock (e.g., biooctet) or a component, and the adjacency means being continuous or separated by a spacer (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides), or - On the 3'-side of the most 3'-terminal of the sequence of a bioblock (e.g., biooctet) or a component on the plus strand, and when the data storage nucleic acid molecule is a double-stranded nucleic acid molecule, on the 5'-side of the most 5'-terminal of the sequence of a bioblock (e.g., biooctet) or a component on the minus strand, and the adjacency means being continuous or separated by a spacer (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides).
[0082] In one embodiment, the data storage nucleic acid molecule includes a number of upstream regions containing cleavage sites equal to the number of downstream regions containing cleavage sites. In one embodiment, the data storage nucleic acid molecule includes at least one upstream region containing a cleavage site and at least one downstream region containing a cleavage site. In one embodiment, the data storage nucleic acid molecule includes one upstream region containing a cleavage site and one downstream region containing a cleavage site. In one embodiment, the data storage nucleic acid molecule includes two upstream regions containing cleavage sites and two downstream regions containing cleavage sites.
[0083] In one embodiment, the data storage nucleic acid molecule includes at least two distinct cleavage sites, the distinct cleavage sites having distinct nucleic acid sequences and preferably being digested by distinct enzymes.
[0084] In another embodiment, the upstream cleavage site and the downstream cleavage site are similar and are cleaved by distinct enzymes. In another embodiment, the upstream cleavage site and the downstream cleavage site are similar and are cleaved by the same enzyme.
[0085] In a preferred embodiment, the upstream cleavage site and the downstream cleavage site are distinct and are cleaved by the same enzyme.
[0086] In one embodiment, the data storage nucleic acid molecule further includes two additional cleavage sites, the first being located upstream of a bioblock (e.g., a biooctet) or component and the second being located downstream of a bioblock (e.g., a biooctet) or component.
[0087] In one embodiment, the two additional cleavage sites are distinct and are cleaved by the same enzyme. In another embodiment, the two additional cleavage sites are distinct and are cleaved by distinct enzymes. In another embodiment, the two additional cleavage sites are similar and are cleaved by the same enzyme. In another embodiment, the two additional cleavage sites are similar and are cleaved by distinct enzymes.
[0088] In one embodiment, the two additional cleavage sites are distinct from other cleavage sites contained on the data storage nucleic acid molecule and are cleaved by an enzyme distinct from the enzyme that cleaves the cleavage sites contained on the data storage nucleic acid molecule. In another embodiment, the two additional cleavage sites are similar to other cleavage sites contained on the data storage nucleic acid molecule and are cleaved by an enzyme similar to the enzyme that cleaves the cleavage sites contained on the data storage nucleic acid molecule.
[0089] In one embodiment, it is contemplated that a bioblock (e.g., a biooctet) or component is released when at least one upstream cleavage site and at least one downstream cleavage site are cleaved (i.e., digested or cleaved).
[0090] In one embodiment, the released bioblock (e.g., a biooctet) comprises (i) one bioblock (e.g., a biooctet), (ii) a portion of the nearest upstream cleavage site, i.e., the upstream fusion site, and (iii) a portion of the nearest downstream cleavage site, i.e., the downstream fusion site. In one embodiment, the portion of the nearest upstream cleavage site, i.e., the upstream fusion site, is an overhanging end (e.g., a 3' overhanging end). In one embodiment, the portion of the nearest downstream cleavage site, i.e., the downstream fusion site, is an overhanging end (e.g., a 5' overhanging end).
[0091] In one embodiment, the released component comprises (i) at least one bioblock (e.g., a biooctet), (ii) a portion of the nearest upstream cleavage site, i.e., the upstream fusion site, and (iii) a portion of the nearest downstream cleavage site, i.e., the downstream fusion site. In a preferred embodiment, the released component comprises (i) y bioblocks (e.g., biooctets), (ii) a portion of the nearest upstream cleavage site, i.e., the upstream fusion site, and (iii) a portion of the nearest downstream cleavage site, i.e., the downstream fusion site.
[0092] In one embodiment, assembling together a plurality of x components involves releasing a bioblock (e.g., a biooctet) or a component. In one embodiment, releasing a bioblock (e.g., a biooctet) or a component involves using one enzyme or two distinct enzymes.
[0093] In one embodiment, each of the regions surrounding each bioblock (e.g., a biooctet) contains a site for a restriction enzyme, and step (d) of the method of the invention includes digesting each of the x data storage nucleic acid molecules with one or two restriction enzymes.
[0094] In another embodiment, each of the regions surrounding each bioblock (e.g., a biooctet) contains a site for a restriction enzyme, and step (d) of the method of the invention includes digesting each of the x data storage nucleic acid molecules with two restriction enzymes.
[0095] In one embodiment, digestion of the upstream restriction enzyme recognition site produces a 3'-overhang or a 5'-overhang, and digestion of the downstream restriction enzyme recognition site produces a 3'-overhang or a 5'-overhang. In one embodiment, the nucleotide sequences of the 3'-overhang and the 5'-overhang are complementary.
[0096] In one embodiment, a restriction enzyme recognition site includes a first nucleotide sequence recognized by a restriction enzyme and a second nucleotide sequence that is digested or cleaved by the enzyme. In one embodiment, the first nucleotide sequence and the second nucleotide sequence are distinct. In some embodiments, the first nucleotide sequence and the second nucleotide sequence are separated by at least one nucleotide. In one embodiment, digestion of the restriction enzyme recognition site separates the first nucleotide sequence from the second nucleotide sequence.
[0097] In one embodiment, the restriction enzyme is selected from the group consisting of or comprising a type I, type II, type III, type IV, or type V restriction enzyme, or a combination thereof. In one embodiment, the restriction enzyme is a type II restriction enzyme. In one embodiment, the type II restriction enzyme is selected from the group consisting of or comprising a type II S, type II G, type II B, type II T, and / or type II C restriction enzyme, or a combination thereof, preferably type II S and / or type II G, more preferably type II S. Non-limiting examples of restriction enzyme II S include BsaI, BbsI, BsmBI, FokI, Alw26I, BbvI, BsrI, EarI, HphI, MboII, SfaNI, and Tth111I. In one embodiment, the restriction enzyme is BsaI and / or BbsI and / or BsmBI.
[0098] In one embodiment, the restriction enzyme is modified. In one embodiment, the restriction enzyme contains at least one mutation in the amino acid sequence as compared to the unmodified (or wild-type) amino acid sequence. In one embodiment, the restriction enzyme is post-translationally modified.
[0099] In some embodiments, the enzyme recognition site consists of a nucleotide sequence selected from GGTCTC and CGTCTC.
[0100] In one embodiment, the cleavage site includes a nucleotide sequence selected from the group consisting of or comprising GTAG, TGAC, TCAG, AATA, TCAA, CTTC, AGTA, ACTG, CACA, CCAG, CAAA, GACC, ACTC, CCAC, GAAC, GCAC, CGGC, CGTA, GTAA, CAAC, GCTA, CCGA, ACGA, AGAA, TAAA, AGCG, ACCT, AACA, GGCA, ACGC, AATC, CGAG, TCCA, CCTA, CTAA, GGGA, AAGG, AAAC, CTAC, and GAGA. In one embodiment, these sequences are overhangs.
[0101] In one embodiment, the fusion site comprises a nucleotide sequence selected from the group consisting of or including GTAG, TGAC, TCAG, AATA, TCAA, CTTC, AGTA, ACTG, CACA, CCAG, CAAA, GACC, ACTC, CCAC, GAAC, GCAC, CGGC, CGTA, GTAA, CAAC, GCTA, CCGA, ACGA, AGAA, TAAA, AGCG, ACCT, AACA, GGCA, ACGC, AATC, CGAG, TCCA, CCTA, CTAA, GGGA, AAGG, AAAC, CTAC, and GAGA.
[0102] In one embodiment, the cleavage site comprises a nucleotide sequence selected from the group consisting of or including GTAG, TGAC, TCAG. In one embodiment, the cleavage site comprises a nucleotide sequence selected from the group consisting of or including AATA, TCAA, CTTC, AGTA, ACTG, CACA, CCAG, CAAA, GACC, ACTC, CCAC, GAAC, GCAC, CGGC, CGTA, GTAA, CAAC, GCTA, CCGA, ACGA, AGAA, TAAA, AGCG, ACCT, AACA, GGCA, ACGC, AATC, CGAG, TCCA, CCTA, CTAA, and GGGA. In one embodiment, the cleavage site comprises a nucleotide sequence selected from the group consisting of or including AATA, AAGG, AAAC, TAAA, ACGA, ACTG, AGCG, GCTA, GGCA, ACCT, CGTA, AACA, CTAC, GAGA, CCAG, AGAA, and GCAC.
[0103] In one embodiment, the fusion site comprises a nucleotide sequence selected from the group consisting of or including GTAG, TGAC, TCAG. In one embodiment, the fusion site comprises a nucleotide sequence selected from the group consisting of or including AATA, TCAA, CTTC, AGTA, ACTG, CACA, CCAG, CAAA, GACC, ACTC, CCAC, GAAC, GCAC, CGGC, CGTA, GTAA, CAAC, GCTA, CCGA, ACGA, AGAA, TAAA, AGCG, ACCT, AACA, GGCA, ACGC, AATC, CGAG, TCCA, CCTA, CTAA, and GGGA. In one embodiment, the fusion site comprises a nucleotide sequence selected from the group consisting of or including AATA, AAGG, AAAC, TAAA, ACGA, ACTG, AGCG, GCTA, GGCA, ACCT, CGTA, AACA, CTAC, GAGA, CCAG, AGAA, and GCAC.
[0104] In one embodiment, step (e) comprises one or more assembly steps using overlap extension polymerase chain reaction (PCR), polymerase cycling assembly, sticky end ligation, BioBrick assembly, Golden Gate assembly, Gibson assembly, recombinase assembly, ligase cycling reaction, template-directed ligation, in vivo assembly, or any other DNA assembly protocol.
[0105] In one embodiment, step (e) includes one or more assembly steps using overlap PCR. In one embodiment, step (e) includes one or more assembly steps using polymerase cycling assembly. In one embodiment, step (e) includes one or more assembly steps using sticky-end ligation. In one embodiment, step (e) includes one or more assembly steps using BioBrick assembly. In one embodiment, step (e) includes one or more assembly steps using Golden Gate assembly. In one embodiment, step (e) includes one or more assembly steps using Gibson assembly. In one embodiment, step (e) includes one or more assembly steps using recombinase assembly. In one embodiment, step (e) includes one or more assembly steps using ligase cycling reaction. In one embodiment, step (e) includes one or more assembly steps using template-directed ligation. In one embodiment, step (e) includes one or more assembly steps using in vivo assembly.
[0106] In one embodiment, step (e) includes using a ligase.
[0107] In a preferred embodiment, cleavage of the region containing the cleavage site results in overhanging ends, which are also referred to as fusion sites. In a preferred embodiment, the closest fusion site to one end (e.g., the 3' end) of the first BioBrick (e.g., BioOctet) or component is complementary to the closest fusion site to the other end (e.g., the 5' end) of the second BioBrick (e.g., BioOctet) or component.
[0108] In one embodiment, the assembly of components including at least one BioBrick (e.g., BioOctet) - The closest fusion site to one end (e.g., the 3' end) of the first bio-block (e.g., a bio-octet), and - The closest fusion site to the other end (e.g., the 5' end) of the second bio-block (e.g., a bio-octet) requires complementarity between them or is facilitated by them.
[0109] In one embodiment, the nucleotide sequence recognized by the enzyme is not included on the nucleotide sequence digested by the enzyme. In one embodiment, after digestion of the cleavage site, the nucleotide sequence recognized by the enzyme is lost, i.e., separated from the cleaved sequence. In one embodiment, the cleavage site between two bio-blocks (e.g., two bio-octets) or two components does not contain the nucleotide sequence recognized by the enzyme.
[0110] In one embodiment, an assembled component comprising y bio-blocks (e.g., bio-octets) - y bio-blocks (e.g., bio-octets) in a fixed order, - y + 1 fusion sites on the sides of the bio-blocks comprises or consists of them.
[0111] In one embodiment, an assembled component comprising y bio-blocks (e.g., bio-octets) - y bio-blocks (e.g., bio-octets) in a fixed order, - y + 1 fusion sites on the sides of the bio-blocks, and - two regions including a cleavage site, and the region including the cleavage site is located at the farthest 5' end and the farthest 3' end of the component.
[0112] The present invention further relates to a data storage nucleic acid molecule comprising at least one bio-block, the bio-block consisting of a nucleic acid sequence of m nucleotides assigned to positions from 0 to m-1, - The bio-block is formed from at least two and at most four (i.e., two, three, or four) distinct nucleotides, - Nucleotides at even positions can be selected from the first and second nucleotides, and nucleotides at odd positions can be selected from the third and fourth nucleotides, and the first, second, third, and fourth nucleotides are distinct.
[0113] In one embodiment, the first, second, third, and fourth nucleotides are referred to as N1, N2, N3, and N4, respectively.
[0114] In one embodiment, N1, N2, N3, and N4 are selected from the group consisting of or comprising adenine, guanine, cytosine, uracil, thymine, and unnatural nucleotides, and N1, N2, N3, and N4 are distinct nucleotides. In one embodiment, N1, N2, N3, and N4 are selected from the group consisting of or comprising adenine, guanine, cytosine, uracil, and thymine, and N1, N2, N3, and N4 are distinct nucleotides.
[0115] In one embodiment, N1, N2, N3, and N4 are selected from the group consisting of or comprising adenine, guanine, cytosine, and thymine, and N1, N2, N3, and N4 are distinct nucleotides. In one embodiment, N1 is adenine, N2 is guanine, N3 is cytosine, and N4 is thymine. In another embodiment, N1 is adenine, N2 is guanine, N3 is thymine, and N4 is cytosine. In another embodiment, N1 is adenine, N2 is cytosine, N3 is thymine, and N4 is guanine. In another embodiment, N1 is adenine, N2 is cytosine, N3 is guanine, and N4 is thymine. In another embodiment, N1 is adenine, N2 is thymine, N3 is cytosine, and N4 is guanine. In another embodiment, N1 is adenine, N2 is thymine, N3 is guanine, and N4 is cytosine. In another embodiment, N1 is guanine, N2 is adenine, N3 is cytosine, and N4 is thymine. In another embodiment, N1 is guanine, N2 is adenine, N3 is thymine, and N4 is cytosine. In another embodiment, N1 is guanine, N2 is cytosine, N3 is adenine, and N4 is thymine. In another embodiment, N1 is guanine, N2 is cytosine, N3 is thymine, and N4 is adenine. In another embodiment, N1 is guanine, N2 is thymine, N3 is adenine, and N4 is cytosine. In another embodiment, N1 is guanine, N2 is thymine, N3 is cytosine, and N4 is adenine. In another embodiment, N1 is cytosine, N2 is adenine, N3 is guanine, and N4 is thymine. In another embodiment, N1 is cytosine, N2 is adenine, N3 is thymine, and N4 is guanine. In another embodiment, N1 is cytosine, N2 is guanine, N3 is adenine, and N4 is thymine. In another embodiment, N1 is cytosine, N2 is guanine, N3 is thymine, and N4 is adenine. In another embodiment, N1 is cytosine, N2 is thymine, N3 is adenine, and N4 is guanine.In another embodiment, N1 is cytosine, N2 is thymine, N3 is guanine, and N4 is adenine. In another embodiment, N1 is thymine, N2 is adenine, N3 is guanine, and N4 is cytosine. In another embodiment, N1 is thymine, N2 is adenine, N3 is cytosine, and N4 is guanine. In another embodiment, N1 is thymine, N2 is guanine, N3 is adenine, and N4 is cytosine. In another embodiment, N1 is thymine, N2 is guanine, N3 is cytosine, and N4 is adenine. In another embodiment, N1 is thymine, N2 is cytosine, N3 is adenine, and N4 is guanine. In another embodiment, N1 is thymine, N2 is cytosine, N3 is guanine, and N4 is adenine.
[0116] In one embodiment, N1, N2, N3, and N4 are selected from the group consisting of or including adenine, guanine, cytosine, and uracil, and N1, N2, N3, and N4 are distinct nucleotides. In one embodiment, N1 is adenine, N2 is guanine, N3 is cytosine, and N4 is uracil. In another embodiment, N1 is adenine, N2 is guanine, N3 is uracil, and N4 is cytosine. In another embodiment, N1 is adenine, N2 is cytosine, N3 is uracil, and N4 is guanine. In another embodiment, N1 is adenine, N2 is cytosine, N3 is guanine, and N4 is uracil. In another embodiment, N1 is adenine, N2 is uracil, N3 is cytosine, and N4 is guanine. In another embodiment, N1 is adenine, N2 is uracil, N3 is guanine, and N4 is cytosine. In another embodiment, N1 is guanine, N2 is adenine, N3 is cytosine, and N4 is uracil. In another embodiment, N1 is guanine, N2 is adenine, N3 is uracil, and N4 is cytosine. In another embodiment, N1 is guanine, N2 is cytosine, N3 is adenine, and N4 is uracil. In another embodiment, N1 is guanine, N2 is cytosine, N3 is uracil, and N4 is adenine. In another embodiment, N1 is guanine, N2 is uracil, N3 is adenine, and N4 is cytosine. In another embodiment, N1 is guanine, N2 is uracil, N3 is cytosine, and N4 is adenine. In another embodiment, N1 is cytosine, N2 is adenine, N3 is guanine, and N4 is uracil. In another embodiment, N1 is cytosine, N2 is adenine, N3 is uracil, and N4 is guanine. In another embodiment, N1 is cytosine, N2 is guanine, N3 is adenine, and N4 is uracil. In another embodiment, N1 is cytosine, N2 is guanine, N3 is uracil, and N4 is adenine.In another embodiment, N1 is cytosine, N2 is uracil, N3 is adenine, and N4 is guanine. In another embodiment, N1 is cytosine, N2 is uracil, N3 is guanine, and N4 is adenine. In another embodiment, N1 is uracil, N2 is adenine, N3 is guanine, and N4 is cytosine. In another embodiment, N1 is uracil, N2 is adenine, N3 is cytosine, and N4 is guanine. In another embodiment, N1 is uracil, N2 is guanine, N3 is adenine, and N4 is cytosine. In another embodiment, N1 is uracil, N2 is guanine, N3 is cytosine, and N4 is adenine. In another embodiment, N1 is uracil, N2 is cytosine, N3 is adenine, and N4 is guanine. In another embodiment, N1 is uracil, N2 is cytosine, N3 is guanine, and N4 is adenine.
[0117] In some embodiments, N1, N2, N3, and N4 are non-natural nucleotides as described above, and N1, N2, N3, and N4 are distinct nucleotides.
[0118] In one embodiment, the data storage nucleic acid molecule is a double-stranded molecule, preferably a DNA molecule.
[0119] In one embodiment, the double-stranded nucleic acid molecule is circular or linear, preferably circular. In one embodiment, the data storage nucleic acid molecule is a circularized linear sequence. Methods for circularizing DNA sequences are known in the art.
[0120] In one embodiment, the data storage nucleic acid molecule is a plasmid, cosmid, fosmid, prokaryotic chromosome (e.g., bacterial artificial chromosome) or eukaryotic chromosome (e.g., yeast artificial chromosome or human artificial chromosome).
[0121] In a preferred embodiment, the data storage nucleic acid molecule is a plasmid. In another embodiment, the data storage nucleic acid molecule is a cosmid. In another embodiment, the data storage nucleic acid molecule is a fosmid. In another embodiment, the data storage nucleic acid molecule is a prokaryotic chromosome. In another embodiment, the data storage nucleic acid molecule is a eukaryotic chromosome.
[0122] In one embodiment, in the data storage nucleic acid molecule, each of the bio-blocks (e.g., bio-octets) or components is surrounded by a region containing at least one cleavage site. In one embodiment, in the data storage nucleic acid molecule, each of the bio-blocks (e.g., bio-octets) or components is surrounded by a region containing one cleavage site. In another embodiment, in the data storage nucleic acid molecule, each of the bio-blocks (e.g., bio-octets) or components is surrounded by a region containing two cleavage sites, and the cleavage sites within the same region are clearly different.
[0123] In one embodiment, digestion of the region containing the cleavage site by a restriction enzyme generates overhanging ends or blunt ends, preferably overhanging ends (i.e., fusion sites). In one embodiment, the overhanging ends are 3'-overhanging ends or 5'-overhanging ends.
[0124] In one embodiment, the data storage nucleic acid molecule comprises at least one component, and each of the components is surrounded by a region containing one or more cleavage sites.
[0125] In one embodiment, the data storage nucleic acid molecule is replicable.
[0126] As used herein, the "replicable" property of the data storage nucleic acid molecule according to the present invention refers to the ability to be replicated one or more times in vivo in an organism, particularly by a polymerase, more specifically a DNA polymerase.
[0127] In one embodiment, the evaluation of the replicable properties of a nucleic acid molecule can be carried out by any standard method in the art or a method derived therefrom. By way of example, the replicable properties can be evaluated by an increase in the copy number of the nucleic acid molecule in and / or by an organism, and / or by the ability of the organism to transmit the nucleic acid to its progeny.
[0128] In one embodiment, the organism is a microorganism, particularly a bacterium, microalgae, archaea, fungus, phage, virus, or yeast. In one embodiment, the organism is a prokaryote. Non-limiting examples of prokaryotes according to the present invention include bacteria such as Actinobacteria, Chlamydia, Cyanobacteria, Firmicutes, Proteobacteria, Spirochaeta, Thermotoga, and archaea such as Euryarchaeota, Crenarchaeota. In one embodiment, the organism is a bacterium, preferably Escherichia coli, more preferably Escherichia coli strain DH5α.
[0129] In some embodiments, the organism is a eukaryote. Non-limiting examples of eukaryotes according to the present invention include protozoa, algae, plants, fungi, animals, and their respective cells.
[0130] In order to be replicated, the data storage nucleic acid molecule according to the present invention has at least one origin of replication, i.e., one or more sequences of nucleotides recognized by the replication initiation mechanism. Exemplarily, the origins of replication of archaea and bacteria include oriC. In fact, most bacteria have their own origin of replication, archaea have one or more origins of replication, eukaryotes have multiple origins of replication, and may have origins of replication particularly in the form of centromeres. Within the scope of the present invention, the expression "multiple origins of replication" refers to at least 2, 3, 4, 5, 10, 15, 20, 25, 50, 75, 100, 150, 200 origins of replication per nucleic acid molecule.
[0131] In one embodiment, the data storage nucleic acid molecule comprises or consists of (i) at least one component as described above herein, and (ii) at least one origin of replication.
[0132] In one embodiment, the data storage nucleic acid molecule does not contain a promoter region. In one embodiment, the data storage nucleic acid molecule does not contain a biological coding sequence.
[0133] In one embodiment, the data storage nucleic acid molecule is non-coding.
[0134] In one embodiment, the size of the data storage nucleic acid molecule is in the range from 100 base pairs (bp) to 1.10 6 bp. As used herein, the expression "in the range from 100 base pairs (bp) to 10 6 bp" includes 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 5000, 10 4 10 5 and 10 6 bp.
[0135] In some embodiments, the data storage nucleic acid molecule further comprises one or more regions carrying metadata, i.e., information that does not encode digital information. Typically, these regions are referred to as "metadata bioblocks" (e.g., metadata biooctets).
[0136] In some embodiments, the metadata region comprises or consists of at least one barcoding region. As used herein, the term "barcoding region" refers to a bioblock (e.g., a biooctet) added at the beginning of a component or group of components. Typically, the barcode encodes numbers (e.g., 0, 1, 2, 3, 4, and the like) using the same encoding system as the bioblock, and the numbering system enables labeling the components or group of components in a defined order.
[0137] In some embodiments, the metadata region comprises or consists of an "end of file" signal. As used herein, the term "end of file signal" refers to a special bioblock (e.g., a biooctet) having a predefined sequence that is located at the end of a sequence and is not shared with any other bioblock. Typically, the "end of file" signal indicates the end of the region encoding the digital data of a file.
[0138] In some embodiments, the metadata region comprises or consists of at least one barcoding region and one "end of file signal" as described above herein.
[0139] The present invention further relates to a library comprising a plurality of data storage nucleic acid molecules according to the present invention, each of the data storage nucleic acid molecules of the library comprising one bioblock (e.g., a biooctet), each data storage nucleic acid molecule of the library comprising the same peripheral region containing a cleavage site, and the library comprising all possible bioblocks of m nucleotides.
[0140] In one embodiment, each data storage molecule of the library comprises exactly one bioblock (e.g., a biooctet). In one embodiment, the total number of data storage nucleic acid molecules in the library is 2 m equal. In one embodiment, m = 8, and thus the size of the library is 256 data storage nucleic acid molecules.
[0141] In one embodiment, each data storage molecule of the library comprises distinct bioblocks (e.g., biooctets). In fact, the library comprises 2 m distinct bioblocks (e.g., biooctets).
[0142] In one embodiment, two distinct libraries contain distinct bio-blocks (e.g., bio-octets). In another embodiment, two distinct libraries may contain at least one common (i.e., identical) bio-block (e.g., bio-octet). In some embodiments, two distinct libraries contain more than two distinct bio-blocks (e.g., bio-octets). m In another embodiment, each data storage molecule comprises components according to the invention, and each component comprises a plurality of bio-blocks (e.g., bio-octets). In one embodiment, each data storage molecule of a library comprises at least one component. In one embodiment, each data storage molecule of a library comprises distinct components. In one embodiment, two distinct libraries comprise distinct components. In another embodiment, two distinct libraries may contain at least one common (i.e., identical) component.
[0143]
[0144] In one embodiment, each data storage molecule of the library comprises from 1 to 32 components. As used herein, the expression from 1 to 32 includes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, and 32. In one embodiment, each data storage molecule of the library comprises from 2 to 32 components. In one embodiment, each data storage molecule of the library comprises from 4 to 32 components. In one embodiment, each data storage molecule of the library comprises from 8 to 32 components. In one embodiment, each data storage molecule of the library comprises from 16 to 32 components. In one embodiment, each data storage molecule of the library comprises from 1 to 16 components. In one embodiment, each data storage molecule of the library comprises from 1 to 8 components. In one embodiment, each data storage molecule of the library comprises from 1 to 4 components. In one embodiment, each data storage molecule of the library comprises from 1 to 2 components.
[0145] In another embodiment, each data storage molecule of the library comprises more than 32 components.
[0146] In some embodiments, a library containing data storage molecules comprising at least one component is assembled using bio-blocks (e.g., bio-octets) released from at least one library containing data storage molecules comprising exactly one bio-block (e.g., bio-octet), as disclosed in the methods of the present invention. In practice, a nucleic acid molecule comprising exactly one set of cleavage sites identical to the cleavage sites flanking the bio-block (e.g., bio-octet), herein referred to as the recipient molecule, is digested using at least one enzyme, preferably one enzyme, and assembled with at least one bio-block (e.g., bio-octet) using the methods described above herein.
[0147] In some embodiments, a library containing data storage molecules that include a plurality of components is assembled using components released from at least one library containing data storage molecules that include exactly one component, by using a method as disclosed in the present invention.
[0148] In one embodiment, the regions containing cleavage sites included on each data storage molecule of the library are identical.
[0149] In one embodiment, data storage molecules of distinct libraries include distinct regions that include cleavage sites.
[0150] In one embodiment, data storage diffusion molecules of distinct libraries include the same region that includes a cleavage site, and a bio-block (e.g., a bio-octet) or component included in a data storage molecule of a first library is not used to assemble a component included in a data storage molecule of a second library, and a bio-block (e.g., a bio-octet) or component included in a data storage molecule of the second library is not used to assemble a component included in a data storage molecule of the first library.
[0151] In one embodiment, a component can be assembled using bio-blocks (e.g., bio-octets) or components from multiple libraries.
[0152] In one embodiment, the data storage nucleic acid molecules included in the library are, by the method of the present invention, - the nucleic acid sequences of the bio-blocks (e.g., bio-octets) and / or components they include, and / or - the nucleic acid sequences or regions that include cleavage sites surrounding the bio-blocks (e.g., bio-octets) and / or components, and / or - Identified and labeled according to an encoding system used to convert a digital subsequence containing m bits (i.e., the value and position of the bits) into a bio-block. In one embodiment, the encoding system is represented in the form of "(N1, N2, N3, N4)", "(N1, N2, N3)" or "(N1, N2)".
[0153] In one embodiment, the labeling information is digital information and / or physical information. In one embodiment, the labeling information is stored in at least one database.
[0154] In one embodiment, the data storage nucleic acid molecules contained in the library are labeled using a code or identifier that provides no information about the content of the data storage nucleic acid molecule. In one embodiment, information regarding the sequence of the data storage nucleic acid molecule and the encoding system is retrieved by searching for a corresponding code or identifier within at least one database.
[0155] In a preferred embodiment, the data storage nucleic acid molecules contained in the library are stored separately.
[0156] In one embodiment, the data storage nucleic acid molecules of the library are stored at a temperature suitable for preventing nucleic acid degradation. In one embodiment, the data storage nucleic acid molecules contained in the library are stored at a temperature within the range from 4°C to -200°C. As used herein, the expression "from 4°C to -200°C" includes 4, 3, 2, 1, 0, -1, -2, -3, -4, -5, -6, -7, -8, -9, -10, -11, -12, -13, -14, -15, -16, -17, -18, -19, -20, -30, -40, -50, -60, -70, -80, -90, -100, -120, -140, -160, -180, -200°C. In one embodiment, the data storage nucleic acid molecules contained in the library are stored at a temperature within the range from 4°C to -80°C. In one embodiment, the data storage nucleic acid molecules contained in the library are stored at a temperature within the range from 4°C to -20°C. In one embodiment, the data storage nucleic acid molecules contained in the library are stored at a temperature within the range from 4°C to 0°C. In one embodiment, the data storage nucleic acid molecules contained in the library are stored at a temperature within the range from 0°C to -200°C. In one embodiment, the data storage nucleic acid molecules contained in the library are stored at a temperature within the range from -20°C to -200°C. In one embodiment, the data storage nucleic acid molecules contained in the library are stored at a temperature within the range from -80°C to -200°C. In one embodiment, the data storage nucleic acid molecules contained in the library are stored at a temperature of -196°C.
[0157] In one embodiment, the data storage nucleic acid molecules contained in the library are stored in a suitable solvent. Solvents suitable for nucleic acid storage are known in the art. Non-limiting examples of solvents used for nucleic acid storage include aqueous solvents such as deionized water or physiological buffers (e.g., phosphate buffered saline, Tris-HCl).
[0158] In one embodiment, the data storage nucleic acid molecules contained in the library are lyophilized.
[0159] The present invention further relates to a nucleic acid-based data storage system comprising at least two libraries according to the present invention.
[0160] In one embodiment, the data storage nucleic acid molecules of at least two libraries comprise bio-blocks (e.g., bio-octets) and / or components. In one embodiment, the data storage nucleic acid molecules of at least two libraries comprise bio-blocks (e.g., bio-octets).
[0161] In one embodiment, the nucleic acid-based data storage system is for storing data included in a digital sequence as described above herein. In one embodiment, the conversion of information carried by a digital sequence into the nucleic acid-based data storage system, i.e., encoding, is performed using the methods of the present disclosure.
[0162] In one embodiment, the digital data consists of binary digital data. In practice, the conversion of digital data into nucleic acid molecules can be automatically performed in a computer by suitable software.
[0163] In one embodiment, the data included in the digital sequence is stored on at least one data storage nucleic acid molecule, and at least one data storage nucleic acid molecule is assembled from a library according to the present invention using the methods according to the present invention.
[0164] In one embodiment, the nucleic acid-based data storage system can store an amount of information corresponding to a range from 2 to 10 21 bytes. As used herein, "2 to 10 21The phrase "within the range up to bytes" means 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 88, 90, 92, 94, 96, 98, 100, 128, 256, 500, 512, 1000, 1024, 2048, 4096, 8192, 10 4 、 10 5 、 10 6 、 10 7 、 10 8 、 10 9 、 10 10 、 10 11 、 10 12 、 10 13 、 10 14 、 10 15 、 10 16 、 10 17 、 10 18 、 10 19 、 10 20 、 10 21 including bytes.
[0165] Another object of the present invention is computer software that implements uses and methods for storing digital data.
[0166] In one embodiment, the method of the present invention is implemented by a microprocessor including software configured to assign at least one nucleic acid molecule to digital data according to the present invention. In some embodiments, the software is configured to prevent the sequence of the composite nucleic acid molecule according to the present invention from encoding one or more RNAs, and preferably, is configured not to encode mRNA. In some embodiments, the software is configured to prevent the sequence of the composite nucleic acid molecule according to the present invention from including one or more start codons within all six reading frames. In some embodiments, the software is configured to prevent the sequence of the composite nucleic acid molecule according to the present invention from including one or more specific restriction enzyme recognition sites. In some embodiments, the software is configured to prevent the sequence of the composite nucleic acid molecule according to the present invention from including one or more repeats of at least five identical nucleotides.
[0167] In one embodiment, information can be retrieved from a nucleic acid-based data storage system by sequencing at least one nucleic acid molecule. Methods for sequencing nucleic acid molecules, particularly high-throughput sequencing, are known in the art and include, inter alia, Illumina (sequencing by synthesis), single molecule real-time (SMRT) sequencing, nanopore sequencing (e.g., the sequencing solution of Oxford Nanopore Technologies), sequencing by ligation or chain termination sequencing (Sanger method).
[0168] In one embodiment, converting the data retrieved from the data storage system into digital data further includes - an encoding system used to convert a digital subsequence including m bits (i.e., the value and position of the bits) into a bioblock according to the method of the present invention, - the sequence of the region including the cleavage site, - the position and type of the metadata bioblock, - The value of m, - requires obtaining the values of n and x.
[0169] In one embodiment, when data retrieved from a data storage system is converted into digital data, as a result, a sequence of bytes containing m bits is retrieved. In one embodiment, when data retrieved from a data storage system is converted into digital data, as a result, a sequence of octets is retrieved.
[0170] In one embodiment, the information necessary to convert data retrieved from a data storage system into digital data is stored in at least one database. In another embodiment, the information necessary to convert data retrieved from a data storage system into digital data is stored in a metadata bioclock.
[0171] In one embodiment, the conversion of data contained in a data storage system into digital data is, i.e., automated by a suitable software or program. In practice, a program into which (i) the sequence of at least one data storage nucleic acid molecule and (ii) the information necessary to convert the data retrieved from the data storage system into digital data (i.e., the encoding system, the sequence of cleavage sites, the position and type of metadata bioblocks, the value of m, and both the values of n and x) are input provides a sequence of bytes containing m bits, optionally a sequence of octets. Typically, the nucleotides corresponding to the cleavage sites are skipped by the program.
[0172] In one embodiment, the said sequence of bytes, optionally the octets, are read as such. In one embodiment, the said sequence of bytes, optionally the octets, are first converted into a file format as described in the present disclosure. In one embodiment, the converted file is read by a suitable program.
[0173] Another object of the present invention is computer software that implements a use and method for retrieving digital data. In one embodiment, the method of the present invention is implemented using a microprocessor that includes software configured to convert at least one nucleic acid sequence into digital data by using the method described above herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0174]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
[0175] Examples The present invention is further illustrated by the following examples of encoding 8-bit biodata.
[0176] Materials and Methods Practical biodata encoding of a text file containing the poem Liberté written by Paul Eluard in 1942 (Table 1).
[0177] [Table 1A]
[0178] [Table 1B]
[0179]
Table 1C
[0180]
Table 1D
[0181] The text is encoded using the ISO8859-1 standard, also known as Latin-1, to generate File A, which contains 2358 octets (Table 2). File A is compressed as a 7z archive using the LZMA2 algorithm to generate File B, which contains 1137 octets (Table 3). File B corresponds to a digital sequence formed from a plurality of 9096 bits. This digital sequence is subdivided into n = 1137 digital subsequences, each containing m = 8 bits. Each of these 1137 digital subsequences of 8 bits is converted into a bio-block of m = 8 nucleotides called a bio-octet.
[0182]
Table 2A
[0183]
Table 2B
[0184]
Table 2C
[0185]
Table 2D
[0186]
Table 2E
[0187]
Table 2F
[0188]
Table 3A
[0189]
Table 3B
[0190]
Table 3C
[0191] For this conversion, the nucleotides are selected from among the four natural nucleotides, namely adenine (A), thymine (T), cytosine (C), and guanine (G). The conversion of each digital subsequence to a bio-octet consists of converting bit 0 in an even position to nucleotide N1 = A, bit 1 in an even position to nucleotide N2 = T, bit 0 in an odd position to nucleotide N3 = C, and bit 1 in an odd position to nucleotide N4 = G.
[0192] The size of the longest assembly, called a track, was limited to 1024 bio - octets. File B contains more than 1024 bio - octets and thus is assembled onto multiple tracks. To be able to rearrange the tracks in the correct order, a binary barcode consisting of 4 bio - octets was added at the beginning of each track. A total of 256 to the fourth power (4,294,967,296) barcodes are available. The first track (track 0) contains barcode 0, which consists of 4 identical bio - octets 0 of the sequence "ACACACAC" (SEQ ID NO: 1107), followed by the first 1020 bio - octets of File B. The second track contains barcode 1, which consists of 3 octets 0 of the sequence "ACACACAC" and 1 bio - octet 1 of the subsequent sequence "ACACACAG", followed by the last 117 bio - octets of File B. To mark the end of the file (EOF), the last special bio - octet named EOF_B of the sequence "CAGTCTGT" is added at the end of track 1. Thus, track 0 contains 1024 bio - octets and track 1 contains 122 bio - octets.
[0193] To generate the DNA molecule corresponding to the two tracks, for example, it is possible to perform three Golden Gate assembly steps to assemble 1146 bio - octets (Figure 1). In step 1, the bio - octets are assembled from two libraries that contain all the bio - octets within a 2 - bio - octet block named BioblockX2. In step 2, a block named BioblockX64 is assembled, which contains 32 BioblockX2s. In step 3, a block named BioblockX1024 is assembled, which contains 16 BioblockX64s.
[0194] Results Two libraries named "Library A" and "Library B" are constructed, each containing all 256 possible biooctets. An EOF biooctet EOF_B is added to Library B, which thus consists of 257 biooctets. In both libraries, each biooctet is surrounded by a region containing an 11-nucleotide BsaI cleavage site and is contained within a double-stranded replicable plasmid. A variable region of the BsaI cleavage site, called the fusion site, is defined for each library. In Library A, each biooctet is surrounded by a GTAG fusion site upstream of the biooctet and a TGAC fusion site downstream of the biooctet. In Library B, each biooctet is surrounded by a TGAC fusion site upstream of the biooctet and a TCAG fusion site downstream of the biooctet. The compositions of Libraries A and B are given in Table 4, and their design is presented in Figure 2.
[0195]
Table 4A
[0196]
Table 4B
[0197]
Table 4C
[0198]
Table 4D
[0199]
Table 4E
[0200]
Table 4F
[0201] If there is a BsaI cleavage site within the library plasmid, it becomes possible to capture 1146 necessary bio-octets surrounded by the fusion site, alternating between library A and library B. The plasmids containing the necessary bio-octets from each library are digested by the restriction enzyme BsaI, thereby releasing the 1146 bio-octets surrounded by the fusion site. After capturing x = 1146 bio-octets surrounded by the cleavage site, they are assembled together in a fixed order in three steps.
[0202] In step 1, a block (BioblockX2) containing two bio-octets is assembled from the 1146 bio-octets surrounded by the fusion site in the double-stranded replicable plasmid. Each plasmid contains two internal BsaI cleavage sites in opposite directions, which enables the release of the fusion sites GTAG and TCAG upstream and downstream of BioblockX2, respectively, after BsaI cleavage. The fusion sites surrounding each bio-octet in libraries A and B enable the assembly of bio-octets from library A in the first position and bio-octets from library B in the second position. BioblockX2 is assembled from a set of 32 double-stranded replicable plasmids that contain the region surrounding BioblockX2 and a cleavage site for the type IIs restriction enzyme BsmBI (Figure 3). The variable region of the BsmBI cleavage site is defined for each of the 32 plasmids, which, thanks to a set of 33 fusion sites, defines the ordered positions for assembling the group of 32 BioblockX2s in step 2 of the assembly process (Table 5).
[0203] [Table 5]
[0204] A total of 573 plasmids are assembled in step 1. The 36-nucleotide sequences of the 573 BioblockX2 and their surrounding fusion sites correspond to SEQ ID NOs: 1 to SEQ ID NOs: 573. The first and last groups of 4 nucleotides correspond to the fusion sites flanking each BioblockX2. The groups of 4 nucleotides at positions 5-8, 17-20, and 29-32 correspond to the fusion sites from the bioblocks derived from libraries A and B.
[0205] As an example, BioblockX2_0 has the following sequence:
Chemical formula
[0206] In step 2, x = 573 BioblockX2 and their surrounding fusion sites are captured by digestion with the BsmBI restriction enzyme and assembled into BioblockX64 containing 32 BioblockX2 in a double-stranded replicable plasmid. Each plasmid contains two internal BsmBI cleavage sites in opposite directions, which allows for the release of the fusion site FS1_0 upstream and the fusion site FS1_32 downstream of BioblockX64, respectively, after BsmBI cleavage. BioblockX2 is assembled in the correct order due to the 33 fusion sites within a set of 16 double-stranded replicable plasmids that contain the region surrounding BioblockX64 and cleavage sites for the type IIs restriction enzyme BsaI (Figure 4). The variable region of the BsaI cleavage site is different for each of the 16 plasmids, which, due to the set of 17 fusion sites, defines the ordered positions for assembling the group of 16 bioblockX64 in step 3 of the assembly process (Table 6).
[0207]
Table 6
[0208] A total of 18 plasmids were assembled in step 2, 17 of which contain 32 BioblockX2s and the last one contains 29 BioblockX2s. The sequences of the 18 BioblockX64s and the fusion sites around them correspond to SEQ ID NO: 574 to SEQ ID NO: 591.
[0209] In step 3, x = 18 BioblockX64s and the fusion sites around them are captured by digestion with the BsaI restriction enzyme and assembled into BioblockX1024 containing 16 BioblockX64s in a double-stranded replicable plasmid. Each plasmid contains two internal BsaI cleavage sites in opposite directions, which enables the release of the fusion sites FS2_0 and fusion site FS2_16 (Table 6) upstream and downstream of BioblockX1024, respectively, after BsaI cleavage. BioblockX64 is assembled in the correct order thanks to the 17 fusion sites (Table 6).
[0210] In step 3, two plasmids corresponding to track 0 and track 1 are assembled (Figure 5). The sequences of the two BioblockX1024s correspond to SEQ ID NO: 592 and SEQ ID NO: 593. Track 0 contains 1024 biooctets (4 barcode biooctets and the first 1020 biooctets of file B). Track 1 contains 122 biooctets (4 barcode biooctets, the last 117 biooctets of file B, and a special EOF_B biooctet).
Claims
1. A nucleic acid-based data storage method for storing information, comprising: a) Restoring data in the form of a digital sequence formed from a plurality of bits, each bit having a value of 0 or 1; b) Subdividing the digital sequence into n digital subsequences, each digital subsequence comprising m bits, where m is in the range from 2 to 16; c) Converting each of the n digital subsequences into a bioblock consisting of a sequence of m nucleotides, wherein the digital subsequence resides in m bits assigned to positions from 0 to m-1, and said conversion of the digital subsequence into a bioblock comprises: - converting bits at even positions to a first nucleotide N1 when the bit has a value of 0, and to a second distinct nucleotide N2 when the bit has a value of 1, and - converting bits at odd positions to a third nucleotide N3 when the bit has a value of 0, and to a fourth distinct nucleotide N4 when the bit has a value of 1, - where N1, N2, N3, and N4 are distinct nucleotides; d) Constructing a plurality of x components, each individual component of said plurality of x components comprising at least one bioblock, said x components together comprising n bioblocks; e) Assembling said plurality of x components together, in one or more steps, in a fixed order. A nucleic acid-based data storage method for storing information, comprising the steps above.
2. The nucleotide-based data storage method according to claim 1, wherein the nucleotide is selected from the group of natural nucleotides consisting of adenine, guanine, cytosine, uracil, and thymine or unnatural nucleotides.
3. The nucleotide-based data storage method according to claim 1 or 2, wherein the x components are x DNA molecules, preferably x double-stranded DNA molecules.
4. In step (d), the construction of the plurality of x components, each containing at least one bioblock, comprises: - A step of selectively capturing x data storage nucleic acid molecules from at least one library of data storage nucleic acid molecules, each data storage nucleic acid molecule containing at least one bioblock surrounded by a region containing a cleavage site. - A step of cleaving each of the x data storage nucleic acid molecules, thereby releasing the at least one bioblock. The nucleotide-based data storage method according to any one of claims 1 to 3, comprising the above steps.
5. In step (d), the construction of the plurality of x components, each containing at least one bioblock, comprises: - A step of selectively capturing n data storage nucleic acid molecules from at least two libraries of data storage nucleic acid molecules, each data storage nucleic acid molecule in each library containing one bioblock surrounded by a region containing a cleavage site, and each library containing all possible bioblocks of m nucleotides. - A step of cleaving each of the n data storage nucleic acid molecules, thereby releasing the n bioblocks. The nucleotide-based data storage method according to claim 4, comprising the above steps.
6. The nucleotide-based data storage method according to claim 4 or 5, wherein the region containing the cleavage site contains 2 to 25 nucleotides.
7. Each of the regions surrounding the respective bioblocks contains a site for a restriction enzyme, and step (d) includes digesting each of the x data storage nucleic acid molecules with one or two restriction enzymes, the nucleic acid-based data storage method according to any one of claims 4 to 6.
8. Step (e) includes one or more assembly steps using overlap extension polymerase chain reaction (PCR), polymerase cycling assembly, sticky end ligation, biobrick assembly, golden gate assembly, Gibson assembly, recombinase assembly, ligase cycling reaction, template-directed ligation, in vivo assembly, or any other DNA assembly protocol, the nucleic acid-based data storage method according to any one of claims 1 to 7.
9. A data storage nucleic acid molecule comprising at least one bioblock, the bioblock consisting of a nucleic acid sequence of m nucleotides assigned to positions from 0 to m-1, - The bioblock is formed from at least two and at most four distinct nucleotides, - The nucleotides at even positions can be selected from the first and second nucleotides, and the nucleotides at odd positions can be selected from the third and fourth nucleotides, the first, second, third, and fourth nucleotides being distinct data storage nucleic acid molecules.
10. The data storage nucleic acid molecule according to claim 9, which is a double-stranded molecule, preferably a DNA molecule.
11. The data storage nucleic acid molecule according to claim 9 or 10, which is a plasmid, cosmid, fosmid, prokaryotic chromosome, or eukaryotic chromosome.
12. Each of the bioblocks is surrounded by a region containing a cleavage site, preferably surrounded by two sites for one restriction enzyme, the data storage nucleic acid molecule according to any one of claims 9 to 11.
13. A data storage nucleic acid molecule according to any one of claims 9 to 12, which is replicable.
14. A library comprising a plurality of data storage nucleic acid molecules according to any one of claims 9 to 12, wherein each of the data storage nucleic acid molecules of the library comprises one bio - block, each data storage nucleic acid molecule of the library comprises the same peripheral region containing a cleavage site, and the library comprises all possible bio - blocks of m nucleotides.
15. A nucleic acid - based data storage system comprising at least two libraries according to claim 14.
Citation Information
Patent Citations
Homopolymer encoded nucleic acid memory
US20190194739A1
Biocompatible nucleic acids for digital data storage
WO2021064095A1
Nucleic acid-based data storage
US20180137418A1