A DNA information storage method based on natural and unnatural bases

Through the DNA storage method based on natural and non-natural bases, the detection of base combinations is solved by mass spectrometer, and the problems of low encoding density and sequencing errors in existing DNA storage technologies are achieved, and efficient DNA information storage and decoding are achieved.

CN116030895BActive Publication Date: 2025-08-29SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211594766.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-13
Publication Date
2025-08-29
Estimated Expiration
2042-12-13

AI Technical Summary

Technical Problem

The existing DNA information storage technology is limited by the natural base encoding mode, has low encoding density, errors in synthesis and sequencing, and the synthesis length and sequence limitations make it difficult to decode information.

Method used

DNA storage methods based on natural and non-natural bases are adopted, and mass spectrometry sequencing analysis is used to map data information with DNA information one by one by designing a coding table, and base combinations are detected by mass spectrometer to directly store and read DNA information, avoiding synthesis and sequencing steps.

Benefits of technology

It improves the encoding density, breaks through the theoretical limits of 4-base encoding, solves the limitations brought by synthesis and sequencing, reduces costs, improves coding efficiency and freedom, and reduces sequencing errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116030895B_ABST
    Figure CN116030895B_ABST
Patent Text Reader

Abstract

The present invention relates to a DNA information storage method based on natural and unnatural bases, comprising the following steps: extracting data information to be stored; designing a coding table, wherein the coding table is formed by mapping DNA information to data information units; sequentially confirming the DNA information corresponding to the split data information in the coding table according to the split data information; and sequentially arranging the DNA information in the holes of a physical medium to obtain a DNA information storage carrier; when reading, using mass spectrometry to read the DNA information and decode it. The storage and reading method of the present invention can fully utilize natural and unnatural bases to achieve multi-base encoding, with a density expected to exceed 8 bits / nt; while avoiding the trouble of synthesizing DNA from scratch for storage information; and also abandoning the practice of using a sequencer to identify recorded information, innovatively using a mass spectrometer to read the base type, overcoming the errors and limitations brought about by the sequencing process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data storage, and in particular relates to a DNA information storage method based on natural and non-natural bases. Background Art

[0002] With the advancement of network technology, the exchange and generation of information have exploded. With such a massive amount of information, the question of how to store it has become a pressing issue. Existing silicon-based storage technology can no longer meet this growing demand, and researchers have turned their attention to other materials, with deoxyribonucleic acid (DNA) being a particularly hot topic. DNA, as an information storage medium, offers advantages such as high storage density, long and stable storage times, and low energy consumption. However, in the early stages of DNA research, there are still some issues that need to be addressed. For example, current methods do not fully utilize the properties of bases, especially non-natural bases, and the coding density is limited to 2 bits / nt. Furthermore, current methods require the bases to be synthesized from scratch each time information is stored, which is costly.

[0003] Furthermore, DNA information storage technology is severely limited by DNA synthesis and sequencing techniques, and errors introduced during these processes are difficult to eliminate. Therefore, how to mitigate or mitigate errors caused by DNA synthesis and sequencing has become a research hotspot in the field of DNA information storage. Patent CN113066534A proposes using a quaternary encoding system based on the four bases ATCG to read data from a biochip, employing a sequencing-by-synthesis method. This method measures the DNA sequence on the chip to retrieve information, and while biochips can be stored at low temperatures for long periods, it still fails to address the issue of errors introduced by sequencing.

[0004] Existing DNA information storage technologies primarily rely on DNA synthesis and sequencing, which often introduce errors during the synthesis and sequencing process, making it difficult to decode the information. Furthermore, these processes impose certain restrictions on DNA sequences, such as requiring the sequence to contain no consecutive repeating bases (i.e., homopolymer length ≤ 4), a GC content of 40% to 60%, and the absence of complementary secondary structures between sequences. Summary of the Invention

[0005] Existing DNA information storage technologies, both writing and reading information, are subject to various technical limitations. For example, information writing currently relies solely on natural DNA encoding, underutilizing the properties of unnatural bases and limiting encoding density. During oligo synthesis, the length of the synthesized sequence is limited to 200 base pairs, and homopolymers exceeding 3 base pairs cannot be present in the synthesized sequence. Furthermore, during sequencing, there are even more limitations, preventing the presence of large numbers of repeats, and the sequencer itself is subject to significant sequencing errors. Furthermore, de novo DNA synthesis is expensive. These issues not only limit the randomness of information encoding but also distort the results, leading to decoding failures. To address these issues, the present invention, based on natural and unnatural bases and leveraging the advantages of mass spectrometry sequencing analysis, establishes a DNA information storage method. This storage process does not require a synthesis step, thus avoiding the various limitations of sequence synthesis. This method not only meets the needs of information storage but also breaks the inherent framework of existing DNA information storage. Furthermore, based on natural and unnatural base storage, the encoding density is high, providing a new direction for the development of DNA information storage technology.

[0006] One aspect of the present invention provides a DNA storage method based on natural and unnatural bases, the DNA storage method comprising the following steps:

[0007] S11): extracting data information of computer data information to be stored;

[0008] S12): Designing a coding table, wherein the coding table is formed by one-to-one mapping of DNA information and data information units;

[0009] S13) splitting the data information in step S11) into data information units, and sequentially confirming the DNA information corresponding to the split data information units in the coding table obtained in step S12);

[0010] S14) obtaining the deoxyribonucleotides or polynucleotide sequences consisting of two or more deoxyribonucleotides corresponding to the DNA information confirmed in step S13), and arranging them sequentially in different wells of the chip to obtain a DNA information storage carrier;

[0011] The DNA information is a deoxyribonucleotide or a polynucleotide sequence consisting of two or more deoxyribonucleotides, or a combination of any two of the deoxyribonucleotides or the polynucleotide sequence consisting of two or more deoxyribonucleotides;

[0012] In the coding table, the molecular weight of each DNA information is different, and the molecular weight difference between each is no less than 10.

[0013] Furthermore, in step S11), the computer data information is text, numbers, pictures, videos, programs, and audio.

[0014] Further, in step S11), the data information is character information, RGB information, binary data information, octal data information, hexadecimal data information, or decimal data information.

[0015] Furthermore, in the coding table, the molecular weight of each DNA information is different, and the difference in molecular weight between each two is not less than 10.

[0016] Furthermore, the deoxyribonucleotide is a natural deoxyribonucleotide or a non-natural deoxyribonucleotide, and the non-natural deoxyribonucleotide is a deoxyribonucleotide with modified bases.

[0017] Furthermore, the deoxyribonucleotides in the polynucleotide sequence composed of two or more deoxyribonucleotides are natural deoxyribonucleotides or non-natural deoxyribonucleotides.

[0018] Furthermore, the polynucleotide sequence composed of two or more deoxyribonucleotides can have different molecular weights by adjusting the combination of the types and quantities of deoxyribonucleotides in the polynucleotide sequence.

[0019] Furthermore, the 32-1024 mapping relationships formed in the coding table are, for example, 32, 64, 128, 256, 512, and 1024 types.

[0020] Furthermore, the data information unit is to divide the data information into different units for recording information. In some instances, the data information unit is RGB pixel information corresponding to a 4-bit or 8-bit binary number, a single character, or a single pixel.

[0021] Furthermore, the bases in the nucleic acid sequence are modified or unmodified. Furthermore, the modified bases are at least one of DBCO modification, AMCA modification, thio modification, amino modification, biotin modification, digoxigenin modification, phosphate group, and sulfhydryl group modification.

[0022] Furthermore, there are 32 types of polynucleotide sequences in the coding table, and 128 different combined polynucleotide sequences are formed by combining two polynucleotide sequences.

[0023] Furthermore, the coding table maps DNA information to 4-bit or 8-bit binary numbers one by one.

[0024] Furthermore, the coding table maps DNA information to the ASCII code table one by one.

[0025] Furthermore, the coding table maps DNA information to 128 characters one by one.

[0026] Furthermore, the coding table maps DNA information to RGB pixel information one by one.

[0027] In some specific embodiments, the DNA information in the coding table is 128 or 256 different nucleotides with modified bases, and 32 or 64 different types of modifications are performed on A, T, C, and G, respectively, to obtain a total of 128 or 256 different nucleotides with modified bases.

[0028] Furthermore, in the coding table, one of the DNA information is mapped one by one to the color information in the RGB pixel information, that is, R\G\B, and the other of the DNA information is mapped one by one to the numbers 0-255 in the color information of the RGB pixel information, and the combination of the two DNA information forms and One-to-one mapping relationship between RGB pixel information.

[0029] In an embodiment comprising only four natural bases, the coding table comprises 128 combined polynucleotide sequences, the lengths of the polynucleotide sequences are divided into 8 groups, the lengths of the 8 groups of polynucleotide sequences are successively extended, each comprising 10-24 bases of nucleotides, with 4 polynucleotide sequences in each group, and the types and or quantities of bases in the 4 polynucleotide sequences in each group are different.

[0030] Another aspect of the present invention provides a method for reading information from a DNA information storage carrier obtained by the above-mentioned DNA storage method, the information reading method comprising the following steps:

[0031] S21) performing mass spectrometry on the sequences to be tested at different well positions in the DNA information storage carrier to obtain molecular weight information of the DNA information in each well position;

[0032] S22) analyzing the base combination information based on the molecular weight information of the DNA information to confirm the DNA information corresponding to the different pore positions;

[0033] S23) decoding the data information unit based on the DNA information obtained in step S22) and the coding table in step S12);

[0034] S24) splicing the data units obtained in S23) and decoding the data to obtain stored computer data.

[0035] Furthermore, in step S21), the mass spectrometry detection method is MALDI mass spectrometry sequencing.

[0036] Furthermore, in step S21), the mass spectrometry detection method includes the following steps:

[0037] S211) performing enzyme digestion and / or purification on the sequence to be tested;

[0038] S212) The purified fragments are subjected to mass spectrometry to obtain molecular weight.

[0039] Furthermore, the purification method in step S211) is ethanol precipitation, microdialysis or MillporeZiptip microchromatography.

[0040] Furthermore, computer data information is data information that can exist on a computer, preferably selected from pictures, texts, programs, audio, and video.

[0041] Another aspect of the present invention provides a method for storing and decoding DNA information based on natural and unnatural bases, the method comprising:

[0042] The DNA storage method based on natural and unnatural bases as described above comprises the following steps:

[0043] S11): extracting binary information of computer data information to be stored;

[0044] S12): Designing a coding table, wherein the coding table is formed by one-to-one mapping of DNA information and data information units;

[0045] S13): splitting the data information in step S11) into data information units, and confirming the DNA information corresponding to the split data information units in the coding table obtained in step S12);

[0046] S14) obtaining the deoxyribonucleotides or polynucleotide sequences consisting of two or more deoxyribonucleotides corresponding to the DNA information confirmed in step S13), and arranging them sequentially in different wells of the chip to obtain a DNA information storage carrier;

[0047] The DNA information is a deoxyribonucleotide or a polynucleotide sequence consisting of two or more deoxyribonucleotides, or a combination of any two of the deoxyribonucleotides or the polynucleotide sequence consisting of two or more deoxyribonucleotides;

[0048] In the coding table, the molecular weight of each DNA information is different, and the molecular weight difference between each two is not less than 10;

[0049] And a method for reading information from a DNA information storage carrier obtained by the DNA storage method as described above, the information reading method comprising the following steps:

[0050] S21) performing mass spectrometry on the sequences to be tested at different well positions in the DNA information storage carrier to obtain molecular weight information of the DNA information in each well position;

[0051] S22) analyzing the base combination information based on the molecular weight information of the DNA information to confirm the DNA information corresponding to the different pore positions;

[0052] S23) decoding the data information unit based on the DNA information obtained in step S22) and the coding table in step S12);

[0053] S24) splicing the data units obtained in S23) and decoding the data to obtain stored computer data.

[0054] Another aspect of the present invention provides a DNA information storage device based on natural and unnatural bases, the device comprising:

[0055] A data information extraction unit, configured to extract computer data to be stored and convert the computer data to be stored into data information corresponding to the information;

[0056] A data information and DNA information conversion unit, configured to split or assemble the data information sequence and convert it into DNA information according to a preset mapping relationship;

[0057] A synthesis and storage unit, used to synthesize the data information and the nucleic acid sequence converted by the DNA information conversion unit, and store the deoxyribonucleotides corresponding to the DNA information, a polynucleotide sequence consisting of two or more deoxyribonucleotides, or a combination thereof in different wells of the storage unit chip in sequence;

[0058] The mapping relationship is a one-to-one mapping relationship between DNA information and data information units;

[0059] The DNA information is a deoxyribonucleotide or a polynucleotide sequence consisting of two or more deoxyribonucleotides, or a combination of any two of the deoxyribonucleotides or the polynucleotide sequence consisting of two or more deoxyribonucleotides;

[0060] In the mapping relationship, the molecular weight of each DNA information is different, and the molecular weight difference between each two is no less than 10.

[0061] The data information and DNA information conversion unit may include a DNA information encoding unit, a DNA information and data information matching unit, and a DNA information conversion unit. The DNA information encoding unit is used to record the combination of base types and quantities corresponding to each type of DNA information. The DNA information and data information matching unit is used to call different DNA information in the DNA information encoding unit for one-to-one matching and correspondence with data information units. The DNA information conversion unit is used to convert the digital information in the data information extraction unit into DNA information one by one based on the information of the DNA information and data information matching unit.

[0062] Another aspect of the present invention provides a decoding device based on natural and unnatural base nucleic acid storage, the device comprising:

[0063] A reading unit, used to detect the sequence to be tested stored in the synthesis and storage unit through a mass spectrometer and confirm its DNA information based on the molecular weight;

[0064] A DNA information and data information conversion unit, configured to convert the DNA information obtained by the reading unit into data information according to a preset mapping relationship, i.e., a mapping relationship between the same data information and DNA information in the DNA information storage device;

[0065] A computer data output unit, used to convert the DNA information and the data information obtained by the data information conversion unit into stored computer data;

[0066] The reading module includes a mass spectrometer for performing mass spectrometry detection on the nucleic acid sequence in each well position in the chip for detecting stored information, and may also include a unit for purifying and / or enzymatically hydrolyzing the nucleic acid sequence in each well position.

[0067] Another aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of the above-mentioned DNA storage method based on natural and non-natural bases or the information reading method of obtaining a DNA information storage carrier using the above-mentioned storage method are implemented.

[0068] Another aspect of the present invention provides a computer device comprising a memory and a processor, wherein a computer program that can be run on the processor is stored in the memory, and when the processor executes the computer program, the steps of the above-mentioned DNA storage method based on natural and non-natural bases or the information reading method of obtaining a DNA information storage carrier using the above-mentioned DNA storage method are implemented.

[0069] Beneficial effects

[0070] 1) The present invention's solution can be applied to any modified or unmodified bases, as well as unnatural base pairs. Furthermore, as the number of available bases increases, the efficiency of DNA storage also increases. By using multiple unnatural bases, each base can encode an 8-bit binary code, resulting in a logical encoding capacity of 8 bits per nt (bit / nt), surpassing the theoretical limit of 2 bits per nt for 4-base encoding.

[0071] 2) The encoding method of the present invention is designed according to the type and number of bases, without considering issues such as repetition and secondary structure in the sequence.

[0072] 3) The present invention solves a series of problems caused by synthesis and sequencing from the source, directly abandoning these two technologies and instead adopting a method of fixed-point base storage and mass spectrometer detection.

[0073] 4) The prior art uses four natural bases to directly map to a quaternary code, but the coding efficiency is still low. The present invention can encode any natural and unnatural bases, and base combinations of different numbers and lengths can also be encoded as a variable of the code, which solves the limitations of coding efficiency and coding methods. At the same time, by introducing a large number of unnatural bases, the coding efficiency and coding freedom are greatly improved. Since there are many types of commercially available unnatural bases and they are highly commercialized, they can fully meet the coding requirements.

[0074] 5) Due to the special way of information reading and encoding, there is no need to synthesize DNA sequences. Instead, grid storage combined with microfluidic technology is used. Only short sequences need to be synthesized, and there is no need to synthesize long sequences, thus breaking the shackles brought by synthesis.

[0075] 6) The storage and reading method of the present invention abandons the existing practice of using sequencers to identify the nucleic acid sequence of recorded information. It innovatively uses a mass spectrometer to read the base type, overcoming the errors and limitations brought by the sequencing process. It can form chain types with different molecular weights by combining different bases in different ratios. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] Figure 1 Schematic diagram of the DNA storage method based on natural and unnatural base matrix spectrum decoding of the present invention.

[0077] Figure 2 This is a schematic diagram of Example 1, part of the specific embodiments of the present invention, which is a method for directly encoding DNA storage based on spectrum decoding of natural and unnatural base matrices.

[0078] Figure 3 Schematic diagram of the information reading method of the DNA information storage carrier of the present invention.

[0079] Figure 4 Schematic diagram of mass spectrometry detection of the present invention.

[0080] Figure 5 This is a schematic diagram of the structure of the encoding device provided in Example 4 of the present invention.

[0081] Figure 6 This is a schematic diagram of the structure of a decoding device provided in Example 4 of the present invention.

[0082] Figure 7 A schematic diagram of the structure of the terminal device provided in Example 4 of the present invention.

[0083] Figure 8 It is the overall technical flow chart of the present invention.

[0084] Figure 9 Schematic diagram of a picture to be encoded according to the present invention. DETAILED DESCRIPTION

[0085] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below, but it should not be understood as limiting the scope of implementation of the present invention.

[0086] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0087] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0088] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0089] References to "one embodiment" or "some embodiments" described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the phrases "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in another way. Some specific embodiments of the present invention are described below in conjunction with the accompanying drawings.

[0090] In existing DNA information storage technologies, both writing and reading information are subject to various technical limitations, especially those based on sequencing technology, which greatly restricts information storage. To address the above problems, the present invention innovatively establishes a DNA storage method based on natural and non-natural bases, replacing sequencing methods with mass spectrometry, overturning traditional information storage and decoding methods. Moreover, the storage process of the present invention does not require excessive synthesis, or even a synthesis step, and therefore does not face the various limitations of sequence synthesis. The method of the present invention is described below in the form of embodiments with reference to the accompanying drawings.

[0091] Example 1 DNA storage method based on natural and unnatural base matrix spectrum decoding:

[0092] S11) extracting data information of computer data information to be stored;

[0093] The computer data information to be stored may be data in any format, such as text, numbers, pictures, videos, programs, audio, etc. In a specific embodiment of the present invention, the data information may be RGB information, binary data information, octal data information, hexadecimal data information, decimal data information, or character information.

[0094] In some specific embodiments, the computer data information to be stored is converted into digital information by any method known in the art.

[0095] In other specific implementations, the image information can be converted into RGD information. The RGD information is RGD pixel information. The image information is converted into RGD pixel information of different pixels. The RGD pixel information consists of color type, i.e., R\G\D information, and color intensity 0-255.

[0096] In some specific implementation schemes, if the information to be stored is character information, the conversion may not be performed, and the character information may be directly used as data information for subsequent conversion and encoding.

[0097] S12) Designing a coding table, wherein the coding table is formed by one-to-one mapping of DNA information and data information units;

[0098] The DNA information is a deoxyribonucleotide or a polynucleotide sequence consisting of two or more deoxyribonucleotides, or a combination of any two of the deoxyribonucleotides or the polynucleotide sequence consisting of two or more deoxyribonucleotides;

[0099] In the coding table, the molecular weight of each DNA information is different, and the molecular weight difference between each is not less than 5.

[0100] In some specific embodiments, illustratively, the molecular weight difference between two DNA information is greater than 10, for example, any number between 10-100.

[0101] In some specific technical solutions, the 32-1024 mapping relationships in the coding table are, for example, 32, 64, 128, 256, 512, and 1024.

[0102] In some specific technical solutions, the DNA information and the data information in the data information coding table are illustratively 4-bit binary data or 8-bit binary data, and are the same as the number of DNA information in the coding table, forming a one-to-one mapping relationship.

[0103] In some specific technical solutions, the DNA information and the data information in the data information coding table are exemplarily character information, and the number is the same as the DNA information in the coding table, forming a one-to-one mapping relationship.

[0104] In some specific technical solutions, the data information in the DNA information and data information coding table is illustratively RGB pixel information, and forms a one-to-one mapping relationship with the DNA information in the coding table.

[0105] In some specific technical solutions, the DNA information may be natural deoxyribonucleotides or non-natural deoxyribonucleotides, and the non-natural deoxyribonucleotides are base-modified deoxyribonucleotides.

[0106] In some specific technical solutions, modified bases refer to at least one or a combination of two or more of the following: DBCO modification, AMCA modification, thio modification, amino modification, biotin modification, digoxigenin modification, phosphate group, sulfhydryl group, amino group, NHBOC modification, Fmoc modification, carboxylic acid modification, Mal modification, NHS modification, azide modification, Cy3 / Cy5 / Cy7 modification, THP modification, benzyl modification, propynyl modification, bromo modification, tert-butyl propionate modification, tert-butyl acetate modification, methyl modification, biotin modification, pentafluorophenol modification, and sulfonate modification. Several commercially available modified bases are available, and those skilled in the art can select based on molecular weight requirements. Detailed modification types can be found on the official websites of commercial companies.

[0107] Exemplarily, the unnatural base is selected from 2-aminoadenin-9-yl, 2-aminoadenine, 2-F-adenine, 2-thiouracil, 2-thiothymine, 2-thiocytosine, 2-propyl and alkyl derivatives of adenine and guanine, 2-amino-adenine, 2-amino-propyl-adenine, 2-aminopyridine, 2-pyridone, 2'-deoxyuridine, 2-amino-2'-deoxyadenosine, 3-deazaguanine, 3-amino-3'-deoxyuridine ... -deazaadenine, 4-thiouracil, 4-thiothymine, uracil-5-yl, hypoxanthine-9-yl (I), 5-methyl-cytosine, 5-hydroxymethylcytosine, xanthine, hypoxanthine, 5-bromo- and 5-trifluoromethyluracil and cytosine; 5-halouracil, 5-halocytosine, 5-propynyl-uracil, 5-propynylcytosine, 5-uracil, 5-substituted, 5-halogenated, 5-substituted pyrimidines, 5 -hydroxycytosine, 5-bromocytosine, 5-bromouracil, 5-chlorocytosine, chlorinated cytosine, cyclocytosine, cytosine arabinoside, 5-fluorocytosine, fluoropyrimidine, fluorouracil, 5,6-dihydrocytosine, 5-iodocytosine, hydroxyurea, iodouracil, 5-nitrocytosine, 5-bromouracil, 5-chlorouracil, 5-fluorouracil and 5-iodouracil, 6-alkyl derivatives of adenine and guanine, 6-azapyrimidine, 6-azo-uracil, 6-azocytosine, azacytosine, 6-azo-thymine, 6-thioguanine, 7-methylguanine, 7-methyladenine, 7-deazaguanine, 7-deazaguanosine, 7-deaza-adenine, 7-deaza-8-azaguanine, 8-azaguanine, 8-azaadenine, 8-halo-, 8-amino-, 8-thiol-, 8-thioalkyl- and 8-hydroxy-substituted adenine and guanine;N4-ethylcytosine, N-2 substituted purine, N-6 substituted purine, O-6 substituted purine, those that increase the stability of duplex formation, universal nucleic acid, hydrophobic nucleic acid, promiscuous nucleic acid, size-extended nucleic acid, fluorinated nucleic acid, tricyclic pyrimidine, phenoxazine cytidine ([5,4-b][1,4]benzoxazin-2(3H)-one), phenothiazine cytidine (1H-pyrimido[5,4-b][1,4]benzothiazin-2(3H)-one), G-clamp, phenoxazine cytidine (9-(2-aminoethoxy) )-H-pyrimido[5,4-b][1,4]benzoxazin-2(3H)-one), carbazole cytidine (2H-pyrimido[4,5-b]indol-2-one), pyridoindole cytidine (H-pyrido[3',2':4,5]pyrrolo[2,3-d]pyrimidin-2-one), 5-fluorouracil, 5-bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xanthine, 4-acetylcytosine, 5-(carboxyhydroxymethyl)uracil, 5-carboxymethylaminomethyl-2-thiouridine , 5-carboxymethylaminomethyluracil, dihydrouracil, β-D-galactosyl quercetin, inosine, N6-isopentenyl adenine, 1-methylguanine, 1-methylinosine, 2,2-dimethylguanine, 2-methyladenine, 2-methylguanine, 3-methylcytosine, 5-methylcytosine, N6-adenine, 7-methylguanine, 5-methylaminomethyluracil, 5-methoxyaminomethyl-2-thiouracil, β-D-mannosyl quercetin, 5'-methoxycarboxymethyluracil, 5-methoxyuracil Pyrimidine, 2-methylthio-N6-isopentenyl adenine, uracil-5-oxoacetic acid, whibutoxoside, pseudouracil, quercetin, 2-thiocytosine, 5-methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, uracil-5-oxoacetic acid methyl ester, uracil-5-oxoacetic acid, 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil, (acp3)w and 2,6-diaminopurine, as well as purine or pyrimidine bases replaced by heterocycles.

[0108] In some specific embodiments, the deoxyribonucleotides in the polynucleotide sequence composed of two or more deoxyribonucleotides in the DNA information are natural deoxyribonucleotides or non-natural deoxyribonucleotides. The polynucleotide sequence composed of two or more deoxyribonucleotides can be modified to have different molecular weights by adjusting the type and quantity of deoxyribonucleotides in the polynucleotide sequence. For example, a combination encoding comprising four natural bases or a mixed encoding of natural and non-natural bases can be used.

[0109] Since the present invention does not require sequencing during the decoding process but uses mass spectrometry for determination, the different masses of chain types composed of ribonucleotides or deoxyribonucleotides in the chain types can be distinguished by mass spectrometry data.

[0110] Exemplarily, the coding table contains 128 combined polynucleotide sequences, and the lengths of the polynucleotide sequences are divided into 8 groups. The lengths of the 8 groups of polynucleotide sequences are successively extended, each containing 10-24 base nucleotides, 4 chain types in each group, and the base types and or quantities in the 4 chain types in each group are different.

[0111] For example, four deoxyribonucleotides can be modified with different modifying groups. Under the action of different modifying groups, for example, 32 modifications can be made to each deoxyribonucleotide, resulting in 128 different deoxyribonucleotides, while 64 modifications can result in 256 different deoxyribonucleotides. If more nucleotide types are added, the number of encoded types can be doubled on the basis of the above.

[0112] In some specific embodiments, a combination of two polynucleotides, a combination of two non-natural deoxyribonucleotides, or a combination of a polynucleotide and a non-natural deoxyribonucleotide can also be used.

[0113] For example, the types of DNA information in the coding table can be increased in the form of combinations, thereby improving coding efficiency. For example, by preparing 32 different nucleotide sequences and combining them in pairs, a maximum of 1024 combinations can be obtained, of which 128 or 256 can be selected for designing the coding table.

[0114] In a specific embodiment, the DNA information in the coding table can be confirmed by the following method. First, in this embodiment, only polynucleotides composed of natural nucleotides are used, that is, polynucleotide sequences of lengths between 10 and 24 composed of deoxyribonucleic acid of A, T, C, and G. A total of 8 length gradients are set, and each length gradient has four different base contents. The four chains of the same length can be arbitrarily combined to form 16 different combinations (as shown in Table 1, for example, the 16 combinations in the first length gradient are a1a1, a1a2, a1a3, a1a4, a2a1, a2a2, a2a3, a2a4, a3a1, a3a2, a3a3, a3a4, a4a1, a4a2, a4a3, a4a4), so there are a total of 8×16=128 combinations. For example, a1a1 corresponds to A4T2C2G2.

[0115] Table 1 Comparison table of polynucleotide length gradient and base composition

[0116]

[0117]

[0118]

[0119] They correspond to the 128 elements of the ASCII code table, and each element is represented by an 8-bit binary number.

[0120] It is understood that the number of nucleotides in a polynucleotide sequence can be adjusted. If the upper limit of the number range selected for the number of polynucleotide sequences is smaller, the number of bases required will be smaller. The above exemplary scheme selects four bases, and if the number of bases involved in encoding is increased, the number of bases required will also be greatly reduced.

[0121] The polynucleotide sequences are mapped one by one to the data information, and a coding table is formed. For example, the polynucleotide sequences described in Table 1 above, i.e., DNA information and 8-bit binary information are mapped one by one to form a DNA information and binary data coding table encoding 128 types of information. The four vertical digits and the four horizontal digits together form an 8-bit binary number. Each group of 8-bit binary numbers corresponds to a group of DNA information, as shown in Table 2:

[0122] Table 2 DNA information and data information coding table

[0123]

[0124]

[0125] For example: 00000000 corresponds to a1a1, and the polynucleotide sequence combination corresponding to a1a1 is a sequence composed of A8T4C4G4.

[0126] In another specific embodiment, four deoxyribonucleotides can be modified with different modification groups. Under the action of different modification groups, for example, 32 modifications are performed on each deoxyribonucleotide, resulting in 128 different deoxyribonucleotides, while 64 modifications are performed, resulting in 256 different deoxyribonucleotides.

[0127] The nucleotides with modified bases are matched with the data information and a coding table is formed.

[0128] Specifically, taking the above 256 nucleotides with modified bases as an example, by matching them with 8-bit binary information, a matching table of 256 nucleotides with modified bases and binary data can be obtained, as shown in Table 3:

[0129] Table 3 Matching table of nucleotides with modified bases and data information

[0130]

[0131]

[0132]

[0133] In Table 3, A1 represents the first modified deoxyadenine nucleotide, and so on, A, T, C, G+numbers represent the Nth modified deoxyribonucleotide.

[0134] For example, 32 different modified bases of ribonucleotides can be selected, with a total of 128 different modified bases, which are directly matched with characters. The encoding table is shown in Table 4 below:

[0135] Table 4

[0136]

[0137]

[0138] In some specific embodiments, direct encoding can also be used to directly map each combination to a character, thereby directly converting a text file into a DNA information file. For example, the DNA information in the encoding table in Table 4 above can be mapped one-to-one with the 128 types of character information in the ASII table to form a coding table. This coding table can then be used to directly convert the character information in the text into DNA information.

[0139] S13) Split the data information in step S11), and sequentially confirm the DNA information corresponding to the split data information using the DNA information obtained in step S12) and the data information coding table.

[0140] Exemplarily, when the data information mapped in the above DNA information and data information coding table is an 8-bit binary number, the data information in step S11) is split into 8-bit binary numbers, and the DNA information corresponding to the 8-bit binary data is searched in the above DNA information and data information coding table in turn.

[0141] S14) obtaining the chain types confirmed in step S13) and sequentially arranging different wells of the chip to obtain a DNA information storage carrier;

[0142] Based on the DNA information confirmed in step S13), the nucleotide or polynucleotide sequence can be obtained by directly using commercially available nucleotides or synthesizing them for different polynucleotide sequence types. Alternatively, large-scale synthesis and storage can be performed based on the types in the coding table in step S12), and then extracted during storage.

[0143] The confirmed nucleotide or polynucleotide sequences can be directly arranged in sequence in different wells of the chip without further ligation, which can reduce the number of synthesis steps.

[0144] The method of the present invention is demonstrated below using several specific data to be stored. The first specific implementation case is storing English words and characters: Hello world!

[0145] The information to be stored is the string "Hello world!". First, the 12 characters in the string are converted sequentially into 12 8-bit binary numbers. These binary numbers are then converted into different polynucleotide sequence combinations according to the encoding tables in Tables 1 and 2. The chains corresponding to the different characters are synthesized and sequentially placed in different wells of the chip. The resulting chip containing all the information becomes the storage medium for this information.

[0146] In addition, the above characters can also be encoded in the form of direct encoding, such as Figure 2 The characters shown are mapped to different DNA sequences or ribonucleotides of non-natural bases with different modification groups, as shown in Table 4 above, and are encoded and stored.

[0147] In a second specific embodiment, the information to be stored is a text file "wssnt10.txt" encoded in Goldman's 2012 article "Toward spractical, high-capacity, low-maintenance information storage in synthesized DNA".

[0148] The wssnt10.txt file is as follows:

[0149]

[0150] First, the 107,738-byte text is directly encoded into a base combination information file. For example, the polynucleotide combination corresponding to "!" is e1e2, which means the base sequence is "4A4T4C6G4A4T4C6G". After conversion to DNA information, an example fragment is shown below:

[0151]

[0152] Further converted to chain types displayed by base type and number, an example fragment is shown below:

[0153]

[0154] According to the obtained DNA information, a biochip is prepared for DNA information storage. In a third specific embodiment, the information to be stored is an image.

[0155] First, convert the image information into binary data. The following is an example snippet:

[0156]

[0157]

[0158] The generated binary file is encoded according to the encoding tables in Tables 1 and 2 above. For example, the polynucleotide combination corresponding to "0100 0001" is e1e2, that is, the base sequence is "4A4T4C6G4A4T4C6G".

[0159] An exemplary fragment after conversion to a polynucleotide is shown below:

[0160]

[0161] Further converted to DNA information displayed in terms of base types and numbers, an example fragment is shown below:

[0162]

[0163] Regarding base encoding, the first three specific implementation cases demonstrate the four natural bases ATCG. However, those skilled in the art can select as needed when employing the method of the present invention for encoding. Polynucleotide sequences, or even combinations thereof, can also be used, as can a single non-natural base deoxyribonucleotide, not limited to natural base deoxyribonucleotides. This innovative use of mass spectrometry as a detection method enables mass spectrometry to distinguish not only natural bases but also non-natural bases, achieving a function that DNA sequencing cannot.

[0164] The following provides several specific implementation plans using non-natural base amino deoxyribonucleic acid as an example.

[0165] The fourth specific implementation case is demonstrated by storing the first chapter of the original version of Pride and Prejudice. The text file size is 4,501 bytes. The binary file size generated by binary encoding text conversion is 36,008 bytes, as shown in the following excerpt:

[0166]

[0167] In this embodiment, the above text is first converted into binary data, and some excerpts of binary data information are shown below:

[0168]

[0169] 2. Convert the binary data into a polynucleotide sequence file of 13,478 bytes according to the encoding table (Table 3). The following is an excerpt of the encoded sequence information:

[0170]

[0171] The nucleotides corresponding to the encoded DNA information are stored in different wells of the biochip according to their order.

[0172] It can be seen from this embodiment that each nucleotide can encode an 8-bit binary code, so the logical coding capacity of the present invention can reach 8 bits / nt, which has broken through the theoretical limit of the existing four-base coding.

[0173] The information to be stored in the fifth embodiment is still the text of Chapter 1 of Pride and Prejudice. However, unlike the fourth embodiment, in this embodiment, the text information is not converted into binary data. Instead, the characters in the text information are directly used as the encoded information data.

[0174] Using the above coding table, Table 4 can be encoded to directly obtain the corresponding DNA information, as shown below:

[0175]

[0176] By comparing the coding efficiency of the above two specific embodiments, it can be seen that although the coding density of direct coding (fifth embodiment) and indirect coding (fourth embodiment) is 8 bit / nt, direct coding uses fewer types of modified bases and is simpler.

[0177] The encoded nucleotides are sequentially stored in different wells of the chip to obtain a storage medium.

[0178] In the sixth embodiment, the information to be stored is a picture file. For the picture information to be stored, see Figure 9 , Figure 9 The original image is a color image. Its RGB information is as follows:

[0179] RGB format information:

[0180]

[0181] The image information is directly encoded into nucleotide information. The encoded sequence information is as follows:

[0182]

[0183] The nucleotides are sequentially stored in different wells of the chip to obtain a storage medium.

[0184] The image information can also be divided into different pixel points according to the pixel, and the RGB pixel information of different pixel points is used as the data information. A coding table is further designed, which contains the DNA information corresponding to different colors, i.e. RGB, at depths of 0-255, and is encoded and stored accordingly.

[0185] Example 2 DNA information decoding method based on natural and unnatural base matrix spectrum decoding

[0186] S21) performing mass spectrometry on the sequences to be tested at different well positions in the DNA information storage carrier to obtain molecular weight information of the DNA information in each well position;

[0187] Unlike the prior art which uses a sequencer and needs to sequence the order of bases, the present invention uses MALDI mass spectrometry sequencing to read out the base combination in the chain type at each position on the chip.

[0188] During the information recording process, different nucleotide or polynucleotide sequences or their combinations are placed in different chip wells, and mass spectrometry detection is performed on the sequences to be tested in different wells, and the peaks of the substances are printed out in the order of different wells.

[0189] The mass spectrometry results show the molecular weight of the sequence to be tested and its fragment peaks after bombardment. Based on these data, the type of DNA information to be sequenced in the coding table can be confirmed.

[0190] Before mass spectrometry sequencing, it can also include:

[0191] S211) performing enzyme digestion and purification on the sequence to be tested;

[0192] S212) The purified fragments are subjected to mass spectrometry to obtain molecular weight.

[0193] In step S211), the purification method is ethanol precipitation, microdialysis or Millpore Ziptip microchromatography.

[0194] In step S21), the mass spectrometry sequencing method is MALDI mass spectrometry sequencing.

[0195] S22) analyzing the base combination information based on the molecular weight information of the DNA information to confirm the DNA information corresponding to the different pore positions;

[0196] MALDI mass spectrometry sequencing can display RNA or DNA with different bases as peaks in different time-of-flight sequences, identify the base combinations contained in these time-series peaks, and directly translate these base combinations into their corresponding different nucleic acid sequences.

[0197] S23) decoding the data information unit based on the DNA information obtained in step S22) and the coding table in step S12);

[0198] The chain type and data information coding table in step S23) is the chain type and data information coding table obtained in step S12) in the above-mentioned DNA storage method based on natural and non-natural bases.

[0199] In a specific embodiment, according to the chain type obtained in step S22), the corresponding binary number is confirmed in Table 2. Since every two chain types in Table 2 correspond to 8 binary digits, the conversion is performed in such a way that every two serial chains correspond to one ASCII character to obtain all binary data;

[0200] S24) splicing the data units obtained in S23) and decoding the data to obtain stored computer data.

[0201] According to the six specific implementation plans for data to be stored shown in Example 1, the steps in the decoding process are respectively given:

[0202] First, for a chip storing the English word and characters "Hello world!", the polynucleotide sequences in each of the 12 wells on the chip were digested and purified. MALDI mass spectrometry was then used to determine the specific nucleotide types and quantities within each strand. The MALDI mass spectrometry results confirmed the type of polynucleotide sequence corresponding to each well. Based on the correspondence between the DNA information used in the encoding process and the data encoding table, the corresponding binary number was determined. This binary number was then directly converted to the corresponding character, yielding the original data information—the string "Hello world!".

[0203] In a second specific embodiment, the stored information is a text file "wssnt10.txt" encoded in Goldman's 2012 article "Toward spractical, high-capacity, low-maintenance information storage in synthesized DNA".

[0204] The polynucleotide sequences in different wells of the chip are first digested and purified separately. MALDI mass spectrometry is then used to determine the specific nucleotide types and quantities within each strand. The MALDI mass spectrometry results confirm the type of DNA information corresponding to each well. Based on the correspondence between the DNA information used in the encoding process and the data information encoding table, the different encoding bases are converted into corresponding characters, resulting in the raw data information, which is then stored as a text file called "wssnt10.txt."

[0205] In the third embodiment, the stored information is picture information, which is different from the first two embodiments only in that the obtained binary data is converted into picture information.

[0206] In a fourth specific embodiment, for the stored data of Pride and Prejudice, Chapter 1, the modified base nucleotides in different wells of the chip are first purified, and then the specific types and quantities of the modified base nucleotides are determined by MALDI mass spectrometry. The MALDI mass spectrometry results can be used to confirm the type of modified base nucleotide corresponding to each well. The corresponding binary number is then determined based on the correspondence between the modified base nucleotides used in the encoding process and the data information matching table. The binary number is then directly converted into the corresponding character, thus obtaining the original data information, namely Pride and Prejudice, Chapter 1.

[0207] In the fifth specific embodiment, the stored information is Chapter 1 of Pride and Prejudice.

[0208] First, the modified base nucleotides in different wells of the chip are purified separately. Then, MALDI mass spectrometry is used to determine the specific nucleotide species within these modified base nucleotides. The MALDI mass spectrometry results confirm the type of modified base nucleotide corresponding to each well. Based on the correspondence between the modified base nucleotides used in the encoding process and the data information matching table, the corresponding characters can be directly identified. These characters are then spliced ​​together to obtain the original computer data information.

[0209] In the sixth embodiment, the stored information is picture information, which is similar to the fourth and fifth embodiments, except that the obtained binary data is converted into picture information.

[0210] Example 3: Encoding device for decoding DNA storage based on natural and unnatural base matrix spectra

[0211] Corresponding to the encoding method described in Example 1 above, Figure 5 FIG. 4 shows a block diagram of a coding apparatus according to Embodiment 3 of the present invention. For ease of explanation, Figure 5 Only the parts related to Example 3 of the present invention are shown.

[0212] Reference Figure 5 , the encoding device may include:

[0213] A data information extraction unit, configured to extract computer data to be stored and convert the computer data to be stored into data information corresponding to the information;

[0214] A data information and DNA information conversion unit, configured to split or assemble the data information sequence and convert it into DNA information according to a preset mapping relationship;

[0215] The synthesis and storage unit is used to synthesize the data information and the DNA sequence converted by the DNA information conversion unit, and store the DNA sequence in different holes of the storage unit chip in order.

[0216] The data information extraction unit may include an information storage unit and a conversion unit. The information storage unit can be used to store and retrieve computer information to be stored, such as text, numbers, pictures, audio, video, etc. The conversion unit can convert computer information into any type of digital information, such as characters, binary data information, octal data information, hexadecimal data information, decimal data information, RGB pixel information, etc. using conventional methods.

[0217] The data information and DNA information conversion unit may include a DNA information encoding unit, a DNA information and data information matching unit, and a DNA information conversion unit. The DNA information encoding unit is used to record the combination of base types and quantities corresponding to each type of DNA information. The DNA information and data information matching unit is used to call different DNA information in the DNA information encoding unit for one-to-one matching and correspondence with data information units. The DNA information conversion unit is used to convert the digital information in the data information extraction unit into DNA information one by one based on the information of the DNA information and data information matching unit.

[0218] In one specific technical solution, at least 128 different chain types can be formed by selecting only four different deoxyribonucleotide bases. This is achieved by combining different numbers of four different deoxyribonucleic acids. Eight length gradients are set for the chain types, ranging from 10 to 24. Each length gradient is further configured with four DNA chains of different base contents. Four chains of the same length can be combined to create 16 different combinations. With eight length gradients, there are a total of 8 × 16 = 128 types. By increasing the number of nucleotides or adjusting the length of the chain types, the number of chain types can be increased or decreased to meet different information storage requirements.

[0219] In another specific technical solution, only four deoxyribonucleotides can be selected, and each deoxyribonucleotide can be subjected to 32 or 64 different modifications, thereby forming at least 128 or 256 different nucleotides with modified bases. By increasing the number of nucleotides or adjusting the type and number of modifications of the nucleotides with modified bases, the number of nucleotides with modified bases can be increased or decreased to meet different information storage requirements.

[0220] The synthesis and storage unit includes a synthesis unit and a storage unit. The synthesis unit is capable of synthesizing DNA information for storage, which is then confirmed by the DNA information conversion unit. The storage unit is capable of storing the DNA sequence containing the data information. The storage unit is a chip with multiple wells, each well accommodating a sequence. In other words, each well corresponds to a type of DNA information. The wells on the chip are arranged in a sequential order.

[0221] Example 4 Decoding device for nucleic acid storage based on natural and unnatural base matrix spectra

[0222] Corresponding to the decoding method described in Example 2 above, Figure 6 FIG. 4 shows a structural block diagram of a decoding device provided by embodiment 4 of the present invention. For ease of explanation, Figure 6 Only the parts related to Example 4 of the present invention are shown.

[0223] Reference Figure 6 , the decoding device may include:

[0224] A reading unit, used to detect the sequence to be tested stored in the synthesis and storage unit through a mass spectrometer and confirm its DNA information based on the molecular weight;

[0225] A DNA information and data information conversion unit, configured to convert the DNA information obtained by the reading unit into data information according to a preset mapping relationship, i.e., a mapping relationship between the same data information and DNA information in the DNA information storage device;

[0226] The reading module includes a mass spectrometer for performing mass spectrometry detection on the nucleic acid sequence in each well position in the chip for detecting stored information, and may also include pre-processing such as purification and / or enzymatic hydrolysis of the nucleic acid sequence in each well position.

[0227] A computer data output unit, used to convert the DNA information and the data information obtained by the data information conversion unit into stored computer data;

[0228] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the embodiment of the method of the present invention. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.

[0229] Figure 7 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Figure 7 As shown, the computer device of this embodiment includes: at least one processor ( Figure 7 Only one is shown), storage and a computer program stored in the memory and executable on the at least one processor, wherein when the processor executes the computer program, any storage method and decoding method of the present invention is implemented.

[0230] The computer device may be a laptop computer, desktop computer, tablet computer, mobile phone or other computing device. The computer device includes at least a processor and a memory. It will be understood by those skilled in the art that Figure 7 It is merely a schematic diagram of a computer device and does not constitute a limitation on the computer device, which may also include other components, such as information input or output components.

[0231] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned various method embodiments can be implemented.

[0232] The computer-readable storage medium may also be any available medium or data storage device that can be accessed by a computer, such as a server or data center integrated with the medium. The available medium may be a magnetic medium, a DVD, or a semiconductor medium.

[0233] An embodiment of the present invention provides a computer program product. When the computer program product is run on a terminal device, the terminal device can implement the steps in the above-mentioned method embodiments when executing the computer program product.

[0234] The terminal device may be a general-purpose computer, a handheld computer, a mobile phone, a dedicated computer, a computer network, or other programmable devices, or a storage device with programming functions. The computer program may be stored in a computer-readable storage medium, or transmitted to another computer-readable storage medium via a network.

[0235] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0236] In the above embodiments, all or part of the above embodiments may be implemented by software, hardware, firmware, or any combination thereof. When the computer program instructions are loaded and executed on a computer, all or part of the steps and methods described in the embodiments of the present invention are generated.

[0237] It is understood that the systems, devices and methods described in this application can also be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the functional units can be re-divided according to actual needs, and does not affect its ability to meet or complete the above-mentioned functions and steps of the present invention. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. The units of the above-mentioned devices can be merged or re-divided according to the storage and decoding methods, and additional functional units can be added according to actual needs to meet the requirements of the above-mentioned steps and methods.

[0238] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A DNA information storage method based on natural and unnatural bases, characterized in that: The DNA information storage method comprises the following steps: S11): extracting data information of computer data information to be stored; S12): Designing a coding table, wherein the coding table is formed by one-to-one mapping of DNA information and data information units; S13) splitting the data information in step S11) into data information units, and sequentially confirming the DNA information corresponding to the split data information units in the coding table obtained in step S12); S14) obtaining the deoxyribonucleotides or polynucleotide sequences consisting of two or more deoxyribonucleotides corresponding to the DNA information confirmed in step S13), and arranging them sequentially in different wells of the chip to obtain a DNA information storage carrier; The DNA information is a deoxyribonucleotide or a polynucleotide sequence consisting of two or more deoxyribonucleotides, or a combination of any two of the deoxyribonucleotides or the polynucleotide sequence consisting of two or more deoxyribonucleotides; In the coding table, the molecular weight of each DNA information is different, and the molecular weight difference between each two is not less than 10; The coding table maps DNA information to 4-bit or 8-bit binary numbers one by one; or The coding table maps DNA information to the ASCII code table one by one; or The coding table maps DNA information to 128 characters one by one; or The coding table maps DNA information to RGB pixel information one by one; or The DNA information in the coding table is 128 or 256 different nucleotides with modified bases, and 32 or 64 different types of modifications are performed on A, T, C, and G, respectively, to obtain a total of 128 or 256 different nucleotides with modified bases; or In the coding table, one of the DNA information is mapped one by one to the color information in the RGB pixel information, that is, R\G\B, and the other of the DNA information is mapped one by one to the numbers 0-255 in the color information of the RGB pixel information, and the combination of the two DNA information forms and A one-to-one mapping relationship between RGB pixel information; or The coding table contains 128 combined polynucleotide sequences, and the lengths of the polynucleotide sequences are divided into 8 groups. The lengths of the 8 groups of polynucleotide sequences are successively extended, each containing 10-24 base nucleotides. There are 4 polynucleotide sequences in each group, and the types and or quantities of bases in the 4 polynucleotide sequences in each group are different.

2. The DNA information storage method according to claim 1, characterized in that: The deoxyribonucleotide is a natural deoxyribonucleotide or a non-natural deoxyribonucleotide, and the non-natural deoxyribonucleotide is a deoxyribonucleotide with modified bases.

3. The DNA information storage method according to claim 1, characterized in that: The deoxyribonucleotides in the polynucleotide sequence composed of two or more deoxyribonucleotides are natural deoxyribonucleotides or non-natural deoxyribonucleotides.

4. The DNA information storage method according to claim 3, characterized in that: The polynucleotide sequence composed of two or more deoxyribonucleotides can have different molecular weights by adjusting the combination of the types and quantities of the deoxyribonucleotides in the polynucleotide sequence.

5. The DNA information storage method according to claim 1, characterized in that: The 32-1024 mapping relationships formed in the coding table.

6. The DNA information storage method according to claim 1, characterized in that: In step S11), the data information is character information, RGB information, binary data information, octal data information, hexadecimal data information, or decimal data information.

7. The DNA information storage method according to claim 1, characterized in that: In the coding table, the molecular weight of each DNA information is different, and the molecular weight difference between each two is not less than 10.

8. The method for reading information from a DNA information storage carrier obtained by the DNA information storage method according to any one of claims 1 to 7, characterized in that: The information reading method comprises the following steps: S21) performing mass spectrometry on the sequences to be tested at different well positions in the DNA information storage carrier to obtain molecular weight information of the DNA information at each well position; S22) analyzing the base combination information based on the molecular weight information of the DNA information to confirm the DNA information corresponding to the different pore positions; S23) decoding the data information unit based on the DNA information obtained in step S22) and the coding table in step S12); S24) splicing the data units obtained in S23) and decoding the data to obtain stored computer data.

9. The information reading method according to claim 8, characterized in that: In step S21), the mass spectrometry sequencing method is MALDI mass spectrometry sequencing.

10. The information reading method according to claim 8, wherein: The mass spectrometry sequencing method in step S21) comprises the following steps: S211) performing enzyme digestion and / or purification on the sequence to be tested; S212) The purified fragments are subjected to mass spectrometry to obtain molecular weight.

11. The information reading method according to claim 8, wherein: The purification method in step S211) is ethanol precipitation, microdialysis or Millpore Ziptip microchromatography.

12. An encoding device based on natural and unnatural base DNA information storage, characterized in that: The DNA information storage device comprises: A data information extraction unit, configured to extract computer data to be stored and convert the computer data to be stored into data information corresponding to the information; A data information and DNA information conversion unit, configured to split or assemble the data information sequence and convert it into DNA information according to a preset mapping relationship; A synthesis and storage unit, used to synthesize the data information and the nucleic acid sequence converted by the DNA information conversion unit, and store the deoxyribonucleotides corresponding to the DNA information, a polynucleotide sequence consisting of two or more deoxyribonucleotides, or a combination thereof in different wells of the storage unit chip in sequence; The mapping relationship is a one-to-one mapping relationship between DNA information and data information units; The DNA information is a deoxyribonucleotide or a polynucleotide sequence consisting of two or more deoxyribonucleotides, or a combination of any two of the deoxyribonucleotides or the polynucleotide sequence consisting of two or more deoxyribonucleotides; In the mapping relationship, the molecular weight of each DNA information is different, and the molecular weight difference between each two is not less than 10; The mapping relationship is a one-to-one mapping of DNA information to 4-bit or 8-bit binary numbers; or The mapping relationship is to map DNA information to the ASCII code table one by one; or The mapping relationship is to map DNA information to 128 characters one by one; or The mapping relationship is to map DNA information to RGB pixel information one by one; or The mapping relationship is that the DNA information is 128 or 256 different nucleotides with modified bases, and 32 or 64 different types of modifications are performed on A, T, C, and G, respectively, to obtain a total of 128 or 256 different nucleotides with modified bases; or The mapping relationship is to map one of the DNA information to the color information in the RGB pixel information, that is, R\G\B, and the other of the DNA information to the numbers 0-255 in the color information of the RGB pixel information, and to form a combination of the two DNA information. and A one-to-one mapping relationship between RGB pixel information; or The mapping relationship includes 128 combined polynucleotide sequences, the lengths of the polynucleotide sequences are divided into 8 groups, the lengths of the 8 groups of polynucleotide sequences are extended successively, each containing 10-24 base nucleotides, 4 polynucleotide sequences in each group, and the types and or quantities of bases in the 4 polynucleotide sequences in each group are different.

13. The DNA information storage device according to claim 12, characterized in that: The data information and DNA information conversion unit includes a DNA information encoding unit, a DNA information and data information matching unit, and a DNA information conversion unit; the DNA information encoding unit is used to record the combination of base types and quantities corresponding to each type of DNA information; the DNA information and data information matching unit is used to call different DNA information in the DNA information encoding unit for one-to-one matching and correspondence with data information units; the DNA information conversion unit is used to convert the digital information in the data information extraction unit into DNA information one-to-one according to the information of the DNA information and data information matching unit.

14. A decoding device for DNA information storage based on natural and unnatural bases, characterized in that: The decoding device comprises: a reading unit for detecting the sequence to be tested synthesized and stored in the storage unit of the DNA information storage device according to claim 12 or 13 by a mass spectrometer, and confirming its DNA information according to the molecular weight; A DNA information and data information conversion unit, configured to convert the DNA information obtained by the reading unit into data information according to a preset mapping relationship, i.e., a mapping relationship between the same data information and DNA information in the DNA information storage device; A computer data output unit, used to convert the DNA information and the data information obtained by the data information conversion unit into stored computer data; The reading unit includes a mass spectrometer for performing mass spectrometry detection on the nucleic acid sequence in each well in the chip for detecting stored information.

15. The decoding device according to claim 14, wherein: The reading unit further comprises a unit for pre-processing the sequence to be tested in each well position.

16. A computer-readable storage medium, characterized in that A computer program is stored thereon, wherein when the computer program is executed by a processor, the steps of the method for obtaining an information reading method of a DNA information storage carrier by the DNA information storage method based on natural and non-natural bases as described in any one of claims 1 to 7 or the DNA information storage method as described in any one of claims 8 to 11 are implemented.

17. A computer device, characterized in that: The invention comprises a memory and a processor, wherein a computer program capable of being run on the processor is stored on the memory, and when the processor executes the computer program, the steps of the method for obtaining an information reading method of a DNA information storage carrier by the DNA information storage method based on natural and non-natural bases as described in any one of claims 1 to 7 or the DNA information storage method as described in any one of claims 8 to 11 are implemented.

Citation Information

Patent Citations

  • Method for writing and reading information by using DNA sequence

    CN113066534A

  • Drug resistance gene identification method and system and electronic equipment

    CN114283886A

  • Device for information encoding and, storage using artificially expanded alphabets of nucleic acids and other analogous polymers

    WO2019040871A1