Data storage technology based on Z-DNA type and related application thereof

By using pre-synthesized DNA fragment libraries and Z-bases, combined with nanopore sequencing technology, the problems of high writing costs and long delays in DNA information storage technology have been solved, achieving low-cost and efficient traceability and anti-counterfeiting of biomass products.

CN120600128APending Publication Date: 2025-09-05TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510259090.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-05
Filing Date
2025-03-05
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing DNA information storage technology has high costs and long delays in writing mode, which makes it difficult to meet the needs of anti-counterfeiting and traceability of biomass products.

Method used

A pre-synthesized DNA fragment library is used, and Z-bases are used as information storage units. Low-cost and efficient information writing is achieved through rapid selective mixing. Nanopore sequencing technology is used for reading, combined with the k-mer algorithm to distinguish Z-DNA from ordinary DNA.

Benefits of technology

It achieves fast and low-cost information writing and reading, ensures the traceability and anti-counterfeiting capabilities of biomass products, reduces storage costs and improves reading efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120600128A_ABST
    Figure CN120600128A_ABST
Patent Text Reader

Abstract

The invention provides an information storage method, an information reading and decoding method, an information storage system, application of a DNA fragment library in product traceability and anti-counterfeiting aspects and a method for traceability or anti-counterfeiting verification of a product. Due to the advantages of the method in the aspects of privacy, writing and reading speed and cost, the method has great potential in the aspects of product traceability and anti-counterfeiting based on DNA.
Need to check novelty before this filing date? Find Prior Art

Description

Field of the Invention

[0001] The present application generally relates to the field of bioinformatics, and more specifically, to a sensitive information protection storage technology and its application. Background of the Invention

[0002] DNA information storage is a cutting-edge information storage technology that utilizes DNA molecules as an information storage medium. As a biological macromolecule that stores genetic information within organisms, DNA offers advantages such as high density, long-term stability, and potentially enormous storage capacity. In DNA information storage, digital data is encoded into DNA sequences, and information is stored and retrieved through processes such as synthesis, storage, and sequencing. DNA information storage technology, with its high density, long-term stability, and excellent data security, has attracted widespread attention from scientists and is considered a promising storage solution. The history of DNA information storage dates back to the 1950s, when scientists first recognized the ability of DNA molecules to carry genetic information. However, it was only in recent years, with the rapid advancement of DNA synthesis and sequencing technologies, that practical applications of DNA information storage became possible. For example, in 2012, researchers at Harvard University successfully encoded and stored 50,000 words of text and an image into DNA sequences. This research not only demonstrated the enormous potential of DNA information storage but also promoted further development in the field.

[0003] Although DNA as a storage medium offers significant advantages in terms of data security, information density, and long-term reliability, its practical application is still limited by high read and write costs and long read latency. The recent discovery of Z-bases has provided new opportunities for the development of DNA information storage technology. By introducing Z-bases, this technology can further increase storage density and develop new information encryption schemes to meet a wider range of application needs. Furthermore, with the continuous advancement of sequencing technology, the efficiency and reliability of DNA information reading have been significantly improved. SUMMARY OF THE INVENTION

[0004] DNA information storage technology demonstrates unique application potential in the field of biomass product anti-counterfeiting and traceability. Biomass products, such as cotton, traditional Chinese medicine, fruit, and grain crops, are vulnerable to counterfeiting and tampering due to their complex production and supply chain processes. Traditional anti-counterfeiting technologies often rely on physical labels or chemical markers, but these methods are susceptible to imitation, high tracking costs, and limited information storage capacity. DNA-based storage technology can provide a unique identifier for each biomass product, ensuring its uniqueness and traceability. By encoding DNA fragments into specific sequences, information such as the product's origin, production time, and batch can be embedded in the DNA. This technology not only effectively prevents product counterfeiting but also enables tracking of products throughout their entire life cycle.

[0005] Storage technology based on Z-DNA movable type further improves the efficiency and security of DNA information storage. By using Z-bases as the core unit of information storage, this technology not only significantly increases storage speed and reduces storage costs, but also prevents the copying of anti-counterfeiting code information. Combined with the characteristics of biomass products, this technology provides a new solution for the protection and high-density storage of sensitive information. For example, in cotton anti-counterfeiting, the cotton's origin, composition information, and production records can be embedded into the DNA sequence by pre-synthesizing a specific DNA fragment library. These DNA fragments can be conveniently sprayed or attached to the cotton surface and quickly read and decoded using sequencing technology, thus achieving anti-counterfeiting traceability.

[0006] In order to solve the high cost and high delay problems of DNA synthesis-based writing mode, this application adopts a method of rapid selective mixing of pre-synthesized DNA fragments with different sequence compositions and different types to achieve fast and low-cost information writing.

[0007] The present application also develops a reliable algorithm to distinguish between normal DNA and Z-DNA fragments containing Z-bases. The present application also makes full use of the advantages of modern sequencing technologies, such as nanopore sequencers, which have fast reading speeds, low costs, and can quickly distinguish Z-DNA from normal DNA.

[0008] In a first aspect, the present application provides an information storage method, comprising: using a plurality of DNA fragments to store the information in the form of a DNA fragment library, wherein the information includes a plurality of bits,

[0009] Each of the plurality of DNA fragments includes an information recording unit, and the information recording unit can be used for indexing and data recording.

[0010] wherein the data recording purpose of the information recording unit is realized by whether the DNA fragment contains a Z-base, wherein the DNA fragment containing the Z-base is set to correspond to information bit 1, and the DNA fragment not containing the Z-base is set to correspond to information bit 0, or vice versa, and

[0011] The indexing is achieved through the specific sequences of the multiple DNA fragments.

[0012] In a second aspect, the present application provides an information reading method, comprising:

[0013] a) sequencing multiple DNA fragments in the DNA fragment library,

[0014] Each of the plurality of DNA fragments includes an information recording unit, and the information recording unit can be used for indexing and data recording.

[0015] wherein the data recording purpose of the information recording unit is realized by whether the DNA fragment contains a Z-base, wherein the DNA fragment containing the Z-base is set to correspond to information bit 1, and the DNA fragment not containing the Z-base is set to correspond to information bit 0, or vice versa, and

[0016] wherein the indexing is achieved by the specific sequences of the multiple DNA fragments;

[0017] b) grouping the multiple DNA fragments according to the sequencing results of step a); and

[0018] c) Calculating the PIK value of each grouped DNA fragment based on the k-mer algorithm, and then determining whether the DNA fragment contains Z-base.

[0019] In a third aspect, the present application provides an information storage method, comprising:

[0020] a) pre-synthesizing multiple DNA fragment libraries, each DNA fragment library containing multiple DNA fragments,

[0021] Each of the plurality of DNA fragments includes an information recording unit, and the information recording unit can be used for indexing and data recording.

[0022] wherein the data recording purpose of the information recording unit is realized by whether the DNA fragment contains a Z-base, wherein the DNA fragment containing the Z-base is set to correspond to information bit 1, and the DNA fragment not containing the Z-base is set to correspond to information bit 0, or vice versa, and

[0023] wherein the indexing is achieved by the specific sequences of the multiple DNA fragments,

[0024] wherein whether each of the plurality of DNA fragments contains a Z-base is predetermined,

[0025] b) randomly combining one or more of the plurality of DNA fragment libraries for information storage.

[0026] In a fourth aspect, the present application provides an information storage system comprising one or more pre-synthesized DNA fragment libraries, each DNA fragment library comprising a plurality of DNA fragments,

[0027] Each of the plurality of DNA fragments includes an information recording unit, and the information recording unit can be used for indexing and data recording.

[0028] wherein the data recording purpose of the information recording unit is realized by whether the DNA fragment contains a Z-base, wherein the DNA fragment containing the Z-base is set to correspond to information bit 1, and the DNA fragment not containing the Z-base is set to correspond to information bit 0, or vice versa, and

[0029] wherein the indexing is achieved by the specific sequences of the multiple DNA fragments,

[0030] Whether each of the plurality of DNA fragments contains a Z-base is predetermined.

[0031] In a fifth aspect, the present application provides a DNA fragment library for use in product traceability and anti-counterfeiting, wherein the DNA fragment library comprises a plurality of DNA fragments.

[0032] Each of the plurality of DNA fragments includes an information recording unit, and the information recording unit can be used for indexing and data recording.

[0033] wherein the data recording purpose of the information recording unit is realized by whether the DNA fragment contains a Z-base, wherein the DNA fragment containing the Z-base is set to correspond to information bit 1, and the DNA fragment not containing the Z-base is set to correspond to information bit 0, or vice versa, and

[0034] wherein the indexing is achieved by the specific sequences of the multiple DNA fragments,

[0035] Whether each of the plurality of DNA fragments contains a Z-base is predetermined.

[0036] In a sixth aspect, the present application provides a method for tracing or verifying the product's origin or anti-counterfeiting, comprising the following steps:

[0037] a) pre-synthesizing a DNA fragment library, wherein the DNA fragment library comprises a plurality of DNA fragments,

[0038] Each of the plurality of DNA fragments includes an information recording unit, and the information recording unit can be used for indexing and data recording.

[0039] wherein the data recording purpose of the information recording unit is realized by whether the DNA fragment contains a Z-base, wherein the DNA fragment containing the Z-base is set to correspond to information bit 1, and the DNA fragment not containing the Z-base is set to correspond to information bit 0, or vice versa, and

[0040] wherein the indexing is achieved by the specific sequences of the multiple DNA fragments,

[0041] wherein whether each of the plurality of DNA fragments contains a Z-base is predetermined;

[0042] b) spraying the DNA fragment library on the surface of the product, and

[0043] c) sequencing the DNA fragment library sprayed on the surface of the product to reconstruct the anti-counterfeiting code information.

[0044] Due to the advantages of this application in writing, reading speed and cost, it has great potential in DNA-based product traceability and anti-counterfeiting. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 A schematic diagram showing the principle of Z-DNA movable type information storage.

[0046] Figure 2 The data writing and reading process of Z-DNA movable type is shown.

[0047] Figure 3 The algorithm for distinguishing between Z-DNA and DNA fragments is shown.

[0048] Figure 4 It shows the product traceability and anti-counterfeiting application process of Z-DNA movable type technology.

[0049] Figure 5 (A) shows the PIK value distribution of sequencing reads of Z-DNA anti-counterfeiting code A; Figure 5 Panel (B) shows the PIK value distribution of sequencing reads from amplified sample B, obtained by amplification of Z-DNA security code A. The x-axis represents the number of individual sequencing reads, and the y-axis (logarithmic scale) shows the corresponding PIK value for each read. Red circles represent DNA reads, and blue circles represent Z-DNA reads. The results show that the bit values ​​of the amplified sample are indistinguishable, indicating information loss.

[0050] Figure 6 Flowchart for grouping sequencing reads according to index sequence.

[0051] Figure 7 Flowchart for decoding bit values ​​of sequenced reads within each group. Detailed Description of the Invention

[0052] Unless otherwise specified, the terms used in this application have the meanings commonly understood by those skilled in the art.

[0053] Unless otherwise indicated, nucleic acids are written left to right in 5' to 3' orientation; nucleotides may be referred to by their commonly accepted single-letter codes.

[0054] Current data protection technologies for copy protection are based on cryptographic algorithms or hardware-level protection for media readouts, rather than underlying (media-level) copy protection. Current data disc protection technologies, such as CD-Cops and Dummy files, employ two mechanisms: adding a protective sleeve to the installation program, verifying a password (usually 8 characters) during installation, and modifying ISO file size-related codes to create oversized fake files, causing the disc's total capacity to display as a maximum of 2GB. This creates the illusion that common 650MB and 700MB CDRs cannot be backed up. Current USB flash drive data protection technologies include, but are not limited to, keystroke encryption, fingerprint encryption, software encryption, hardware encryption, and a combination of software and hardware encryption.

[0055] Key encryption type

[0056] Professionally manufactured copy-protected USB flash drives feature built-in physical digital keys, allowing users to manually enter a pre-set password to encrypt and decrypt data. Key-encrypted copy-protected USB flash drives store the password in the encryption chip. Furthermore, key-encrypted copy-protected USB flash drives can encrypt and decrypt data without a computer and support multiple accounts and permissions.

[0057] Fingerprint encryption type

[0058] A technologically sophisticated anti-copy USB flash drive can have a built-in fingerprint collection and recognition device. Since each person's fingerprint is unique and remains unchanged throughout life, a fingerprint-encrypted anti-copy USB flash drive can rely on the uniqueness and stability of the fingerprint to match the user's fingerprint, thereby verifying the user's true identity, and thus achieve file data encryption and decryption functions through this method.

[0059] Software encryption type

[0060] According to our understanding of the product, software-encrypted anti-copy USB flash drives encrypt the contents of the USB flash drive through built-in and accompanying software, and generally use ASP encryption. Therefore, software-encrypted anti-copy USB flash drives can partially avoid the disadvantage that the files on the original USB flash drive can be read out on other PCB boards through password cracking tools.

[0061] Hardware encryption

[0062] When this type of technology reads files, the decryption is performed on the USB flash drive. The files transmitted through the USB port are already decrypted, and the files on the disk can be easily obtained using a USB packet capture tool. Hardware encryption technology is actually not suitable for "anti-copying" scenarios.

[0063] Software and hardware combined encryption

[0064] This type of technology uses an original VNAS encryption solution, which can achieve active authorization control and ensure security without loopholes. It does not use any non-standard hacker technology, has no drivers or hooks, and has no risk of being mistakenly killed by antivirus software. Files are highly encrypted and decryption is almost impossible.

[0065] However, all of the above protection methods require a password / fingerprint as a key for reading information, which means that these methods cannot solve the protection problem of the key itself.

[0066] DNA modifications vary in form and function, but generally do not alter Watson-Crick base pairing. Diaminopurine (Z-base) is a unique exception because, in cyanobacterial phages, it completely replaces adenine and forms three hydrogen bonds with thymine. During in vitro DNA synthesis, deoxyribonucleotides with Z-bases can be added to in vitro amplification systems such as PCR systems, replacing the A and T pairing in conventional DNA, thereby incorporating Z-bases into the synthesized DNA.

[0067] In the present application, DNA containing Z-bases is referred to as Z-DNA, and for the sake of distinction, conventional DNA that does not contain Z-bases is referred to as DNA. The inventors of the present application cleverly developed a DNA-based password storage system based on the property that T can pair with both A and Z. It is not only safe and reliable, but also not easy to be stolen. This is because when the specific sequence of the DNA is not known in advance, it is difficult to amplify the DNA containing Z-bases through an in vitro amplification system, and thus it is impossible to decipher the implicit information in the DNA sequence. Therefore, the applicant has pioneered the application of biological methods to the field of electronics, specifically, to the field of informatics. On the other hand, the present application pre-synthesized a plurality of Z-DNA movable type pairs, which can be used repeatedly, and the Z-DNA movable type pairs can be automatically prepared on a large scale, so the storage capacity of Z-DNA encryption is expandable.

[0068] The present application aims to provide a movable type information storage technology based on DNA fragments containing Z-bases (Z-DNA) and its application in product traceability and anti-counterfeiting. Its basic principle is to record information by distinguishing Z-bases from ordinary bases, or by sequences containing Z-bases. The rapid and low-cost writing of information is achieved by selectively using multiple pre-synthesized "movable type-like" fragments containing Z-bases and not containing Z-bases and reusing these fragments. The information writing process of this method only requires a simple mixing of the required fragments, and each fragment can be used repeatedly, with the significant advantages of a fast writing process and low cost. And the reading process can be reliably read by nanopore sequencing technology. The fast writing and portable reading characteristics of this technology make this technology have great application value in product traceability based on DNA anti-counterfeiting code technology. In addition, this technology also has great scalability potential in big data storage.

[0069] like Figure 1 As shown in A, this application realizes the writing and storage of information through the "movable type pair" of ordinary DNA and Z-DNA. The sequence information of the movable type pair is used as the sorting basis (index sequence), and whether the sequence contains Z-bases represents the digital information "0" or "1", thereby storing information. Figure 1 As shown in B, since the number of movable type pairs can be freely expanded as needed, the method of the present application has great scalability in terms of information storage capacity.

[0070] The entire Z-DNA movable type storage writing and reading process is as follows Figure 2 As shown. Z-DNA type pairs can be reused after pre-synthesis. The data writing process of Z-DNA type storage is achieved by quickly selecting the corresponding ordinary DNA or Z-DNA type of the Z-DNA type pair, which can be completed in seconds. Therefore, the storage capacity of Z-DNA encryption is scalable, especially considering that indexes and data fragments can be prepared on a large scale through automation and reused. The stored data can be sequenced and read using a nanopore sequencer. This process includes three main steps. In the first step, all type fragments are sequenced in parallel using nanopore sequencing. In the second step, the sequenced reads are grouped and sorted according to the sequence after base recognition. In the final step, a k-mer-based algorithm is used to calculate the PIK value of each group, determine the Z-base insertion status of each type fragment, and then determine the bit value according to the preset rules based on the Z-base insertion status, thereby achieving complete reading and decoding of the written information.

[0071] In order to read the data accurately, we developed a Figure 3Figure A shows the k-mer-based algorithm. The basic steps include first constructing k-mer set A using the designed DNA sequence, and then constructing k-mer set B using the grouped Nanopore sequencing sequence clusters generated by the Guppy basecaller. The percentage of k-mers in set A (PIK) representing the intersection of sets A and B is used to distinguish between Z-DNA and normal DNA.

[0072] For the Z-DNA with the same sequence design, except that all A's were replaced by Z's, the PIK values ​​showed significant deviations ( Figure 3 B) This process was repeated 100,000 times to determine the optimal threshold to distinguish DNA from Z-DNA fragments. Figure 3 (B shows a partial result). Based on these 100,000 calculations, we obtained an optimal threshold of 72.7%. It should be noted that this optimal threshold is best optimized based on the actual performance of the nanopore sequencing chip used.

[0073] An odd number of sequencing reads are randomly selected, and the PIK value for each sequencing read and the data region sequence is calculated. The PIK value is then used to calculate the bit value of each sequencing read. The bit value of the entire set of sequencing reads is then calculated using a majority voting mechanism. Based on this threshold, the error rate for a single read of 11 random sequencing reads is as low as approximately 3.90E-08. Nanopore sequencing can quickly and inexpensively generate a large number of sequencing reads, sufficient to support tens of thousands of searches. By performing multiple (n) searches, the error rate can be further reduced.

[0074] This application has the advantages of fast reading and writing speed and low cost, and is very suitable for tracing the origin of agricultural products such as cotton, traditional Chinese medicine, fruit, and grain. Figure 4 The application process of Z-DNA movable type storage for product traceability and anti-counterfeiting is demonstrated. First, a specific security code is converted into a Z-DNA movable type combination scheme based on specific security code encoding technology. Then, large quantities of Z-DNA movable type are produced through large-scale preparation technology. During use, the Z-DNA movable type security code is sprayed onto a specific product, such as cotton. This process uniquely associates the Z-DNA movable type security code with the specific product. This information can be quickly read and reconstructed using third-generation sequencing technology. Combined with a traceability record database, this enables historical traceability and anti-counterfeiting capabilities for specific products.

[0075] Specifically, this application provides the following technical solutions:

[0076] In a first aspect, the present application provides an information storage method, comprising: using a plurality of DNA fragments to store the information in the form of a DNA fragment library, wherein the information includes a plurality of bits,

[0077] Each of the plurality of DNA fragments includes an information recording unit, and the information recording unit can be used for indexing and data recording.

[0078] wherein the data recording purpose of the information recording unit is realized by whether the DNA fragment contains a Z-base, wherein the DNA fragment containing the Z-base is set to correspond to information bit 1, and the DNA fragment not containing the Z-base is set to correspond to information bit 0, or vice versa, and

[0079] The indexing is achieved through the specific sequences of the multiple DNA fragments.

[0080] In a preferred embodiment, the library of DNA fragments is presynthesized.

[0081] In a second aspect, the present application provides an information reading method, comprising:

[0082] a) sequencing multiple DNA fragments in the DNA fragment library,

[0083] Each of the plurality of DNA fragments includes an information recording unit, and the information recording unit can be used for indexing and data recording.

[0084] wherein the data recording purpose of the information recording unit is realized by whether the DNA fragment contains a Z-base, wherein the DNA fragment containing the Z-base is set to correspond to information bit 1, and the DNA fragment not containing the Z-base is set to correspond to information bit 0, or vice versa, and

[0085] wherein the indexing is achieved by the specific sequences of the multiple DNA fragments;

[0086] b) grouping the multiple DNA fragments according to the sequencing results of step a); and

[0087] c) Calculating the PIK value of each grouped DNA fragment based on the k-mer algorithm, and then determining whether the DNA fragment contains Z-base.

[0088] In a preferred embodiment, the library of DNA fragments is presynthesized.

[0089] In some specific embodiments, the sequencing can be performed by a method selected from the group consisting of massively parallel sequencing, ion semiconductor sequencing, and nanopore sequencing.

[0090] Preferably, the sequencing is performed using nanopore sequencing. Using nanopore sequencing, single DNA or RNA molecules can be sequenced without the need for PCR amplification or chemical labeling of the sample. At least one of the above steps is required in any previously developed sequencing method. Nanopore sequencing has the potential to provide relatively low-cost genotyping, high test mobility, and rapid sample processing, with real-time display of results.

[0091] In some specific embodiments, step b) of the second aspect comprises:

[0092] Constructing a k-mer set A using the originally designed sequences not containing Z-bases of the various DNA fragments in the DNA fragment library;

[0093] Each of the DNA fragments is used to construct k-mer sets B1 to B n , where for each DNA fragment, the percentage of its intersection with set A in set A is calculated in turn as the PIK value, and the DNA fragment with the highest PIK value is selected as the grouping result, and n is the number of types of the DNA fragments.

[0094] In a preferred embodiment, the size of the k-mer is 8-18, preferably 10-16, more preferably 13.

[0095] In order to eliminate low-quality reads, a threshold standard for the PIK value was set: for the R9.4.1 sequencing chip, a PIK value below 0.15 was considered a low-quality read and discarded; for the R10.4.1 chip, a PIK value below 0.35 was also discarded.

[0096] In some specific embodiments, step c) of the second aspect comprises:

[0097] Constructing a k-mer set A using the originally designed sequences not containing Z-bases of the various DNA fragments in the DNA fragment library;

[0098] Each sequencing read segment corresponding to the specific sequence of the DNA fragment after grouping is used to construct k-mer sets B1, B2...B n , wherein n is the number of randomly selected sequencing reads, preferably, n is an odd number greater than or equal to 3, more preferably, n=3, 5, 7, 9 or 11; and

[0099] Combine set A with set B n The percentage of the intersection of the two sets in set A is used as the indicator PIK to distinguish whether the sequencing read is Z-DNA or normal DNA. n value,

[0100] In a preferred embodiment, the size of the k-mer is 8-18, preferably 10-16, more preferably 13.

[0101] In some specific embodiments, the PIK n The value is compared with a threshold to determine whether the DNA fragment contains a Z-base, wherein the threshold is about 72.7%. Sequencing reads with a value less than the threshold are judged to be Z-DNA, and the bit value is 1 or 0.

[0102] After the bit values ​​of all sequencing reads are determined, the final bit value of the group of sequencing reads is calculated based on the bit value distribution of the n sequencing reads in the group and the principle of majority voting.

[0103] In a third aspect, the present application provides an information storage method, comprising:

[0104] a) pre-synthesizing multiple DNA fragment libraries, each DNA fragment library containing multiple DNA fragments,

[0105] Each of the plurality of DNA fragments includes an information recording unit, and the information recording unit can be used for indexing and data recording.

[0106] wherein the data recording purpose of the information recording unit is realized by whether the DNA fragment contains a Z-base, wherein the DNA fragment containing the Z-base is set to correspond to information bit 1, and the DNA fragment not containing the Z-base is set to correspond to information bit 0, or vice versa, and

[0107] wherein the indexing is achieved by the specific sequences of the multiple DNA fragments,

[0108] wherein whether each of the plurality of DNA fragments contains a Z-base is predetermined,

[0109] b) randomly combining one or more of the plurality of DNA fragment libraries for information storage.

[0110] In some embodiments of the method described in any of the above aspects, the information includes n cryptographic bits, where n is greater than or equal to 2. Preferably, n is greater than or equal to 16. Due to the extremely large capacity of DNA libraries, the information may include an extremely large number of cryptographic bits, making it almost impossible to decipher the information.

[0111] In some embodiments of the method of any of the above aspects, the one or more DNA fragments are each 20 bp to 2000 bp in length, preferably 100 bp to 250 bp in length. Neither the length nor the composition of the DNA fragments is critical for information storage. Those skilled in the art can select DNA fragments of appropriate lengths as needed to generate a library of DNA fragments ultimately used for information storage.

[0112] In a fourth aspect, the present application provides an information storage system comprising one or more pre-synthesized DNA fragment libraries, each DNA fragment library comprising a plurality of DNA fragments,

[0113] Each of the plurality of DNA fragments includes an information recording unit, and the information recording unit can be used for indexing and data recording.

[0114] wherein the data recording purpose of the information recording unit is realized by whether the DNA fragment contains a Z-base, wherein the DNA fragment containing the Z-base is set to correspond to information bit 1, and the DNA fragment not containing the Z-base is set to correspond to information bit 0, or vice versa, and

[0115] wherein the indexing is achieved by the specific sequences of the multiple DNA fragments,

[0116] Whether each of the plurality of DNA fragments contains a Z-base is predetermined.

[0117] In a fifth aspect, the present application provides a DNA fragment library for use in product traceability and anti-counterfeiting, wherein the DNA fragment library comprises a plurality of DNA fragments.

[0118] Each of the plurality of DNA fragments includes an information recording unit, and the information recording unit can be used for indexing and data recording.

[0119] wherein the data recording purpose of the information recording unit is realized by whether the DNA fragment contains a Z-base, wherein the DNA fragment containing the Z-base is set to correspond to information bit 1, and the DNA fragment not containing the Z-base is set to correspond to information bit 0, or vice versa, and

[0120] wherein the indexing is achieved by the specific sequences of the multiple DNA fragments,

[0121] Whether each of the plurality of DNA fragments contains a Z-base is predetermined.

[0122] The product can be selected from cotton, traditional Chinese medicine, fruit, food crops, such as rice, wheat, corn, sorghum, and any product that needs to be traced and verified for anti-counterfeiting.

[0123] In a sixth aspect, the present application provides a method for tracing or verifying the product's origin or anti-counterfeiting, comprising the following steps:

[0124] a) pre-synthesizing a DNA fragment library, wherein the DNA fragment library comprises a plurality of DNA fragments,

[0125] Each of the plurality of DNA fragments includes an information recording unit, and the information recording unit can be used for indexing and data recording.

[0126] wherein the data recording purpose of the information recording unit is realized by whether the DNA fragment contains a Z-base, wherein the DNA fragment containing the Z-base is set to correspond to information bit 1, and the DNA fragment not containing the Z-base is set to correspond to information bit 0, or vice versa, and

[0127] wherein the indexing is achieved by the specific sequences of the multiple DNA fragments,

[0128] wherein whether each of the plurality of DNA fragments contains a Z-base is predetermined;

[0129] b) spraying the DNA fragment library on the surface of the product, and

[0130] c) sequencing the DNA fragment library sprayed on the surface of the product to reconstruct the anti-counterfeiting code information.

[0131] The product can be selected from cotton, traditional Chinese medicine, fruit, food crops, such as rice, wheat, corn, sorghum, and any product that needs to be traced and verified for anti-counterfeiting.

[0132] In a specific embodiment, the anti-counterfeiting code can be read and decoded according to the information reading method described in the second aspect. Example

[0133] The following examples are illustrative only and are not intended to limit the scope of the embodiments of the present application or the scope of the appended claims.

[0134] Example 1: Sequence Information and Anti-Counterfeiting Code Examples of Z-DNA Type Pairs

[0135] In this example, we designed 64 Z-DNA word pairs. These pairs, through the introduction of Z-bases, enable high-density encoding and storage of information. In practical applications, the number of Z-DNA word pairs can be flexibly adjusted based on specific needs, with no upper limit, thus providing customized solutions for diverse application scenarios.

[0136] 1. Design of Z-DNA active type sequence

[0137] The index and information region sequences of the Z-DNA active type used in this embodiment are shown in Table 1. The length of the Z-DNA active type in this embodiment is 217 bp. In actual application, the length can be adjusted according to specific circumstances.

[0138] Table 1 Sequence information and anti-counterfeiting code examples of 64 Z-DNA movable type pairs

[0139]

[0140]

[0141]

[0142]

[0143]

[0144]

[0145]

[0146] Note: 1 means the information sequence contains Z-base, and 0 means the information sequence does not contain Z-base (the reverse can also be used in actual use);

[0147] The Z-base in the Z-DNA active sequence is also represented by the A base, and the active sequence containing the Z-base is produced by replacing dATP with dZTP during sequence synthesis.

[0148] 2. Construction of Z-DNA active unit sequence

[0149] The index and information region sequences in Table 1 were randomly selected from the E. coli BL21(DE3) genome. Each sequence is 217 bp in length with a GC content between 50% and 60%. The PCR primer pairs used to amplify the 64 index fragments are listed in Table 2. Genomic DNA extracted from 4 mL of an overnight culture of E. coli BL21(DE3) cells using the TIANamp Bacterial DNA Extraction Kit was used as a template to amplify the index and information region sequences. All PCR reactions used Q5 High-Fidelity DNA Polymerase. All subcloning experiments used the ClonExpress II One-Step Cloning Kit, TIANprep Mini Plasmid Extraction Kit, and E. coli DH5α competent cells. To create 5'-GCA-3' sticky ends and avoid palindromes, a BglI restriction site (5'-GCCNNNNNGGC-3') was added to the 3' end of the index fragment and the 5' end of the information region fragment, respectively. The PCR primer pairs used to amplify the 64 index fragments and one information region are listed in Table 2. The digestion reaction was performed at 37°C for 5 hours in a 20 μL reaction system containing 10 units of BglI enzyme. To construct the storage unit, 64 ligation reactions (10 μL total) containing each of the 64 index fragments, the information region fragment, and 200 units of T4 DNA ligase were incubated at 25°C for 25 minutes and then purified using the DNAClean & Concentrator-5 kit. The ligation products were then individually inserted into the pUC19 vector and transformed into Escherichia coli DH5α cells. Plasmids were extracted and verified by sequencing.

[0150] 3. Construction of Z-DNA and conventional DNA movable type units and construction of 64-bit Z-DNA movable type storage anti-counterfeiting code

[0151] Conventional DNA storage units were obtained by conventional PCR using the aforementioned pUC19-based plasmid as a template. Z-DNA storage units were prepared using a different protocol. The index fragment was amplified by conventional PCR, while the information region fragment was amplified using dZTP instead of dATP. Z-DNA storage units were then obtained by ligating the index fragment with the corresponding Z-DNA information region fragment. All primers used to amplify the storage units are listed in Table 2. All storage units were purified using AgencourtAMPure XP magnetic beads before use. According to the bit values ​​listed in Table 1, the corresponding Z-DNA or conventional DNA storage units were selected and mixed in equal amounts to construct the Z-DNA movable type storage security code A.

[0152] 4. Reading and decoding of Z-DNA security codes

[0153] First, the DNA ends were repaired and A-tailed. After A-tailing, specific barcoded adapters were ligated to the DNA sample. The resulting samples were mixed and used for library construction. The constructed sequencing library was sequenced on the latest version of the Nanoporeer 10.4.1 sequencing chip, and basecalling was performed using Dorado v0.8.2 (https: / / github.com / nanoporetech / dorado / ).

[0154] The decoding of Z-DNA anti-counterfeiting codes includes two main processes: sequencing read grouping and bit value decoding. The read grouping step separates the sequencing reads according to the index sequence. Then the bit value decoding is performed on each read group. Both processes use the percentage of intersection k-mer (PIK) as an indicator. Specifically, in the grouping stage, a k-mer set (A) was constructed from 5'-220bp of the sequencing read using a k-mer size of 13. A series of k-mer sets (B1 to B n The PIK value for each index is calculated sequentially, and the index with the highest PIK value is selected as the grouping result. In addition, a PIK threshold is applied to discard low-quality reads. Sequencing reads show PIK values ​​below 0.15 for the R9.4.1 sequencing chip; reads with PIK values ​​below 0.35 for the R10.4.1 chip are discarded.

[0155] The anti-counterfeiting code samples were sequenced using a Nanopore R10.4.1 chip. An odd number of sequencing reads were randomly selected for each retrieval, and the PIK value of each sequencing read relative to the data sequence was calculated. A PIK threshold of 0.727 was applied to distinguish between Z-DNA and conventional DNA. If the PIK value was lower than 0.727, the read group was classified as Z-DNA and the bit value was decoded as "1." Otherwise, the bit value was assigned to "0." A majority voting process was then used to determine the final bit value. That is, if the number of sequencing reads with a bit value of 1 was greater than the number of sequencing reads with a bit value of 0, the overall bit value of the sequencing read group was 1, otherwise the bit value was 0. The reliability of information reading can be significantly enhanced by increasing the number of sequencing reads and a multi-layer nested majority voting mechanism.

[0156] 5. Experimental verification of the non-replicability of Z-DNA security codes

[0157] The Z-DNA security code A sample was used as a template. Equal amounts of dATP and dZTP were added to the dNTP raw materials for PCR. PCR was performed using the above template and raw materials, as well as universal primers Sall-F / R. The PCR product was recovered using a DNA Clean & Concentrator-5 kit to obtain the amplified security code B. Figure 5As shown, in sharp contrast to the original anti-counterfeiting code A sample, the data bits of the amplified anti-counterfeiting code B can no longer be decoded, fully demonstrating the unique advantage of the movable type storage system of the present application that data cannot be cloned.

[0158] Table 2 Primers used in this example

[0159]

[0160]

[0161]

[0162]

[0163]

[0164]

[0165] Example 2: Anti-counterfeiting and tracing of cotton

[0166] In this example, cotton was used as the research object, and an information storage and anti-counterfeiting traceability solution based on Z-DNA movable type was designed. Specifically, the 64 Z-DNA movable type pairs designed in Example 1 were allocated 64 bits and used to encode key information such as the cotton's origin, growing year, and planting environment. Each bit encodes information based on the presence of a Z-base in the DNA fragment: the presence or absence of a Z-base corresponds to an information bit of 1 or 0 (or vice versa).

[0167] 1. Information Coding Design

[0168] The 64 Z-DNA word pairs were designed to encode 64 bits of information related to cotton, specifically:

[0169] Bits 0 to 9 are used to encode the production base information of the cotton (e.g., specific farm or region).

[0170] Bits 10 to 15 are used to encode the year of cotton planting (for example, the binary code corresponding to 2023).

[0171] Bits 16 to 31 are used to encode cotton growth environment parameters (eg, temperature, humidity, etc.).

[0172] Bits 32 to 64 are used to encode the batch information and quality certification information of the cotton.

[0173] 2. Generation and Storage of DNA Fragment Libraries

[0174] Each Z-DNA character corresponds to a DNA fragment with specific sequence characteristics. The presence of Z-bases is precisely controlled through the DNA synthesis process, ensuring that each DNA fragment accurately reflects the pre-defined bit value. The resulting DNA fragment library contains a variety of DNA sequences, each carrying specific index information and data record information.

[0175] 3. Application and Verification

[0176] The synthesized Z-DNA fragment library is prepared into a solution and evenly sprayed onto the surface of the cotton product. During the subsequent anti-counterfeiting verification process, the Z-DNA fragments are detected using high-throughput nanopore sequencing technology. By analyzing the sequencing data, the PIK value of each sequence read is calculated to determine the presence of the Z-base and ultimately reconstruct the complete 64-bit information.

[0177] 4. Sequencing and Verification

[0178] The decoding process mainly includes two key steps: reading grouping and bit value decoding.

[0179] Read grouping: Sequencing reads are grouped according to the index sequence. Specifically, the sequence is first extracted from the 5' end to 220bp of each sequencing read, and a k-mer set A is constructed with a k-mer size of 13. Subsequently, a series of k-mer sets (B1 to B 64 For each index, the PIK value (percentage of intersection k-mers) with set A was calculated, and the index with the highest PIK value was selected as the grouping result. In addition, to eliminate low-quality reads, a PIK value threshold was set: for the R9.4.1 sequencing chip, PIK values ​​below 0.15 were considered low-quality reads and discarded; for the R10.4.1 chip, PIK values ​​below 0.35 were also discarded. Figure 6 Flowchart for grouping sequencing reads according to index sequence.

[0180] Bit value decoding: After completing the read grouping, the sequence reads within each group are decoded for bit values. In each group, 3 to 11 sequence reads (n = 3, 5, 7, 9, or 11) are randomly selected to construct k-mer sets. Subsequently, these sets are analyzed for intersection with the preset original k-mer set A that does not contain Z-bases, and the percentage of the intersection in the original set A (PIK value) is calculated. By comparing the PIK value with a preset threshold (e.g., 72.7%), it is determined whether each read contains a Z-base: if the PIK value is lower than the threshold, the read is judged to contain a Z-base, and the corresponding bit value is 1; otherwise, the corresponding bit value is 0. Finally, based on the PIK value distribution of the n reads in each group, the final bit value of the group is calculated using the majority voting principle. Figure 7 Flowchart for decoding bit values ​​of sequenced reads within each group.

[0181] Information decoding: After obtaining the correct 64-bit bit value, information such as origin, planting year, growth environment parameters, batch and quality certification can be decoded according to the preset encoding rules.

[0182] The above describes exemplary embodiments of the various inventions of the present application. However, without departing from the essence and scope of the present application, those skilled in the art will be able to modify or improve the exemplary embodiments described in the present application, and the resulting variations or equivalents also fall within the scope of the present application.

Claims

1. An information storage method, comprising: The information is stored in the form of a DNA fragment library using multiple DNA fragments, wherein the information includes multiple bits. Each of the plurality of DNA fragments includes an information recording unit, and the information recording unit can be used for indexing and data recording. wherein the data recording purpose of the information recording unit is realized by whether the DNA fragment contains a Z-base, wherein the DNA fragment containing the Z-base is set to correspond to information bit 1, and the DNA fragment not containing the Z-base is set to correspond to information bit 0, or vice versa, and wherein the indexing is achieved by the specific sequences of the multiple DNA fragments, Optionally, wherein the DNA fragment library is pre-synthesized.

2. A method for reading information, comprising: a) sequencing multiple DNA fragments in the DNA fragment library, Each of the plurality of DNA fragments includes an information recording unit, and the information recording unit can be used for indexing and data recording. wherein the data recording purpose of the information recording unit is realized by whether the DNA fragment contains a Z-base, wherein the DNA fragment containing the Z-base is set to correspond to information bit 1, and the DNA fragment not containing the Z-base is set to correspond to information bit 0, or vice versa, and wherein the indexing is achieved by the specific sequences of the multiple DNA fragments; b) grouping the multiple DNA fragments according to the sequencing results of step a); and c) calculating the PIK value of each grouped DNA fragment based on the k-mer algorithm, and then determining whether the DNA fragment contains Z-base, Optionally, wherein the DNA fragment library is pre-synthesized.

3. The method of claim 2, wherein the sequencing is performed by a method selected from the group consisting of massively parallel sequencing, ion semiconductor sequencing, and nanopore sequencing, and optionally, step b) comprises: Constructing a k-mer set A using the originally designed sequences not containing Z-bases of the various DNA fragments in the DNA fragment library; Each of the DNA fragments is used to construct k-mer sets B1 to B n , where for each DNA fragment, the percentage of its intersection with set A in set A is calculated in turn as the PIK value, and the DNA fragment with the highest PIK value is selected as the grouping result, n is the number of types of the DNA fragments, Preferably, the size of the k-mer is 8-18, preferably 10-16, more preferably 13.

4. The method according to claim 2 or 3, wherein step c) comprises: Constructing a k-mer set A using the originally designed sequences not containing Z-bases of the various DNA fragments in the DNA fragment library; Each sequencing read segment corresponding to the specific sequence of the DNA fragment after grouping is used to construct k-mer sets B1, B2...B n , wherein n is the number of randomly selected sequencing reads, preferably, n is an odd number greater than or equal to 3, more preferably, n=3, 5, 7, 9 or 11; and Combine set A with set B n The percentage of the intersection of the two sets in set A is used as the indicator PIK to distinguish whether the sequencing read is Z-DNA or normal DNA. n value, Preferably, the size of the k-mer is 8-18, preferably 10-16, more preferably 13.

5. The method of claim 4, wherein the PIK of the DNA fragment is n The value is compared with a threshold value to determine whether the DNA fragment contains a Z-base, wherein the threshold value is about 72.7%. Optionally, sequencing reads less than the threshold value are determined to be Z-DNA, and the bit value is 1 or 0. Optionally, after the bit values ​​of all sequencing reads are determined, the final bit value of the sequencing reads in the group is calculated according to the bit value distribution of the n sequencing reads in the group according to the principle of majority voting.

6. An information storage method, comprising: a) pre-synthesizing multiple DNA fragment libraries, each DNA fragment library containing multiple DNA fragments, Each of the plurality of DNA fragments includes an information recording unit, and the information recording unit can be used for indexing and data recording. wherein the data recording purpose of the information recording unit is realized by whether the DNA fragment contains a Z-base, wherein the DNA fragment containing the Z-base is set to correspond to information bit 1, and the DNA fragment not containing the Z-base is set to correspond to information bit 0, or vice versa, and wherein the indexing is achieved by the specific sequences of the multiple DNA fragments, wherein whether each of the plurality of DNA fragments contains a Z-base is predetermined, b) randomly combining one or more of the plurality of DNA fragment libraries for information storage.

7. The method of any one of claims 1 to 6, wherein the information comprises n cryptographic bits, wherein n is greater than or equal to 2, preferably, n is greater than or equal to 16, and optionally, each of the one or more DNA fragments has a length of 20 bp to 2000 bp, preferably 100 bp to 250 bp.

8. An information storage system comprising one or more presynthesized DNA fragment libraries, each DNA fragment library comprising a plurality of DNA fragments, Each of the plurality of DNA fragments includes an information recording unit, and the information recording unit can be used for indexing and data recording. wherein the data recording purpose of the information recording unit is realized by whether the DNA fragment contains a Z-base, wherein the DNA fragment containing the Z-base is set to correspond to information bit 1, and the DNA fragment not containing the Z-base is set to correspond to information bit 0, or vice versa, and wherein the indexing is achieved by the specific sequences of the multiple DNA fragments, Whether each of the plurality of DNA fragments contains a Z-base is predetermined.

9. The DNA fragment library is used for product traceability and anti-counterfeiting purposes. The DNA fragment library contains multiple DNA fragments. Each of the plurality of DNA fragments includes an information recording unit, and the information recording unit can be used for indexing and data recording. wherein the data recording purpose of the information recording unit is realized by whether the DNA fragment contains a Z-base, wherein the DNA fragment containing the Z-base is set to correspond to information bit 1, and the DNA fragment not containing the Z-base is set to correspond to information bit 0, or vice versa, and wherein the indexing is achieved by the specific sequences of the multiple DNA fragments, wherein whether each of the plurality of DNA fragments contains a Z-base is predetermined, Optionally, the product is selected from cotton, traditional Chinese medicine, fruit, and food crops.

10. A method for product traceability or anti-counterfeiting verification, comprising the following steps: a) pre-synthesizing a DNA fragment library, wherein the DNA fragment library comprises a plurality of DNA fragments, Each of the plurality of DNA fragments includes an information recording unit, and the information recording unit can be used for indexing and data recording. wherein the data recording purpose of the information recording unit is realized by whether the DNA fragment contains a Z-base, wherein the DNA fragment containing the Z-base is set to correspond to information bit 1, and the DNA fragment not containing the Z-base is set to correspond to information bit 0, or vice versa, and wherein the indexing is achieved by the specific sequences of the multiple DNA fragments, wherein whether each of the plurality of DNA fragments contains a Z-base is predetermined; b) spraying the DNA fragment library on the surface of the product, and c) sequencing the DNA fragment library sprayed on the surface of the product to reconstruct the anti-counterfeiting code information, and optionally reading and decoding the anti-counterfeiting code using the information reading method according to any one of claims 2 to 5, Optionally, the product is selected from cotton, traditional Chinese medicine, fruit, and food crops.