Security protection using DNA
By using heterogeneous nucleic acid data packets and topoisomerase-mediated ligation technology, highly diverse DNA mixtures are generated, solving the problem of easy counterfeiting in existing identification and authentication methods, and achieving efficient and low-cost product authentication and anti-counterfeiting protection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-12
- Publication Date
- 2026-04-10
AI Technical Summary
Existing identification and authentication methods are easily counterfeited, and the cost of storing and retrieving DNA data is high. The quality and quantity of synthesized DNA are also limited, making it difficult to achieve effective anti-counterfeiting protection.
Using heterogeneous nucleic acid data packets, DNA sequences are synthesized through topoisomerase-mediated ligation, data is encoded and integrated into items, highly diverse DNA mixtures are generated using heterogeneous box data writing technology, and blockchain technology is used for verification.
It provides a highly unique and difficult-to-counterfeit method for item authentication and traceability, reducing the time and resource costs of DNA data storage and retrieval, and improving the reliability of anti-counterfeiting protection.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
[0001] Cross-references to related applications The following commonly owned U.S. Provisional Patent Applications Nos. 63 / 582,199 and 63 / 623,085 contain content relevant to the subject matter described herein, and each of these applications is incorporated herein by reference in its entirety to the fullest extent permitted by applicable law.
[0002] Invention Field This invention relates generally to the field of synthetic biology, and more specifically to methods for anti-counterfeiting protection, identification and / or data embedding using DNA sequences. Background of the Invention Numerous markets demand reliable, durable, and accurate identification and authentication methods. These markets may include, but are not limited to, luxury goods, collectibles, art, wines, spirits, raw materials (such as primary minerals, processed minerals, and intermediate materials), currency, or any other physical item requiring inherent embedded identification, authenticity, data, and / or traceability information. Counterfeit goods result in lost revenue, reputational damage, brand dilution, and circumvention of safety and sustainability standards. Appropriate authentication methods enable the tracking of goods imported or exported and / or the verification of the origin of items. Previous methods for meeting this need have included adding external markings to product packaging or the product itself. External markings include watermarks, holograms, serialization marks, engraving, microprinting, smart labels (such as QR codes), specialty inks, guilloché patterns, and micro-coatings (such as dust identification). However, external markings remain vulnerable to counterfeiting. As an alternative, intrinsic markings embedded within products have been developed to further increase the difficulty of counterfeiting, such as radio frequency identification (RFID) tags, near field communication (NFC) tags, spectral and / or isotopic fingerprinting, and blockchain tracking. Unfortunately, the application of intrinsic markings is limited, and their anti-counterfeiting capabilities may become more vulnerable as technology advances.
[0004] It is known that DNA can be encapsulated in nanoscale silica microbeads, which can be incorporated into various materials for printing or casting articles of arbitrary shapes, and can subsequently be recycled. For example, see Koch J et al.'s paper... A DNA-of-things storage architecture to create materials with embedded memory. Molecular code "Nat. Biotechnol. (2020) 38(1):39-43; for example, U.S. Patent No. 9,850,531, " systems A hydrofluoric acid-free method to dissolve and For example, Bossert et al., quantify silica nanoparticles in aqueous and solid matrices Terminator-Sci. Rep. (2019) 9:7938, the contents of each of which are incorporated herein by reference. However, embedding or integrating identifying DNA into an item is not an effective anti-counterfeiting method if the DNA can be easily extracted, amplified, and embedded into counterfeit goods.
[0005] Furthermore, the synthesis and retrieval of data stored in DNA can be costly in terms of time, resources, and cost. Methods developed to optimize time, reagent usage, and decoding efficiency can provide improved DNA synthesis and anti-counterfeiting protection methods. Furthermore, DNA data storage with single-base precision, while enabling high-density data storage, requires additional time and material resources to encode and retrieve user-defined data. Such processes can limit the quality and quantity of synthesized DNA. However, some data encoding methods have been developed that do not require single-base precision. See, e.g., Lee, HH, et al., “ free template-independent enzymatic DNA synthesis for digital information storage. Figure 1 ”, Nat. Commun. (2019) 10:2383, the contents of which are incorporated by reference.
[0006] There remains a need for improved methods of authentication and anti-counterfeiting protection for goods.
[0007] SUMMARY DNA can be a useful material for item authentication and item provenance tracking, where data is encoded in one or more DNA sequences, integrated into a target item, and subsequently extracted and analyzed. Analysis of such DNA sequences can provide a “fingerprint”; for example, product fingerprints can be generated by various methods such as (A+T) / (G+C) ratio determination, restriction fragment length polymorphism (RFLP), mass spectrometry (MS), and DNA sequencing, enabling detection, tracking, and / or authentication of one or more DNA sequences integrated into an item. Indeed, in our previous work (e.g., U.S. Patent Application Nos. 63 / 582,199 and 63 / 623,085, the contents of which are incorporated by reference), the advantages and applications of the near-infinite design space provided by DNA marker combinations have been discussed.
[0008] In one aspect, the present disclosure relates to a new population of deoxyribonucleic acid (DNA) sequences encoding data that can be used for item authentication and anti-counterfeiting protection, the population of DNA sequences comprising nucleic acid data packets ("nackets"), wherein each nucleic acid data packet comprises a plurality of DNA molecules encoding the same data, wherein the sequences of the DNA molecules are heterogeneous. Herein, a nucleic acid data packet can also be referred to as a DNA (or polymer) storage string or storage chain. For example, a nucleic acid data packet can be prepared by heterogenous (or heterogeneous or diverse) cassette data writing, wherein two or more cassette sequences correspond to (or are associated with or represent) a single bit or bit combination in a machine-readable code (e.g., binary code), such that all or nearly all of the DNA molecules in the nucleic acid data packet encode the same data, but the sequence representation of each molecule exhibits extremely high diversity, e.g., due to the use of heterogenous cassettes encoding the same data bit, e.g., wherein the percentage abundance of different cassette variants used to write the nucleic acid data packet provides a unique and distinguishable signature of the nucleic acid data packet.
[0009] In some embodiments, a nucleic acid data packet is synthesized using one or more transferases (e.g., terminal deoxynucleotidyl transferase (TdT)). For example, a nucleic acid data packet can be prepared by stepwise addition of non-identical nucleotides forming homopolymer extensions within a DNA sequence, wherein the transition from a first homopolymer extension to a second homopolymer extension comprises a transition between non-identical nucleotides, and the transition between non-identical nucleotides corresponds to (or is associated with or represents) a single bit or bit combination in a machine-readable code (e.g., ternary code), such that the population of DNA molecules encodes a desired data string.
[0010] In some embodiments, a nucleic acid data packet is synthesized using topoisomerase-mediated ligation. For example, synonymous cassettes having different sequences but encoding the same information can be added at each addition step to construct a set of DNA polymers, wherein each polymer has a series of information cassettes encoding substantially the same information, but the polymers are heterogeneous at the sequence level.
[0011] Nucleic acid data packets can be integrated into or associated with an item for identification and authentication of the item. In certain embodiments, nucleic acid data packets are adsorbed to (or encapsulated in) silica microbeads or particles (optionally coated with a polymer), integrated into an item, e.g., for identification and authentication of the item. In certain embodiments, nucleic acid data packets are added to an ink, e.g., a water-soluble ink, optionally comprising a polymer, e.g., for identification and authentication of signatures, documents, and printed matter. In certain embodiments, nucleic acid data packets are adsorbed to (or encapsulated in) silica microbeads or particles after synthesis or preparation. In alternative or additional embodiments, nucleic acid data packets are adsorbed to (or encapsulated in) silica microbeads or particles during synthesis or preparation, e.g., during one-pot synthesis of the nucleic acid data packet and the silica microbead or particle. In certain embodiments, nucleic acid data packets are integrated into ceramic or silica microbeads or particles by a sol-gel process comprising reacting a molecular precursor (e.g., a silicate, e.g., tetraethyl orthosilicate) with water in an alcoholic solution comprising the nucleic acid data packet, and condensing the product to form a cross-linked particle structure, wherein the nucleic acid data packet is contained within the cross-linked structure, e.g., a Stöber nanoparticle reaction.
[0012] In another aspect, the present disclosure relates to methods of marking, identifying, and authenticating an item, the methods comprising: (i) marking an item to be identified or to be authenticated by integrating into or associating with the item a nucleic acid data packet as described herein; and (ii) identifying and authenticating the item so marked by extracting the nucleic acid data packet and sequencing it, identifying the item based on encrypted data (e.g., binary encoded data, e.g., ternary encoded data) in the nucleic acid data packet, and authenticating the item by (i) measuring the relative amounts of different cassette variants and / or (ii) analyzing the DNA sequence (e.g., DNA "fingerprint," e.g., transitions between non-identical nucleotides), (iii) and / or sequencing and decoding the encoded data.
[0013] BRIEF DESCRIPTION OF DRAWINGS Figure 2 A process for topoisomerase-mediated ligation using DNA cassettes with complementary overhangs and 5' phosphate groups for blocking and unblocking and phosphatases is schematically depicted to enable controlled single cassette addition.
[0014] Figure 1 A DNA molecule comprising cassettes connected by the process shown is depicted. Figure 3
[0015] Figure 4 The potential for high diversity of topogation cassettes is illustrated.
[0016] Figure 5 A comparison of two-bit, multi-base encoding versus two-bit, single-base encoding is depicted.
[0017] Figure 6 An illustration of how multi-base encoding (heterogeneous cassettes) is used to generate a highly diverse set of combinations for product fingerprints (i.e., characteristics of the manufacturing process).
[0018] Figures 7-13 An example of cassettes that can be used for homogeneous cassette data writing and heterogeneous cassette data writing is shown, which uses two unique, non-interacting overhangs (A and B) to allow for the addition of one cassette per reaction.
[0019] Figure 14 An illustration of how heterogeneous cassette data writing generates a unique DNA mixture is shown schematically.
[0020] Figure 15 An illustration of the advantages of heterogeneous cassette data writing compared to single-base writing is shown.
[0021] Figure 16 An example of how a 32-byte NFT is encoded into a 16, 12 cassette sequence chain is provided.
[0022] Figure 17 An overview of how nucleic acid data packets are prepared and integrated into a product is provided.
[0023] Figure 18 An overview of how nucleic acid data packets are extracted and analyzed to verify authenticity is provided.
[0024] Figure 19 An illustrative overview of the different roles in the verification process is provided.
[0025] Figure 20 is a schematic diagram showing topocassettes representing various combinations of binary bits, in accordance with embodiments of the present disclosure.
[0026] Figure 21 is a schematic diagram showing the number of potential topocassettes based on the number of positions and the number of different DNA bases, in accordance with embodiments of the present disclosure.
[0027] Figure 22 is a schematic diagram showing how multiple different cassettes can be used to specify the same underlying binary information, in accordance with embodiments of the present disclosure.
[0028] Figure 23 is a schematic diagram showing a comparison of homogeneous cassette data writing and heterogeneous cassette data writing using multiple topocassettes combined in a predetermined recipe or mixture, in accordance with embodiments of the present disclosure.
[0029] Figure 22is a schematic showing the loading of a heterogeneous mixture of topological cassettes of Figure 24 into a print head of an inkjet DNA printer.
[0030] Figure 25A is a schematic showing the process of writing a two-bit binary code into a substrate or matrix surface according to embodiments of the present disclosure.
[0031] Figure 26 , 25B , 25C, 25D, 25E, 25F, 25G, 25H, 25I, and 25J are schematics showing the process of writing a storage string at one site on a substrate using a pre-set cassette formulation or cassette mixture for each 2-bit pair according to embodiments of the present disclosure.
[0032] Figure 27A is a schematic showing the cassettes along a storage string and the cassettes assigned to each two-bit code in a storage string according to embodiments of the present disclosure.
[0033] Figure 28 , 27B , 27C, and 27D are schematics showing the process of verifying a storage string or nucleic acid data package using a pre-determined cassette mixture associated with a given two-bit binary code according to embodiments of the present disclosure.
[0034] Figure 29A is a schematic showing the randomness and verification of a storage string (or nucleic acid data package) in two dimensions along a single storage string and for all storage strings for a given site according to embodiments of the present disclosure.
[0035] Figure 30A , 29B , and 29C are tables showing various assignments between batch number-based binary codes and cassettes and associated cassette mixtures / formulations according to embodiments of the present disclosure.
[0036] Figure 30B is a block diagram of an inkjet printing system showing print head control and wafer array / platform control logic and instrumentation for fluids / reagents according to embodiments of the present disclosure.
[0037] Figure 30A is a block diagram of a computer system of Figure 31A according to embodiments of the present disclosure.
[0038] Figure 31B is a flowchart showing the writing (printing) and offloading of encoded polymer storage strings in an inkjet writing system according to embodiments of the present disclosure.
[0039] Figure 32A is a flow chart showing writing (printing) of two-bit codes to DNA / polymer storage strings in an inkjet writing system according to embodiments of the present disclosure.
[0040] Figure 32B is a side view schematic showing a plurality of sites with encoded DNA and a cleaving fluid for removing encoded DNA strands from a substrate surface according to embodiments of the present disclosure.
[0041] Figure 33 is an array schematic showing sites with encoded DNA, the array having columns (X) of redundant sites written with the same encoded DNA data and rows (Y) of sites written with different encoded DNA according to embodiments of the present disclosure.
[0042] Figure 34 is a schematic showing transfer of spotted DNA from a substrate surface to a collection well and reading and decoding of the DNA collection according to embodiments of the present disclosure.
[0043] Figure 35A is a flow chart showing decoding and validation of polymer storage string data according to embodiments of the present disclosure.
[0044] Figure 36A and 35B is a schematic showing an example of a box constituting an address, data and error check of a written DNA / polymer storage string according to embodiments of the present disclosure.
[0045] Figure 36B is a schematic showing a method for creating a unique encrypted DNA fingerprint according to embodiments of the present disclosure.
[0046] Figure 37 is a schematic showing three layers of data derived from a common DNA sequence according to embodiments of the present disclosure.
[0047] Figure 38A is a schematic showing an encoding / decoding system for encoding digital files into DNA and decoding digital files from DNA according to embodiments of the present disclosure.
[0048] Figure 37 , 38B , 38C, 38D, 38E, 38F are schematics showing a method for encoding a digital file into DNA for writing according to embodiments of the present disclosure. Figure 39A is a schematic showing a system for encoding a digital file into DNA for writing according to embodiments of the present disclosure.
[0049] Figure 37 , 39B , 39C, 39D, 39E, 39F, 39G are schematics showing a method for encoding a digital file into DNA for writing according to embodiments of the present disclosure.Figure 40A schematic of the method of the system of the present disclosure to decode the written DNA back into the original digital file.
[0050] Figure 37 40B and 40C are data graphs showing results data from the encoding / decoding system of the present disclosure using Figure 41
[0051] Figure 42 schematically depicts ternary encoding map or scheme.
[0052] Figure 43 schematically depicts the variable space of unique DNA sequences written using heterologous DNA cassette data.
[0053] Figure 44 shows DNA stability and recovery after 2 weeks and 6 weeks after writing DNA onto paper using pen ink.
[0054] Figure 45 shows DNA stability and recovery after 8 weeks after writing DNA onto paper using pen ink.
[0055] Figure 46 shows the variable space of unique DNA sequences with synonymity when encoding NFT codes.
[0056] Figure 47 shows the recovery efficiency of DNA after accelerated aging of samples written onto paper using pen ink.
[0057] Figure 48 shows the relative frequency of double stranded DNA breaks during accelerated aging of samples written onto paper using pen ink.
[0058] Figure 49 shows that sequence error rate is relatively stable while sequence efficiency decreases over time during accelerated aging of samples written onto paper using pen ink.
[0059] Figure 50 shows the change in sequence length distribution over time.
[0060] Figure 51A is a schematic showing a print head group for a laser jet DNA printer according to embodiments of the present disclosure, the print head group having individual topological cassette nozzles within one head group, and the printer having multiple head groups.
[0061] Figure 51B is a schematic showing an array of sites on a chip / array with encoded DNA according to embodiments of the present disclosure, the array having rows (Y) of sites written with different encoded DNA, each row having a computer-generated random proportion of boxes (Cs) associated with each two-bit code.
[0062] Figure 52 is a schematic showing an array of sites on a chip / array with encoded DNA according to embodiments of the present disclosure, the entire chip having a computer-generated random proportion of boxes (Cs) associated with each two-bit code of a given lot number.
[0063] Figure 53A is a schematic showing a print head set for a laser jet DNA printer according to embodiments of the present disclosure, the print head set having individual topological box nozzles within one head set.
[0064] Figure 53B is a flow chart showing writing (printing) and offloading encoded polymer storage strings in an inkjet writing system using computer-based randomness for box write selection according to embodiments of the present disclosure.
[0065] Figure 54 is a flow chart showing writing (printing) two-bit codes to DNA / polymer storage strings in an inkjet writing system using computer-based randomness for box write selection according to embodiments of the present disclosure.
[0066] Figure 1 is a flow chart showing decoding and confirming polymer storage string data when using computer-based randomness for box write selection according to embodiments of the present disclosure. DETAILED DESCRIPTION The following description of different embodiments is merely exemplary in nature and is in no way intended to limit the present application, its application, or uses.
[0068] We have previously described techniques for information storage using a charged polymer (e.g., DNA) comprising at least two different monomers or oligomers, where the information is encoded in a machine-readable code (e.g., binary code). For example, U.S. Patent No. 11505825, U.S. Patent No. 11655465, and U.S. Patent Application No. 18 / 358,861 filed July 25, 2023 (each incorporated herein by reference), among others, disclose methods for synthesizing DNA molecules using a topoisomerase-mediated ligation reaction, adding information boxes to the DNA strand in the 3’ to 5’ direction.
[0069] The following commonly owned issued patents contain subject matter relevant to this document and are hereby incorporated by reference in their entirety to the maximum extent permitted by applicable law: U.S. Patent No. 10,438,662, U.S. Patent No. 10,640,822. The aforementioned commonly owned patents discuss methods of writing (or storing) data in a charged polymer (e.g., DNA) using plus “0” and plus “1” enzymes and a deblocking enzyme.
[0070] The following commonly owned U.S. Patent Application Serial Nos. 18 / 358,861 and 18 / 444,662 contain subject matter relevant to this document and are hereby incorporated by reference in their entirety to the maximum extent permitted by applicable law. The aforementioned commonly owned patent applications discuss other methods of writing (or storing) data in a charged polymer (e.g., DNA), such as using AB adapters instead of deblocking enzymes, using “A0B” and “A1B” as plus “0” and plus “1” reagents, respectively, and methods of writing DNA cassette strands using an inkjet reaction format, among other things.
[0071] As described herein, the present disclosure provides a novel system for storing (or writing, printing) information (or data) using a charged polymer (e.g., DNA), the monomers of which correspond to machine-readable code (e.g., binary, ternary, or other base code) and can be synthesized in a variety of ways, including using a piezoelectric inkjet printing system, such as the system disclosed in U.S. Patent Application No. 18 / 444,662, filed February 17, 2024, which is hereby incorporated by reference in its entirety to the maximum extent permitted by applicable law.
[0072] Topoisomerases are enzymes that can recognize and cleave at least one strand of a nucleic acid duplex within a nucleic acid segment known as a site-specific recombination sequence. Vaccinia topoisomerase is a type I DNA topoisomerase that can cleave a DNA strand at the 3' end of its recognition sequence 5'-(C / T)CCTT-3' (e.g., 5'-CCCTT-3') and religate or join DNA back together. Oligonucleotide cassettes containing digital information can be ligated together by topoisomerases. In this method, the DNA base cassettes contain a topoisomerase recognition sequence, allowing them to be “loaded” by the topoisomerase, such that the DNA strand is cleaved by the enzyme, forming a transient covalent bond with the topoisomerase at the 3' end. When a suitable DNA acceptor is found, the topoisomerase ligates the DNA cassette to the DNA acceptor strand, a process known as “bit addition” or “topogation.” After the DNA cassette is ligated to the DNA acceptor strand, the topoisomerase is no longer bound to the DNA.
[0073] Figure 2 A topoisomerase-mediated ligation method is depicted that utilizes DNA cassettes with complementary overhangs, and 5' phosphate blocking and deblocking with phosphatase to achieve controlled single cassette addition. In other embodiments, blocking and deblocking can be achieved by thermal reactive moieties, photo-reactive moieties, enzyme-reactive moieties, or combinations thereof. In a simple embodiment, there are two cassette libraries that can be added one-by-one to a DNA strand to form a binary code sequence, such as sequence X or sequence Y. Thus, if sequence X is designated as 1 and sequence Y is designated as 0, then a binary sequence of 1001 can be encoded by forming a strand containing the series of cassettes X-Y-Y-X. Each cassette can further contain a spacer region, and / or the cassettes can be separated by one or more spacer regions, where the spacer regions can contain a topoisomerase recognition sequence and a short complementary sequence as a residual sequence for the topological ligation process, as shown. Figure 3 The cassettes can contain multiple bits (e.g., XX, XY, YX, YY) to construct information sequences with fewer operations. But in these cases, the library of cassette sequences to extract is homogeneous - e.g., all "X"s have one signature sequence and all "Y"s have a different signature sequence.
[0074] In this disclosure, a variety of determined sequences can encode a particular bit or combination of bits. Topoisomerase cassettes can be highly diverse. As shown, Figure 4 different cassettes of different lengths and base compositions can be made to encode the same or different bits. While the linker sequence is conserved, the sequence used to convey information need not be. For example, bit X can be encoded by different sequences X1, X2, X3, or X4, and bit Y can be encoded by Y1, Y2, Y3, or Y4. This allows for heterogenous cassette data writing, such that a very large variety of different sequences can encode the same data. This allows for multiple layers of information - binary code information superimposed on a more complex mixture of sequences, forming layered data suitable for product identification. For example, in a product identification, the first layer of data can be considered the product appearance and label (easiest to copy), the second layer the binary code encoded by a series of cassettes (more difficult to copy), and the third layer the precise mixture of heterogenous cassettes used to encode the binary data (very difficult to copy). Specific examples are shown in Figure 5 and Figure 5 .
[0075] In particular, Figure 4 Two different cassette formulations or mixtures (mixl, mix2) are shown. See Figure 5 and Figure 6When employing dual-bit binary encoding, each dual-bit combination can be represented by Y different cassettes in a particular formulation. Formulations can be formulated in different proportions to increase the complexity of the combinations. For example, assuming each potential sequence in the formulation is employed in integer percentages, (100^Y)^4formulations are possible. Furthermore, if there are N encoding bases (or sites) in each 15-base representation unit, and all 15 base sites are used, 4^N or 4^15 variants are possible.
[0076] Figure 6 An example of cassettes that can be used for homologous cassette data writing and heterologous cassette data writing is shown, employing two unique and non-interacting overhangs (A and B). Here, the A overhangs (top strand CACT, bottom strand GTGA) are complementary to each other, and the B overhangs (top strand GGCA, bottom strand CCGT) are complementary to each other, but the A overhangs are not complementary to the B overhangs, allowing for the addition of one cassette in each reaction without the protection / deprotection steps generally described in U.S. Patent Application No. 18 / 358,861. In this system, four cassettes are needed to provide binary (0, 1) codes, e.g., A0B, A1B, B0A, and B1A. But in the heterologous cassette data writing example, there are two different information sequences for 1, and two different information sequences for 0, so there are eight different cassettes in total. Furthermore, the proportions of these cassette types can be numbered (e.g., 50% / 50% or 25% / 75% as shown), resulting in DNA populations with the same binary encoded information, but different sequences, and different proportions of each type of cassette.
[0077] The use of heterologous cassette data writing provides significant room for identification and anti-counterfeiting protection. Each data writing fluid contains two or more unique cassette sequences that are different across the entire writing set, e.g., as in Figure 6 D, d, M, m. The data represented by the cassettes in a given fluid can be the same (e.g., for copyright protection / anti-counterfeiting protection, and to facilitate reading on short read sequencers), or different (for building cassette-based UMIs or for arbitrary random number generation, e.g., for random number-dependent applications). The sequences of the cassettes can be abbreviated as a single letter, where the case of the letter represents AB (lower case) or BA (upper case). In Figures 7-13In this representation, D, d, E, and e all represent 0, and M, m, N, and n all represent 1, as illustrated in the example. Therefore, you can significantly shorten complex DNA sequences and simplify the intuitive interpretation of results. Furthermore, standard processing methods for sequence files (such as FASTA, FASTQ formats, string manipulation, matching, etc.) are compatible with this "data sequence" representation, making a large set of mature tools applicable to heterogeneous data layers. For data writing purposes, for example, a convention is to use the first letter of a set of letters to represent the liquid used to write the relevant nucleic acid data packet during the encoding stage. During decoding, all symbols are used, and the software finds the best match between the component sequence and the most relevant "data sequence" letter. The occurrence rate of each box can be controlled through writing (based on the relative number of different boxes) and subsequently measured by sequencing. This occurrence rate can be used as a unique fingerprint of the reagents used for data writing. Data can also be encoded to the proportion level of each component in the liquid (e.g., batch code). In addition to batch coding, fingerprint information can also be obtained.
[0078] In some embodiments, one or more cartridges may be synthesized using a sequential single-base addition method, such as phosphoramide synthesis. In some embodiments, one or more cartridges may be synthesized using an enzymatic method, such as using one or more DNA polymerases, one or more valve-shaped endonucleases, one or more DNA ligases, or one or more topoisomerases. In some embodiments, one or more cartridges may be synthesized using a sequential single-base addition method and / or an enzymatic method, and then the cartridges may be amplified to increase DNA yield, for example, using PCR (polymerase chain reaction) or RCA (rolling circle amplification). In some embodiments, one or more cartridges are linked together using a method incorporating single-base addition technology, such as phosphoramide chemistry, enzymatic methods (e.g., DNA polymerases, valve-shaped endonucleases, DNA ligases, topoisomerases, or combinations thereof). In some embodiments, the one or more cartridges are synthesized using non-natural nucleotides or nucleobases. In some embodiments, the one or more cartridges may be further modified after synthesis, optionally after being linked with one or more other cartridges, for example, using small molecule moieties, polymers, click-reactive reagents, fluorescent labels, etc.
[0079] Figure 14 This illustration demonstrates how heterogeneous cassette data writing generates unique DNA mixtures. In this example, the binary data for all molecules in the nucleic acid data packet (“nacket”) is 011011, where each cassette represents a single binary bit (0, 1). However, due to the heterogeneity of the cassettes in the writing solution, all written molecules are unique within the same nucleic acid data packet: !EnMeNn#, !EnNdMm#, !EnNeNn#, !EnNeNm#, !DmNeMn#, !DmNeNn#, DmNdNm and DnMeNn are the sequences of the eight molecules generated, where "!" is the start or acceptor string and "#" is an end cap at the end of the nucleic acid package or storage string. The start and end strings can contain other features useful for data storage or authentication; for example, these regions can contain unique "primer regions." The number of permutations is approximately: (number of unique strands or cassettes in each liquid)^(number of rounds of cassette addition). For example, if six rounds of addition are performed (i.e., six cassettes) and each liquid contains two unique strands (or cassettes), then each nucleic acid package has 2 6 or 64 unique molecular permutations. For a strand containing 150 cassettes, if each liquid contains four unique strands (or cassettes), then each nucleic acid package has approximately 4 50 or approximately 2 x 10 90 unique molecules.
[0080] Each read of the nucleic acid package produces three layers of data: a. Package data layer (in this case 011011): Each package ID corresponds to a numerical value.
[0081] b. Generation batch fingerprint: The percentage abundance of different cassette variants used in the writing process is determined. The original fingerprint can be stored in a blockchain.
[0082] c. Item fingerprint: The list of random sequences obtained from each read. The decoded sequences have unique values. During verification of the read, a certain number of numbers must match the original test sequence.
[0083] This approach has several advantages: • All three layers of data are in the same DNA sequence, and they are inseparable.
[0084] • The top data layer allows the sequence to carry numerical data, which can be associated with one or more elements of a blockchain, a public key infrastructure, a digital identifier for any proprietary or public information system, and / or digital data of any size.
[0085] • External systems such as a blockchain, a public key infrastructure system, or other data systems can store information related to the verification of the other two layers of data.
[0086] • A public blockchain can be used, so even if the company that synthesized the DNA goes out of business, the data will remain.
[0087] • There are many technical means of DNA sequencing. Other systems can be iteratively replaced, but DNA can always be read.
[0088] DNA is created by a series of cassette connections (e.g., topological connections) rather than single base additions, which has significant advantages in terms of longer chain length and greater arrangement space. Figure 15 The effect of changing the efficiency of single base chemical coupling on the yield of synthesis is demonstrated in comparison to the cassette data writing described herein. When using single base synthesis, blocks shorter than 6 base pairs are prone to sequencing errors, ligation yield issues, and are easily forged; and for all single base synthesis processes, blocks longer than 7 base pairs do not achieve ideal yield. Long sequences of DNA or groups of DNA sequences can be prepared by, for example, amplification using PCR or phage, and used as identification markers, but such markers are easily forged because the sequences can be easily separated, amplified, and applied to a counterfeit.
[0089] The nucleic acid data packets described herein are particularly well suited for efficient analysis by conventional DNA sequencers (e.g., short read sequencers and / or long read sequencers (e.g., Illumina sequencers)). One skilled in the art can easily know the advantages of each method and the most suitable scenario for short read sequencing and / or long read sequencing: for example, nucleic acid data packets containing 12 or fewer cassettes are suitable for short read sequencing, and nucleic acid data packets containing more than 12 cassettes are suitable for long read sequencing. These nucleic acid data packets are about 2 to 6 kilobases in length, and due to the reuse of cassettes, there are repeated sequences on multiple data chains. Each cassette is about 20 bases in length, which means that a typical read can accommodate about 100 to 300 cassettes. When using heterogenous cassette data writing (e.g., each data writing solution is equipped with 4 types of cassettes), each chain is completely unique before amplification. After amplification, a certain number of chains are "selected" and enriched.
[0090] In certain embodiments, a coding scheme compatible with short read sequencers is used. For example, a nucleic acid data packet containing 10-12 cassettes, each cassette being about 20 base pairs in length, and complete short read sequencing primers in the starting and ending chains each being about 30 base pairs in length, resulting in a nucleic acid data packet of 260-300 base pairs in length. Such a nucleic acid data packet can easily be adapted to various short read sequencers. In this application scenario, the heterogeneity of the cassettes needs to be greater than 4 in typical applications in order to obtain a fingerprint with sufficient complexity. For example, when the heterogeneity is 10, 10 0 ~10¹² unique arrangement ways can be generated. Therefore, the application of long read sequencers has an advantage in reading molecules with more types of variations.
[0091] In certain embodiments, a "rapid fingerprinting" technique can be employed to analyze the nucleic acid data package. In certain embodiments, a rapid fingerprinting technique can provide a preliminary assessment of the nucleic acid data package without sequencing the entire sequence thereof. In certain embodiments, a rapid fingerprinting technique can provide an analysis result in less than 30 minutes, e.g., less than 15 minutes, e.g., less than 10 minutes, e.g., less than 5 minutes, e.g., less than 3 minutes, e.g., less than 2 minutes, e.g., less than 60 seconds, e.g., less than 50 seconds, e.g., less than 45 seconds, e.g., less than 40 seconds, e.g., less than 35 seconds, e.g., less than 30 seconds, e.g., less than 25 seconds, e.g., less than 20 seconds, e.g., less than 15 seconds, e.g., less than 12 seconds, e.g., less than 10 seconds, e.g., less than 9 seconds, e.g., less than 8 seconds, e.g., less than 7 seconds, e.g., less than 6 seconds, e.g., less than 5 seconds. In certain embodiments, rapid fingerprinting includes exposing the nucleic acid data package to a fluorescent probe, an azido-alkyne cycloaddition reagent, an antibody, a microsatellite sequence, or a combination thereof, and / or by employing a restriction fragment length polymorphism (RFLP), amplified fragment length polymorphism (AFLP), or a combination thereof technique. In certain embodiments, a chip platform (e.g., nanochannel or microchannel array) is provided that contains the complementary reagents required to perform rapid fingerprinting, e.g., for field-portable analysis. In certain embodiments, the chip platform contains capture sequences and / or PCR primer sequences complementary to the nucleic acid data package, enabling subsequent authentication, and can optionally contain amplification of the nucleic acid data package. In certain embodiments, the nucleic acid data package contains a terminal nucleotide sequence or a DNA "cap," enabling capture / isolation of the nucleic acid data package, binding to the chip platform, and subsequent authentication.
[0092] In one embodiment, to enable rapid fingerprinting, a mixture of initiator and / or terminator molecules can be used in the write process, where each molecule carries a unique primer sequence that can be recognized by a rapid nucleic acid amplification test (NAAT). This can be a single target validation or a complex molecular fingerprint. Each set of initiator and / or terminator molecules can be mixed and associated with the authentication data, either directly or through a hash function. In one embodiment, the system can contain 32 unique initiator molecules, all of which can be attached to the surface and undergo the first topological attachment reaction, but react with different primers in the NAAT test. When a sample is taken and subjected to the NAAT test, a fingerprint of 32 YES / NO answers can be obtained, which can be translated into a unique ID of 32 bits or 4 billion unique combinations. The ID is not the same for each data write process. In another embodiment, this can be achieved by using 32 initiator molecules and 32 terminator molecules, resulting in 64 bits, 1.8 x 1019arrangements or possibilities. 9 In one embodiment, to enable rapid fingerprinting, a mixture of initiator and / or terminator molecules can be used in the write process, where each molecule carries a unique primer sequence that can be recognized by a rapid nucleic acid amplification test (NAAT). This can be a single target validation or a complex molecular fingerprint. Each set of initiator and / or terminator molecules can be mixed and associated with the authentication data, either directly or through a hash function. In one embodiment, the system can contain 32 unique initiator molecules, all of which can be attached to the surface and undergo the first topological attachment reaction, but react with different primers in the NAAT test. When a sample is taken and subjected to the NAAT test, a fingerprint of 32 YES / NO answers can be obtained, which can be translated into a unique ID of 32 bits or 4 billion unique combinations. The ID is not the same for each data write process. In another embodiment, this can be achieved by using 32 initiator molecules and 32 terminator molecules, resulting in 64 bits, 1.8 x 1019arrangements or possibilities.
[0093] The nucleic acid data package can encode a non-fungible token (NFT), which is a unique digital identifier recorded in a blockchain that verifies ownership and authenticity. It cannot be copied, substituted, or split. Figure 16 An example of how a 32-byte NFT is encoded into a 16-bit, 12-box containing chain is provided. Figure 17 An overview of the process of making a nucleic acid data package and integrating it into a product is provided. Figure 18 An overview of the process of obtaining and analyzing a nucleic acid data package to verify product authenticity is provided. Figure 16 An overview of the different roles in the verification process is provided.
[0094] Specifically, referring to Figure 17 In some embodiments, the first step is to mint an NFT or create a blockchain NFT token and binary code, which can use a public or private blockchain. Next, the second step is to synthesize a DNA strand or string carrying the binary code described herein, which can include a blockchain NFT token, product metadata, and cryptographic fingerprint information. Next, the third step, the DNA can be encapsulated in materials such as silica microbeads or plasmids. Specifically, the DNA is in a stable dry form, the silica further stabilizes the DNA, the optical properties of the item are not affected by the microbeads, they are safe for human consumption, and the microbeads can be extracted from the material and sequenced, and if necessary, the plasmid can be introduced into a living organism. In addition, the plasmid carrying the DNA code can be easily transfected into bacteria, cells, plants, animals, or fungi. Next, referring to Figure 41 , the fifth step is to sample the item embedded with microbeads, etc. Next, the sixth step is to extract the DNA-carrying microbeads from the item and elute the DNA strand or string. Specifically, this step can use known and reliable methods for extracting microbeads from materials to extract silica microbeads or plasmids and separate DNA strands; methods for eluting DNA from microbeads are well known, well studied, and published, and methods for extracting plasmids are also well known to those skilled in the art. Next, the seventh step is to sequence the extracted DNA strand. Any known commercial sequencer (such as Illumina, Oxford Nanopore, or other companies) can be used to read the DNA, and the box sequence can be designed for any sequencing chemistry to achieve optimal performance, and third-party sequencing can also be performed using the global commercial sequencing laboratory network. Next, the eighth step is to verify the binary code encoded by the DNA, which can be in the form of an NFT or NFT hash. Specifically, this step can verify the presence of a blockchain NFT token, product metadata, and / or cryptographic fingerprint, as applicable.
[0095] The nucleic acid data packet can encode one or more public key infrastructure elements that can originate from a private and / or public certificate authority. The use of a certificate authority can be used to mediate the validity of underlying items, e.g., if an item is known to be stolen, the associated certificate can be revoked. Authenticity information can further be stored in a public information system, where the information can be accessed online, e.g., using a public key infrastructure (PKI) to verify the authenticity of a remote server used to authenticate an entity item.
[0096] In another aspect, the present disclosure relates to a nucleotide polymer, such as deoxyribonucleic acid (DNA), synthesized using a terminal deoxynucleotidyl transferase (TdT) in a de novo enzymatic synthesis process. TdT is a template-independent polymerase that extends a DNA "primer strand" by adding one or more deoxyribonucleotide triphosphate (dNTP) monomers to the 3' end of the primer strand. Adenosine triphosphatase (Apyrase) is an enzyme that mediates the degradation of nucleic acid substrates, where the adenosine triphosphatase degrades a nucleotide triphosphate to the corresponding diphosphate or monophosphate precursor; the precursor is inactive to TdT. By optimizing the relative concentrations of TdT and adenosine triphosphatase in the reaction mixture, the enzymes can compete with each other, enabling kinetic control over the stepwise addition of dNTPs to the 3' end of a DNA primer strand. In this way, by iteratively adding dNTPs to one or more primer strands, DNA strands can be produced with short homopolymer stretches, where data (e.g., user-defined data) is encoded in the nucleotide polymer (e.g., DNA), forming a nucleic acid data packet ("nacket"). With this approach, data is not encoded in the specific nucleotide sequence itself, but rather in the transitions between non-identical nucleotides in the polymer.
[0097] In some embodiments, a primer strand is contacted with a reaction mixture comprising TdT and adenosine triphosphatase, where non-identical species of dNTP monomers are added to the reaction mixture in an iterative, stepwise fashion. In some embodiments, the dNTP species include adenosine triphosphate (ATP), guanosine triphosphate (GTP), cytidine triphosphate (CTP), thymidine triphosphate (TTP), and optionally, uridine triphosphate (UTP). In some embodiments, ATP, GTP, CTP, TTP, and UTP can be referred to by their respective nucleobases, i.e., A, G, C, T, and U, respectively; one skilled in the art will appreciate that the use of nucleobase terminology to describe nucleotides in various phosphorylated states is readily understood in the context of the use of the base terminology. For example, if a first addition step includes adding A (i.e., adenosine triphosphate) to the 3' end of a DNA strand in the reaction mixture, a next stepwise addition can include, for example, G, C, or T, but the next stepwise addition cannot be adding A, as it would not encode any additional information on the DNA strand relative to the first addition step.
[0098] In some embodiments, the reaction mixture, DNA priming strand, and dNTPs are contacted under flow conditions. In some embodiments, the reaction mixture, DNA priming strand, and dNTPs are contacted under mixing conditions. In some embodiments, the reaction mixture, DNA priming strand, and dNTPs are contacted in solution, e.g., droplet solution or bulk solution, e.g., without active mixing.
[0099] In some embodiments, the stepwise addition of dNTPs at the 3' end of the DNA priming strand will result in homopolymer stretches of unequal length. For example, in a single addition reaction step, a first DNA strand can be extended by the addition of one or more dNTP monomers (e.g., 2 dNTP monomers), while a second DNA strand can be extended by the addition of one or more dNTP monomers (e.g., 3 dNTP monomers). In some embodiments, the homopolymer stretches of DNA strands can comprise 1 or more dNTP additions, e.g., 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, 35 or more, 40 or more, 45 or more, 50 or more dNTP additions, etc. In some embodiments, the homopolymer stretches of a first DNA strand are independent of the homopolymer stretches of a second, third, fourth, etc. DNA strand.
[0100] In some embodiments, the synthesis reaction results in a set of synthesis strands comprising a series of homopolymer stretches of unequal length. In some embodiments, the set of synthesis strands all comprise the same number and sequence of nucleotide transitions between homopolymer stretches, while the homopolymer stretches are of unequal length.
[0101] In some embodiments, the reaction mixture can comprise aqueous conditions. In some embodiments, the reaction mixture can comprise buffered conditions. In some embodiments, the reaction mixture can comprise other additives, e.g., ions (e.g., cations, e.g., divalent cations, e.g., cobalt ions).
[0102] In some embodiments, the reaction mixture can comprise TdT and apyrase in a ratio of about 10,000: 1 to about 100: 1, for example, the reaction mixture can comprise TdT and apyrase in a ratio of about 5,000: 1 to about 500: 1, for example, 4,000: 1 to about 800: 1, for example, 4,000: 1 to about 1,000: 1, for example, about 4,000: 1 or about 1,000: 1. In some embodiments, the concentration of TdT in the reaction mixture is about 0.1 U / μL to about 10 U / μL, for example, about 0.5 U / μL to about 5 U / μL, for example, about 0.7 U / μL to about 3 U / μL, for example, about 0.8 U / μL to about 2 U / μL, for example, about 0.9 U / μL to about 1.5 U / μL, for example, about 1 U / μL to about 1.2 U / μL, for example, about 1 U / μL. In some embodiments, the concentration of apyrase in the reaction mixture is about 0.1 mU / μL to about 10 mU / μL, for example, about 0.1 mU / μL to about 5 mU / μL, for example, about 0.2 mU / μL to about 2 mU / μL, for example, about 0.25 mU / μL to about 1.5 mU / μL, for example, about 0.25 mU / μL to about 1 mU / μL, for example, about 0.25 mU / μL or about 1 mU / μL.
[0103] In some embodiments, the reaction mixture comprises dNTPs, for example, dATP, dCTP, dGTP, and / or dTTP. In some embodiments, the concentration of dNTPs in the reaction mixture is about 1 μΜ to about 100 mM, for example, about 1 μΜ to about 100 μΜ, about 1 μΜ to about 20 μΜ, about 5 μΜ to about 20 μΜ, about 5 μΜ to about 15 μΜ, about 1 mM to about 100 mM, about 1 mM to about 20 mM, about 4 mM to about 16 mM. In some embodiments, each dNTP is added in the reaction mixture at a concentration independent of the others.
[0104] In some embodiments, user-defined data is encoded in the transition between non-identical nucleotides within a single nucleotide polymer, forming a nucleic acid data packet ("nacket"). In some embodiments, the nucleotides used to synthesize the nucleic acid data packet comprise A, T, C, and G. In some embodiments, the nucleotides used to synthesize the nucleic acid data packet comprise A, T, C, G, and U, optionally, the nucleotides can be further modified, for example, with epigenetic markers, for example, methylation, acetylation, phosphorylation, etc. In some embodiments, one or more non-natural nucleotides can be substituted for or in addition to A, T, C, G, and optionally U. In some embodiments, the sugar groups and / or backbone of the nucleotide polymer can comprise modifications, for example, natural and / or non-natural modifications.
[0105] In some embodiments, data is encoded in the transitions between non-identical nucleotides such that the number of "bits" available is always one less than the number of nucleotides used to encode the data. For example, using the standard nucleotides A, T, C, G as the nucleotides used to encode the nucleic acid data packet, the four available nucleotides allow for three possible transitions from one nucleotide to the next, creating a ternary system, i.e., "trits." For example, if only 3 nucleotides are used to encode the nucleic acid data packet, the three available nucleotides allow for only two possible transitions between nucleotides, creating a binary system, i.e., "bits." For another example, if five nucleotides are used to encode the nucleic acid data packet, the five available nucleotides allow for four possible transitions between nucleotides, creating a quaternary system, i.e., "quits." In some embodiments, the nucleic acid data packet is encoded using three or more nucleotide species, e.g., four, five, six nucleotides. In some embodiments, the nucleic acid data packet is encoded using four nucleotides.
[0106] In some embodiments, to convert user-defined data into a population of nucleotide polymers (e.g., DNA), the information is mapped to a template sequence that comprises an encoding space corresponding to the number of nucleotide species used in the synthesis. For example, if four standard DNA nucleotides are used, the user-defined data is mapped to a trit-based template sequence. To begin encoding data using such a trit-based template sequence, a ternary encoding scheme is first constructed, e.g., Figure 41 The skilled artisan will appreciate that this scheme is merely one example of an available encoding space, and that the scheme shown herein should not be construed as a limiting example. Using such a scheme, a data string can be encoded from trits to DNA nucleotide transitions. For example, if the data string to be encoded comprises, e.g., 0211201, the corresponding transitions between non-identical nucleotides can be represented by the nucleotide sequence CTGTCTATC, where the data string 0211201 is encoded using the ternary scheme Figure 41 However, the skilled artisan will appreciate that the choice of nucleotide sequence is dependent in part on the 3' end of the available DNA strand in the reaction mixture. For example, if the 3' end of the DNA strand available for the reaction is not C, but A, as described above, then using the same ternary scheme Figure 41 shown above, the nucleotide sequence AGCGAGTGA will encode the data string 0211201. Thus, the skilled artisan will appreciate that it is the transitions between non-identical nucleotides that encode the user-defined data string, not the nucleotide sequence itself.
[0107] Furthermore, if non-palindromic data strings are encoded into nucleotide sequences, the complement of the directly encoded nucleotide sequence can yield an inverted data string upon decoding. For example, the data string 10221201 can be directly encoded into the transitions between different nucleotides in the sequence 5'-CTGTAGTGA-3' using a ternary encoding scheme of A DNA-of-things storage architecture to create materials with embedded memory. The complement of the directly encoded nucleotide sequence is 3'-GACATCACT-5', which can be reoriented to 5'-TCACTACAG-3'. Decoding the complement sequence 5'-TCACTACAG-3' using the encoding scheme will yield the data string 10212201, which is the inverse of the original encoded data string 10221201. In some embodiments, the inverted data string can be identified by comparison to a database (e.g., a data string database, an item identification code database). In some embodiments, the encoded data string includes a directional sequence that provides a segment of encoded data to assist in identifying the correct orientation of the encoded data string. In other embodiments, the nucleotide sequence directly encoding the data string and / or its complementary nucleotide sequence includes one or more nucleotide sequences and / or identifying modifications that physically and / or chemically mark the nucleotide sequence and assist in determining the correct orientation of the encoded data string.
[0108] The present disclosure relates, in part, to the synthesis of DNA sequences for encoding data that can be used for item authentication to prevent counterfeiting. The method includes first synthesizing one or more DNA sequences, integrating the DNA sequences into an item, extracting the DNA sequences from the item when authentication is required, and analyzing the DNA sequences to confirm the authenticity and / or origin of the item. By encoding an identification code into a DNA sequence, an encryption system with high entropy (i.e., large permutation space) can be obtained for item identification. For example, if a heterogeneous DNA cassette library containing four different oligonucleotide sequences is used to encode a single bit of data, and 150 rounds of cassette addition are performed, each round using a different set of four oligonucleotide sequences, 4^150 (i.e., 2 x 10^90) different permutations of DNA sequences can be synthesized, each encoding the same item identification code. This process can provide a large permutation space, effectively eliminating random or estimated counterfeiting. Furthermore, if the DNA sequence is extracted from the item and amplified and then attempted to be integrated into a counterfeit product, amplification bias inherent to the DNA replication process will be introduced; such bias is easily identified when further analyzing the potential counterfeit item.
[0109] Accordingly, the present disclosure provides methods for confirming the authenticity and / or origin of an item by integrating a DNA sequence that can be subsequently extracted from the item and identified.
[0110] DNA is a relatively stable molecule that can be easily integrated into or associated with an item of commerce to enable identification and authentication of the item of commerce. In certain embodiments, nucleic acid data packets are adsorbed to silica microbeads or particles (which can optionally be coated with a polymer) and integrated into an item of commerce, e.g., for identification and authentication of the item of commerce. For example, DNA nucleic acid data packets can be integrated into silica microbeads, e.g., using the methods described in Koch J, et al. “ Figure 41 Terminator-free template-independent enzymatic DNA synthesis for digital ” Nat. Biotechnol. 2020, 38 (1):39-43), the contents of which are incorporated herein by reference.
[0111] In certain embodiments, nucleic acid data packets are integrated into an article of manufacture by direct surface conjugation. In other embodiments, nucleic acid data packets are encapsulated in microcontainers or molecular assemblies. In certain embodiments, these encapsulated DNA sequences are integrated into components or materials used to manufacture articles of manufacture, e.g., textiles, fabrics, leathers, biomaterials articles, polymers, plastics, wood, metals, inks, paints, solutions, suspensions, and raw materials. In certain embodiments, nucleic acid data packets are inserted into one or more cells, or into larger DNA constructs and / or genomes, e.g., into yeast, bacterial, fungal, plant, or animal cells, e.g., wherein the cells are used to produce food, drink, biological articles, or materials, such as cheese, beer, wine, vegan leather, pharmaceuticals.
[0112] In certain embodiments, nucleic acid data packets, optionally integrated into (e.g., adsorbed and / or encapsulated in) microbeads (e.g., silica microbeads) are embedded, adhered, or mixed into any physical material. For example, sprayed into minerals, ores, or intermediate raw materials; embedded in polymer films and used to manufacture any device or product; embedded in adhesives and used in product manufacture or labeling; embedded in inks (e.g., for embossing, writing, printing, inkjet printing, screen printing, or otherwise transferred to another substrate); embedded in perfumes; embedded in the ink used by notaries to sign documents; embedded in paper currency paper and / or ink; embedded in the packaging of alcoholic beverages, spirits, and / or food; embedded in the food itself (e.g., wine, cheese, spirits); embedded in animals to track their origin (commercial or biological conservation purposes); embedded, sprayed, or applied to wood products to track origin; sprayed or integrated into seeds to trace seed origin / authenticity; embedded in pharmaceuticals and / or printed on the surface of pharmaceuticals to enable authenticity, drug typing, identification, tracking, and / or embedded authentication; embedded in aviation components to enable tracking; embedded in lock-tite or equivalent thread lockers to identify authenticity, part number, locker, and / or locker time. The skilled artisan will readily think of numerous other and / or alternative applications.
[0113] In certain embodiments, the nucleic acid data package integrated into the item is extracted from the item; this extraction can be done before or after production, transportation, sale, offer for sale, import, or export of the item. In certain embodiments, the extraction is used for identification, authentication, and / or valuation of the item.
[0114] In certain embodiments, the nucleic acid data package integrated into the item is extracted from the item by physical and / or chemical means, such as cutting, grinding, scoring, dicing, shredding, pulverizing, dissolving, or cleaving the nucleic acid data package from one or more portions of the item.
[0115] In certain embodiments, the nucleic acid data package extracted from the item is isolated and / or purified; this can be achieved by chromatography, electrophoresis, centrifugation, or a combination thereof.
[0116] In certain embodiments, the nucleic acid data package extracted from the item is analyzed using mass spectrometry and / or high-throughput DNA sequencing. In certain embodiments, the analyzed DNA sequence is compared to a database of item identification codes, wherein a match between the extracted DNA sequence and an item identification code confirms the identity, authenticity, provenance, and / or security of the item. In certain embodiments, the results of the analysis of the DNA sequence can be compared to previous DNA sequence analysis results of the same or similar items.
[0117] In certain embodiments, the analysis of the nucleic acid data package can yield a "fingerprint," wherein the specific DNA sequence, the specific box sequence, the transition sequence between different nucleotides, the occurrence rate of each individual nucleotide and / or box, the relative occurrence rate of nucleotides and / or boxes, and / or the specific molecular mass of the DNA sequence and its degradation products can be compared to a database of item identification codes.
[0118] In certain embodiments, the nucleic acid data package is analyzed to identify its specific nucleotide sequence, such that the identified nucleotide sequence can be used in conjunction with the original encoding scheme to decode the original encoded data string. For example, a nucleic acid data package comprising a series of transitions of different nucleotide homopolymer stretches can be sequenced. Such nucleic acid data packages (e.g., synthesized using the methods described above) can vary in total length and comprise homopolymer stretches of variable length. However, upon sequencing of the nucleic acid data package, its nucleotide sequence can be compressed, wherein each homopolymer stretch is represented as a single nucleotide corresponding to the identity of the nucleotide comprising the homopolymer stretch. For example, continuing the synthesis example described above, the nucleic acid data package sequences, e.g., CCCCCCTTGGGGGGGGGGTTTTTCCCTTTTTTTTAAAAAAAATTTTTTTCC and / or AAAAGGGCCCGGGAAAAGGGGTTTTTGGGGGGGGAAAAAA, can be simplified to the compressed representative sequences CTGTCTATC and AGCGAGTGA, respectively. Continuing this example, if it is known that the item is a certain brand of a certain model of a certain make of a certain year, then the compressed representative sequence AGCGAGTGA can be used to decode the original encoded data string, e.g.,information storage. For example, the exemplary scheme of FIG. 1, the compressed representative sequences can be decoded to the original data string 10211201.
[0119] In certain embodiments, one or more nucleic acid data packets can comprise synthetic errors, such as one or more mismatched nucleotides, one or more inserted nucleotides, one or more deleted nucleotides, or a combination thereof. In certain embodiments, two or more nucleic acid data packet populations are sequenced and analyzed. In certain embodiments, two or more nucleic acid data packet populations are sequenced, simplified to compressed representative sequences, and then computer analyzed. In certain embodiments, the compressed representative sequences are sorted by length, e.g., where the longest sequence is considered a “perfect” sequence when it matches the original encoding template sequence, and then decoded to obtain the original data string. Alternatively or additionally, the compressed representative sequences can be sorted by abundance, where the most abundant compressed representative sequence is selected and analyzed, optionally where the most abundant compressed representative sequence is further analyzed using statistical inference methods and / or models, such as those described in Lee H.H. et al. “ Figures 7-13 Figure 19 ”Nat. Commun. 2019, 10:2383), the contents of which are incorporated by reference herein.
[0120] The present disclosure provides methods of confirming the authenticity and / or provenance of an item by integrating into a DNA sequence that can be subsequently extracted and identified from the item.
[0121] In one aspect, the present disclosure provides a method of item authentication comprising: i. synthesizing a nucleic acid data packet having a heterologous sequence but encoding the same data in a machine-readable code, such as a binary or ternary code; ii. integrating the nucleic acid data packet into or onto an item; iii. extracting the nucleic acid data packet from the item; and iv. analyzing the extracted nucleic acid data packet; v. optionally, comparing the analyzed nucleic acid data packet to a database of DNA sequences, an authentication database, or a cryptographic hash value; vi. optionally, confirming the authenticity of the item.
[0122] In certain embodiments, the cassettes used in the above methods to synthesize the nucleic acid data package are DNA oligonucleotide sequences comprising a 5' overhang of one or more nucleotides, a region encoding identification code data, a region complementary to an adjacent cassette on one or both sides of the current cassette, a topoisomerase recognition sequence, and / or a 3' overhang of one or more nucleotides. In certain embodiments, the region encoding identification code data comprises one or more bits of data, optionally two or more bits of data, optionally three or more bits of data, optionally five or more bits of data. In further embodiments, the region encoding identification code data comprises one or more bytes of data, optionally two or more bytes of data, optionally three or more bytes of data.
[0123] In certain embodiments, the cassettes are conjugated together using a ligase. In further embodiments, the cassettes are conjugated together using a topoisomerase, optionally wherein the topoisomerase is a Type I topoisomerase (e.g., Type IA, Type IB, Type IC, or a combination thereof), optionally wherein the topoisomerase is a Type II topoisomerase (e.g., Type IIA, Type IIB, or a combination thereof).
[0124] Thus, in one aspect, the disclosure provides the above-described method of article authentication, wherein in the step of synthesizing a nucleic acid data package having heterologous sequences but encoding the same data in a machine-readable code (e.g., binary or ternary code), the nucleic acid data package is synthesized by a method comprising a series of topoisomerase-mediated ligation steps, wherein in each step, heterologous cassettes having at least two different sequences but both encoding the same data in a machine-readable code (e.g., binary or ternary code) are ligated to a DNA strand population by a topoisomerase-mediated ligation reaction, thereby providing a nucleic acid data package having heterologous sequences but encoding the same data in a machine-readable code, wherein the nucleic acid data package comprises a series of topoisomerase-ligated heterologous cassettes.
[0125] In another aspect, the disclosure provides the above-described method of article authentication, wherein the nucleic acid data package having heterologous sequences but encoding the same data in a machine-readable code (e.g., binary or ternary code) is synthesized using a transposase-based synthesis and data encoding method.
[0126] In certain embodiments, one or more DNA sequences are synthesized to encode data designed as an article identification code. In certain embodiments, the identification code is manually written. In further embodiments, the identification code is one or more numbers generated at random.
[0127] In certain embodiments, the nucleic acid data packets are synthesized from a point of attachment on a surface, or in solution. In certain embodiments, the nucleic acid data packets are synthesized in a well plate, droplet, or chamber, wherein each well / droplet / chamber is used to synthesize one or more unique DNA sequences, wherein the DNA has a unique sequence profile but retains the data (e.g., binary or ternary code) encoded in the nucleic acid data packet. In certain embodiments, the nucleic acid data packets are amplified and / or replicated, optionally wherein amplification bias is utilized to further make the set of DNA sequences unique. In other embodiments, the one or more DNA sequences are not amplified and / or replicated, and are used directly for integration into the article.
[0128] Molecules produced using topological ligation of heterogenous components (e.g., molecules produced during surface conjugation of topological ligation) can form a large number of unique molecules (if the arrangement space is large enough). Thus, even using the exact same reagents, procedures, and data, two production runs will not produce the same population of molecules. However, to enable effective authentication, the population of molecules produced must be known and securely stored for later authentication. To this end, an aliquot of the population of molecules produced is isolated and amplified using nucleic acid techniques (e.g., PCR, LAMP, isothermal amplification, and / or RCA). The result of this amplification is a solution containing a large number of copies of the original unique molecules produced in the nucleic acid data packet. This enables the marking of a very large number of articles and / or very large areas of material, while always maintaining a consistent and complex fingerprint. Using the methods described herein, a large amount of data can be embedded in an article through a large number of molecules. However, this method can be vulnerable to amplification cloning attacks, i.e., a forger can sample DNA embedded in a real article, amplify the sample, and embed a forged copy of it in a forged article. The methods described herein resist such attacks through sample complexity; however, higher levels of security can be deployed when amplification attacks are of concern.
[0129] Preventing amplification attacks can be accomplished by writing a single nucleic acid data package on a very large surface area. This single nucleic acid data package can be determined from the authentication database (i.e. the specific data being sought) or reference the data file embedded in the item from which the data amplification fragments are decoded. In one embodiment, an NFT can be written to DNA, amplified to large volume, and the resulting nucleic acid data package embedded in the item. A hash value, CRC, or other hash function can be calculated and then written on a very large surface area with molecules of different length (even very close in length) than the original molecule. These hash molecules do not amplify, so there are no replicates of the hash molecules. These hash molecules are then applied to the item as a second step, or covertly applied to only specific areas that need to be sampled, such that the amplification material is spread throughout the item, but only the specific areas covertly contain the hash molecules. Alternatively, the hash molecules are mixed with the original amplification material and embedded at low abundance (e.g. 0.01%, 0.1%, or 1%). At authentication, the ID is read and authenticated. Then, the authentication program calculates the file hash value and then searches for matching sequences against that hash value or other unique data string. The actual sequence of the molecule found should not be duplicated. If enough sequencing is done and a sufficient number of unique hash molecules are found, then any duplicated hash molecules indicate an amplification attack has occurred. From a sequence and molecular perspective, such molecules are nearly indistinguishable from the correct molecule; therefore, traditional molecular biology methods cannot filter or resolve the hash molecules alone.
[0130] An authentication database can be used to verify the sequence; however, there are several considerations in the authentication design that can be mitigated by the information system design. The considerations are as follows: • Privacy: A user can not want explicit data corresponding to their NFT and / or item to be included in the authentication database that is public and / or privately published. The item itself can have a serial number and / or unique fingerprint such that it can be verified in a public ledger, but not identify which items are in the ledger.
[0131] • DNA Sequence Attack Security: Authentication databases containing actual sequences are vulnerable to attacks, much like insecure password tables. Modern IT systems no longer store plaintext passwords; instead, they use hash tables for password management to prevent user password leaks. This introduces two potential risks: 1) Plaintext sequences can be used to create forged authentication sequence files and transmitted electronically via endpoints or "man-in-the-middle" attacks; or 2) these sequences can be used to synthesize molecules, thereby creating counterfeit molecules. Authentication fingerprints are secure by storing hash values of these sequences instead of actual sequences, much like how passwords are stored in many digital systems. Furthermore, such databases are resistant to lookup-based attacks by using a salted hash algorithm on each record. In these attacks, attackers obtain the salt value and / or hash algorithm, then use brute-force to calculate the hash values of all possible sequences, and then use a reverse lookup to break the database. By setting a unique salt value for each record, this database is resistant to reverse lookup attacks. In an authentication database, the "username" or lookup value is the hash value of the item data. The data embedded within an item may contain a randomized data segment called the item salt, which ensures that no two files will have the same signature. This is a piece of digital information residing at the file layer, randomly generated during manufacturing, encoded into the item's data layer, and not retained in manufacturing logs, authentication databases, or any other location. In some implementations, the digital information exists only within the item after the manufacturing record is cleared. Therefore, this ensures that only the item holder can verify its authenticity. The authentication hash is calculated using all item data (across all items) and the nucleic acid packet ID of the chain to be inspected. This results in each written nucleic acid packet having a unique salt, and an item can have multiple nucleic acid packets, thus forming a large number of entries.
[0132] • Amplification Attack: The authentication database may also maintain another table containing a "Used Unique Reads" table. This table is calculated using only item hashes, not nucleic acid packet IDs, since nucleic acid packet IDs do not exist. If a new authentication request authorizes a detected molecule, this table can be used to invalidate the authentication request and / or issue a warning about a conflict. This prevents replay attacks and ensures that only the first authentication request is approved based on the provided sequence file.
[0133] In certain embodiments, multiple nucleic acid data packets encoding unique, different codes can be placed in a single item. This is a form of molecular encryption, as one must know the encoding scheme to decode. For example, hundreds of unique IDs can be written using different encoding schemes (e.g., different lengths, different starting sequences, and / or different ending sequences). Thus, decoding such nucleic acid data packets requires that the decoder know the encoding scheme used in advance. This method mimics a zero trust security system. Furthermore, a user can be required to provide the encoding scheme and a list of nucleic acid data packet IDs in a particular order. This information is the "key" to obtaining a particular set of information from the item. The strength of this encoding depends on the number of unique entries and the order in which these entries need to be arranged to decode the target file (or key). This method is very effective, with functionality similar to a zero trust security system.
[0134] In still other embodiments, one or more codes can be read from a given sample using multiple encoding schemes. For example, when running an authentication process that involves multiple links in a chain of trust, each link can have its own unique code and encoding scheme. This allows, for example, the unique code (and fingerprint) of a sampling kit, the unique code (and fingerprint) of an amplification kit, and the unique code (and fingerprint) of a target item to be read. When this information is combined, it can be used to ensure that a given combination of kits and items occurs only once. This further enhances the ability of the authentication process to resist replay attacks, man-in-the-middle attacks, and / or the ability to counterfeit authentication test reagents.
[0135] In certain embodiments, one or more cassettes are synthesized on a chip, e.g., a chip comprising a plurality of wells and / or ligation sites on a surface. For example, a chip can comprise a plurality of wells and / or ligation sites on a surface that allow for the synthesis of a plurality of heterologous sequences corresponding to one or more information sequences, e.g., a plurality of sequences (heterologous sequences) corresponding to a "0". A similar chip can allow for the synthesis of a plurality of sequences (heterologous sequences) corresponding to a "1". In certain embodiments, the cassettes synthesized on a chip comprise a replication / amplification primer region (e.g., a PCR primer region) to allow for amplification. In certain embodiments, the chip comprises a replication / amplification primer region (e.g., a PCR primer region) on the acceptor strand prior to addition / synthesis of the cassettes. In certain embodiments, the cassettes comprise sticky ends or overhanging ends to facilitate ligation of the cassettes. For example, a first plurality of cassettes can be synthesized on a "0" chip, and a second plurality of cassettes can be synthesized on a "1" chip, wherein all of the cassettes comprise independently selected overhanging ends (the independently selected overhanging ends of each cassette can be the same, similar, or unique), and a binary coded sequence is synthesized by sequentially adding and ligating cassettes from either the first plurality of cassettes ("0" cassettes) or the second plurality of cassettes ("1" cassettes).
[0136] In some embodiments, one or more DNA sequences are integrated into an article via direct surface conjugation. In other embodiments, one or more DNA sequences are encapsulated in microcontainers (e.g., microspheres, such as silica microspheres). In some embodiments, these microcontainers are integrated into components or materials used to manufacture the article, optionally said components or materials being textiles, fabrics, leather, biomaterials, polymers, plastics, wood, metals, inks, coatings, solutions, suspensions, and raw materials. In some embodiments, one or more DNA sequences are inserted into one or more cells, optionally into larger DNA constructs and / or genomes, optionally into yeast, bacteria, fungi, plant, or animal cells, optionally said cells being used to produce food, beverages, bioproducts, or materials, such as cheese, beer, wine, vegan leather, or pharmaceuticals.
[0137] In some embodiments, one or more DNA sequences integrated into the article are extracted from the article. In some embodiments, the extraction is performed before or after the article's production, transportation, sale, offer for sale, import, or export. In some embodiments, the extraction is used for the identification, authentication, and / or valuation of the article.
[0138] In some embodiments, one or more DNA sequences integrated into the article are extracted from the article by physical means, such as cutting, grinding, scoring, shaving, tearing, or shredding one or more portions of the article. In other embodiments, one or more DNA sequences integrated into the article are extracted from the article by chemical means, such as dissolving or cleaving the DNA sequences from one or more portions of the article.
[0139] In some embodiments, one or more DNA sequences extracted from the article are isolated and / or purified. In some embodiments, one or more DNA sequences extracted from the article are isolated and / or purified using chromatography (e.g., ion exchange chromatography, size exclusion chromatography, normal-phase or reversed-phase high-performance liquid chromatography (HPLC), affinity chromatography (e.g., antibody affinity chromatography) or combinations thereof). In some embodiments, one or more DNA sequences extracted from the article are isolated and / or purified using electrophoresis (e.g., polyacrylamide gel electrophoresis, two-dimensional electrophoresis, pulsed-field electrophoresis, Southern blotting, or combinations thereof). In other embodiments, one or more DNA sequences extracted from the article are isolated and / or purified using centrifugation. In some embodiments, one or more DNA sequences extracted from the article are isolated and / or purified using a combination of chromatography, electrophoresis, and / or centrifugation.
[0140] In certain embodiments, the one or more DNA sequences extracted from the item are analyzed using mass spectrometry and / or high-throughput DNA sequencing. In certain embodiments, the analyzed DNA sequences are compared to a database of item identification codes, wherein a match of the extracted DNA sequence to an item identification code confirms the identity, authenticity, provenance, and / or security of the item. In certain embodiments, the results of the analysis of the DNA sequence can be compared to prior results of analysis of DNA sequences of the same or similar item.
[0141] In certain embodiments, the analysis of the extracted DNA sequence can yield a "fingerprint," wherein the occurrence of particular DNA sequences, individual nucleotides, and / or cassettes, the relative occurrence of nucleotides and / or cassettes, and / or the particular molecular mass of the DNA sequence and its degradation products can be compared to a database of item identification codes. In further embodiments, the DNA cassette sequence can be analyzed and used to determine an item identification code. In other embodiments, the heterogeneous DNA sequence and / or the nucleotide sequence within the cassette can be analyzed and used to determine an item identification code.
[0142] A key feature of the DNA nucleic acid data packets produced herein is that, although they encode the same digital information (e.g., binary encoded information), they are highly heterogeneous. In other words, a large number of DNA sequences can encode the same data. This vast array space provided by the use of heterogenous (or heterogeneous or diverse) cassettes can be expressed using a Heterogeneity Index (HI): wherein the HI is defined as the ratio of the number of DNA sequences encoding a machine-readable code or data packet to the number of machine-readable codes or data packets. In traditional DNA sequence data encoding methods (e.g., genomic information or binary code-based DNA storage), a single code is represented by a single DNA sequence or DNA sequences highly similar to the single DNA sequence (only considering accidental silent mutations and / or single nucleotide polymorphisms (SNPs) that do not affect the encoded protein amino acid sequence). For natural DNA, considering variations due to silent mutations or DNA replication fidelity errors (natural error rate is about 1 error per 1000 bases), the ratio of the number of data packets (e.g., protein amino acid sequences) to the number of DNA sequences encoding the data packets is 1 or slightly higher. In the present disclosure, a single data packet can be represented by multiple synonymous heterogeneous DNA sequences. For example, as described above, using the heterogenous (or heterogeneous or diverse) cassette data writing approach, the binary data nucleic acid data packet 011011 is added for 6 rounds with 2 unique cassettes in each addition step, then 1 machine-readable code (011011) is represented by 64 unique sequences in the list, resulting in an HI of 64, i.e., 1 data packet corresponds to 64 different synonymous sequences. If the sequence is longer or the number of possible cassettes is larger (e.g., encoding a 100-bit sequence, where each bit can be encoded by any one of 4 different cassettes), the HI will become extremely large, on the order of 4^100. For reference, 4^100 is greater than 10^60, while there are about 10^80 atoms in the universe. Therefore, an HI of 4^100 means that each nucleic acid data packet molecule in a given sample (or writing point) can have a different DNA sequence, although they all encode the same data packet. In contrast, when using homogenous cassette data writing (cassettes roughly 1 : 1 corresponding to bit values), the HI is about 1 regardless of whether the sequence encodes 1 bit or 100 bits.
[0143] One of the results of this extremely high heterogeneity is that a forger can hardly forge the DNA features in the goods marked according to the present disclosure by amplification, analysis, and replication of the DNA features. First, the heterogenous cassette data writing will result in extremely high sequence heterogeneity. For longer sequences, any two DNA molecules are not the same, making it much more difficult for a forger to crack the code than in a system where all DNA chains are the same. Second, even if the forger can read the data packet, although there are obstacles in detecting the code in the noise produced by the highly variable sequences, the forger cannot easily replicate and provide a counterfeit DNA marker with unique features produced by the relative levels (or proportions or mixtures) of different heterogenous (or heterogeneous) cassettes. For example, Figure 20As shown, the proportion of different synonymous cassettes can vary in each round of cassette addition (e.g., 50 / 50, 25 / 75, 75 / 25, etc.), so the varying proportions of cassettes add additional combinatorial complexity to the final mixture. Without prior knowledge of the sequence, it is nearly impossible to detect, much less predict or fake, the unique fingerprint provided by a particular proportion of cassettes. The cassettes use fingerprint can be varied in a number of ways, such as by changing the relative amounts of cassettes used in each cassette addition step in a single batch, or using two or more different large batches (e.g., mixed in different proportions after amplification) to create a “hash” that provides a unique profile for the particular DNA population used to mark each item.
[0144] In one embodiment, the present disclosure provides a novel population of deoxyribonucleic acid sequences that encode data (DNA 1) that can be used for item authentication and anti-counterfeiting, comprising nucleic acid packets (nackets), wherein each nucleic acid packet comprises a plurality of DNA molecules that encode the same data, and the sequences of the DNA molecules have heterogeneity. For example, the present disclosure provides: 1.1. DNA 1 made by heterogenous cassette data writing, wherein two or more cassette sequences are provided for a single bit or combination of bits in a machine-readable code, such that all or nearly all DNA molecules in a nucleic acid packet encode the same data, but the sequence of each molecule exhibits extremely high diversity, wherein the nucleic acid packet comprises a plurality of heterogenous cassettes.
[0145] 1.2. Any preceding DNA, wherein the data is binary code.
[0146] 1.3. Any preceding DNA, made from heterogenous cassettes that encode the same bit or bits of data, wherein the percentage abundance of different cassette variants used in writing the DNA provides a unique distinguishable characteristic of the DNA.
[0147] 1.4. Any preceding DNA, wherein the heterogeneity of the DNA molecule sequences is represented by a Heterogeneity Index, HI (HI equals the ratio of the number of synonymous sequences to the number of packets), greater than 10, such as greater than 100, such as greater than 1000, such as greater than 10000, such as between 10^6 and 10^100.
[0148] 1.5. Any preceding DNA, wherein the data carried by the DNA is a non-fungible token (NFT).
[0149] 1.6. Any preceding DNA, wherein one or more DNA sequences and / or cassettes comprise one or more topoisomerase recognition sequences, such as the topoisomerase recognition sequence is 5’-CCCTT-3’, 5’-TCCTT-3’, 5’-CCCTG-3’, or 5’-TGACT-3’.
[0150] 1.7. Any foregoing DNA, wherein one or more topoisomerase recognition sequences encode data, e.g., 5'-CCCTT-3' encodes "1", and / or 5'-TCCTT-3' encodes "0".
[0151] 1.8. Any foregoing method, wherein the DNA comprises a cassette, each cassette comprising: (i) an information field corresponding to one or more bits in a machine-readable code, and (ii) a topoisomerase recognition sequence, wherein the cassette is 18-25 nucleotides in length.
[0152] 1.9. Any foregoing DNA, wherein the DNA is integrated into or associated with an item of merchandise to enable identification and authentication of the item of merchandise.
[0153] 1.10. Any foregoing DNA, wherein the DNA is adsorbed to, integrated into, or encapsulated in a silica microbead or particle.
[0154] 1.11. Any foregoing DNA, wherein the DNA is adsorbed to, integrated into, or encapsulated in a silica microbead or particle, and is embedded or integrated into an item of merchandise, e.g., for identification and authentication of the item of merchandise.
[0155] In one embodiment, the present disclosure provides a population of deoxyribonucleic acid sequences encoding data (DNA 2) useful for item authentication and anti-counterfeiting, comprising nucleic acid packets (nackets), wherein each nucleic acid packet comprises a plurality of DNA molecules encoding the same data, and the sequences of the DNA molecules are synthesized using one or more transferases, e.g., terminal deoxynucleotidyl transferase (TdT). For example, the present disclosure provides: 2.1. DNA 2, wherein the data is user-defined data, e.g., a user-defined string of data.
[0156] 2.2. Any foregoing DNA, wherein the data is computer-generated, e.g., not manually defined, e.g., computer-randomly generated.
[0157] 2.3. Any foregoing DNA, wherein the data is a ternary code.
[0158] 2.4. Any foregoing DNA, wherein the data is encoded and / or decoded using a scheme, e.g., a user-defined scheme, a computer-generated scheme.
[0159] 2.5. Any foregoing DNA, wherein the one or more transferases comprise terminal deoxynucleotidyl transferase (TdT).
[0160] 2.6. Any foregoing DNA, wherein the DNA sequences comprise a DNA initiation strand or initiation sequence.
[0161] 2.7. Any foregoing DNA, wherein the DNA is integrated into or associated with the item of goods to enable identification and authentication of the item of goods.
[0162] 2.8. Any foregoing DNA, wherein the DNA initiation strand or initiation sequence comprises data useful for authentication of the article, such as a batch number, lot number, production number, data code, customer number, and the like.
[0163] 2.9. Any foregoing DNA, wherein the DNA sequence comprises a series of homopolymer stretches, wherein each homopolymer stretch consists of repeated identical nucleotides, and each homopolymer stretch comprises a different nucleotide relative to any adjacent homopolymer stretch.
[0164] 2.10. The foregoing DNA, wherein the homopolymer stretches comprise one or more repeated identical nucleotides, such as 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, 35 or more, 40 or more, 45 or more, 50 or more nucleotides, and the like.
[0165] 2.11. Any foregoing DNA, wherein the DNA comprises one or more standard nucleotides, such as adenosine, guanosine, thymidine, and cytidine.
[0166] 2.12. Any foregoing DNA, wherein the DNA comprises the standard nucleotides adenosine, guanosine, thymidine, and cytidine.
[0167] 2.13. Any foregoing DNA, wherein the DNA comprises one or more non-natural or non-standard nucleotides.
[0168] 2.14. Any foregoing DNA, wherein the DNA comprises further modifications, such as polyadenylation, conjugation to small molecule and / or polymeric moieties.
[0169] 2.15. Any foregoing DNA, wherein the DNA is single-stranded.
[0170] 2.16. Any foregoing DNA, wherein the DNA is double-stranded.
[0171] 2.17. Any foregoing DNA, wherein the DNA is linear.
[0172] 2.18. Any foregoing DNA, wherein the DNA is circular and / or circularized.
[0173] 2.19. Any foregoing DNA, wherein the data carried by the DNA is a non-fungible token (NFT).
[0174] 2.20. Any foregoing DNA, wherein the DNA is integrated into or associated with an item of merchandise to enable identification and authentication of the item of merchandise.
[0175] 2.21. Any foregoing DNA, wherein the DNA is adsorbed to, integrated into, or encapsulated in silica microbeads or particles.
[0176] 2.22. Any foregoing DNA, wherein the DNA is adsorbed to, integrated into, or encapsulated in silica microbeads or particles, and is embedded or integrated into an item of merchandise, e.g., for identification and authentication of the item of merchandise.
[0177] In one embodiment, the present disclosure provides an ink comprising a population of DNA according to any of DNA 1 and its subsequent sequences and / or DNA 2 and its subsequent sequences (e.g., an aqueous ink, which optionally comprises one or more pigments (e.g., carbon black or other pigments), binders (e.g., polymers, oils, or resins), solvents (water and optionally alcohol or organic solvents), and / or additives (e.g., drying agents or chelating agents)), which comprises a population of DNA according to any of DNA 1 and its subsequent sequences and / or DNA 2 and its subsequent sequences. For example, such inks can be used to authenticate signatures, documents, or printed matter. In certain embodiments, the population of DNA according to any of DNA 1 and its subsequent sequences and / or DNA 2 and its subsequent sequences in the ink encodes a non-fungible token (NFT) associated with a blockchain. Preliminary experiments suggest that the DNA can be well-preserved in inks and paper. DNA stored in dried blood spots collected on FTA cards or even Guthrie filter paper can be accurately analyzed after years of storage without special protective conditions.
[0178] In one embodiment, the present disclosure provides a polymer (e.g., a plastic token or item) comprising a population of DNA according to any of DNA 1 and its subsequent sequences and / or DNA 2 and its subsequent sequences.
[0179] In another aspect, the present disclosure provides a method of DNA synthesis, e.g., synthesis of DNA as described in any of DNA 1 and its subsequent sequences, by topoisomerase-mediated ligation, wherein the DNA comprises a series of boxes corresponding to a series of bits in a machine-readable code (e.g., binary or ternary code), the method comprising adding boxes to a DNA strand, the boxes selected from: a first library of boxes, wherein the boxes each encode a first single-bit or multi-bit data, but are a mixture of at least two different sequences; and a second library of boxes, wherein the boxes each encode a second single-bit or multi-bit data, and each have the same sequence or are a mixture of at least two different sequences; until a desired sequence of bits is obtained, e.g. a population of DNA molecules that provides a nucleotide sequence that is highly heterogeneous but all provide the same sequence of data.
[0180] In another aspect, the present disclosure relates to methods of marking, identifying and authenticating an item of merchandise, e.g. (1) a method of marking an item of merchandise by integrating DNA comprising a nucleic acid data packet as described herein, e.g. any of DNA 1 and its subsequent sequences and / or DNA 2 and its subsequent sequences, into or associated with the item of merchandise to be identified or authenticated, and (2) a method of identifying and optionally authenticating an item of merchandise so marked by extracting and sequencing the nucleic acid data packet, identifying the item of merchandise based on the encrypted data, e.g. binary code data, in the extracted and sequenced nucleic acid data packet, and optionally authenticating the item of merchandise by measuring the relative amounts of different cassette variants in the extracted and sequenced nucleic acid data packet.
[0181] Accordingly, the present disclosure provides a method of item authentication (Method 1) comprising: i. synthesising a DNA sequence comprising nucleic acid data packets (nackets), wherein each nucleic acid data packet comprises a plurality of DNA molecules encoding the same data, wherein the sequences of the DNA molecules are heterogeneous; ii. integrating the DNA sequence into or onto an item; iii. extracting the DNA sequence from the item; and iv. analysing the extracted DNA sequence; v. optionally, aligning the analysed DNA sequence to a database of DNA sequences; vi. optionally, confirming the item’s authenticity.
[0182] For example, in particular embodiments, the present disclosure provides: 1.1. Method 1, wherein the DNA sequence encodes data that provides an item identification code.
[0183] 1.2. Method 1.1, wherein the data that provides the identification code is randomly generated.
[0184] 1.3. Any preceding method, wherein the DNA sequence comprises any of DNA 1 and its subsequent sequences.
[0185] 1.4. Any preceding method, wherein the DNA sequence comprises any of DNA 2 and its subsequent sequences.
[0186] 1.5. Any preceding method, wherein the DNA sequence is synthesised by sequentially adding one or more cassettes, wherein each cassette comprises a plurality of nucleotides.
[0187] 1.6. Method 1.5, wherein the plurality of cassettes are conjugated together using a ligase.
[0188] 1.7. Method 1.5, wherein the plurality of cassettes are conjugated together using a topoisomerase.
[0189] 1.8. Method 1.5, 1.6, or 1.7, wherein the cassettes are heterogenous cassettes having at least two different sequences but both encoding the same data in a machine readable code (e.g. binary or ternary code).
[0190] 1.9. Any preceding method, wherein the DNA sequence is synthesized by sequentially adding DNA cassettes to a DNA acceptor strand, wherein in each sequential addition step, the cassette comprises a population of heterogenous synonymous cassettes such that the cassette has at least two different sequences encoding the same data in a machine readable code (e.g. binary or ternary code).
[0191] 1.10. Any preceding method, wherein the DNA sequence is synthesized using a transferase-based synthesis with data encoding method.
[0192] 1.11. Any preceding method, wherein the conjugation of DNA cassettes comprises adding a population of heterogeneous DNA cassettes, wherein each DNA cassette encodes the same one or more bits or bytes of data within a different DNA oligonucleotide sequence.
[0193] 1.12. Any preceding method, wherein the DNA sequence is integrated into an article by directly surface conjugating one or more DNA sequences to the article.
[0194] 1.13. Any preceding method, wherein the DNA sequence is integrated into an article component or material used to produce the article, optionally into a textile, fabric, leather, biomaterial article, polymer, plastic, wood, metal, ink, paint, solution, suspension, and raw material.
[0195] 1.14. Any preceding method, wherein the DNA sequence is encapsulated in a microcontainer prior to integration into an article, the microcontainer optionally being a microsphere, optionally a silica microsphere.
[0196] 1.15. Any preceding method, wherein the DNA sequence is encapsulated in a molecular assembly, such as a lipid nanoparticle, protein complex or aggregate, or crystal lattice.
[0197] 1.16. Any preceding method, wherein the DNA sequence is inserted into one or more cells, optionally into a larger DNA construct and / or genome, optionally into a yeast, bacterial, fungal, plant, or animal cell, optionally wherein the cell is used to produce a food, drink, biomaterial, or material, such as cheese, beer, wine, vegan leather, pharmaceutical.
[0198] 1.17. Any foregoing method, wherein the integrated DNA sequence is extracted from the item by physical means, optionally cutting, grinding, scoring, dicing, shredding, or pulverizing one or more portions of the item.
[0199] 1.18. Any foregoing method, wherein the integrated DNA sequence is extracted from the item by chemical means, optionally dissolving or lysing the DNA sequence and / or one or more portions of the item.
[0200] 1.19. Any foregoing method, wherein the extracted DNA sequence is isolated and / or purified, optionally by chromatography, ion exchange chromatography, size exclusion chromatography, normal or reverse phase high performance liquid chromatography (HPLC), antibody affinity chromatography, or a combination thereof.
[0201] 1.20. Any foregoing method, wherein the extracted DNA sequence is isolated and / or purified, optionally by electrophoresis, polyacrylamide gel electrophoresis, two-dimensional electrophoresis, pulsed field electrophoresis, Southern blotting, or a combination thereof.
[0202] 1.21. Any foregoing method, wherein the extracted DNA sequence is isolated and / or purified, optionally by centrifugation.
[0203] 1.22. Any foregoing method, wherein the extracted DNA sequence is analyzed using mass spectrometry and / or high-throughput DNA sequencing.
[0204] 1.23. Any foregoing method, wherein the extracted DNA sequence is compared to a database comprising item identification codes that were originally synthesized for the item.
[0205] 1.24. Any foregoing method, wherein the extracted DNA sequence is compared to one or more previous analysis results of extracted DNA sequences of the same or similar item.
[0206] 1.25. Any foregoing method, used in combination with any of Method 2 and its subsequent methods, Method 3 and its subsequent methods, Method 4 and its subsequent methods, Method 5 and its subsequent methods, Method 6 and its subsequent methods, and / or Method 7 and its subsequent methods.
[0207] Figure 20A schematic diagram of topological cassettes (i.e., cassettes suitable for topoisomerase binding and / or topoisomerase-mediated conjugation) representing various binary bit combinations, according to embodiments of the present disclosure, is shown. The chemistry underlying the topological cassettes is particularly suitable for data storage. As shown in the dashed box portion of the figure, the bases are labeled "N" and each topological cassette can vary in length. Not only can the length L of the topological cassettes vary, but so can their composition, e.g., DNA bases or other bases. Regardless of length or composition, each topological cassette can represent a single bit, two bits, four bits, or eight bits, providing broad flexibility for codec development. Any number of bits per cassette can be used. However, the greater the number of bits represented, the fewer the total number of available heterogeneous cassettes that can represent a given bit pattern.
[0208] Figure 20 A schematic diagram of the number of potential topological cassettes based on the number of positions and the number of different DNA bases, according to embodiments of the present disclosure, is shown. Specifically, each position in the cassette can be represented by any of the four (4) different DNA bases (G, C, A, T). The number of potential cassettes is equal to 4ΛN, where N is the number of positions (or base pairs) in the cassette. For example, with respect to the data encoding portion of the cassette (described herein as distinct from the topoisomerase recognition portion and / or overhang portion, although these regions can encode additional information), the number of potential cassettes for a 10 base pair (or position) cassette = 4Λ10, i.e., 1,048,576 unique cassettes. Other exemplary sizes and numbers of potential cassettes are also shown herein, e.g.: 4Λ20 = 1,099,511,627,776 (20 position cassettes); 4Λ19 = 274,877,906,944 (19 position cassettes); 4Λ18 = 68,719,476,736 (18 position cassettes); etc. This flexibility in cassette length and composition provides a nearly unlimited library of cassettes for data writing. For example, if a cassette has ten positions, any single position can be any of the four chemical bases of DNA, as shown in the top portion of the figure. Thus, a 10 position topological cassette can have over 1 million potential unique cassettes. In some embodiments, the size of the topological cassette ranges from 18-20 base pairs (bp). Figure 21 The potential library of cassettes is shown. A 20 bp cassette size can provide approximately 1.1 trillion potential unique cassettes to choose from. The space of potential cassettes can be further exponentially grown using additional synthetic bases (e.g., Q and R), which, when added to the four bases, jumps the number of combinations from 4Λ20 to 6Λ20. Figure 21 The potential library of cassettes is shown. A 20 bp cassette size can provide approximately 1.1 trillion potential unique cassettes to choose from. The space of potential cassettes can be further exponentially grown using additional synthetic bases (e.g., Q and R), which, when added to the four bases, jumps the number of combinations from 4Λ20 to 6Λ20.
[0209] Figure 22An illustration of how multiple different (or unique) boxes can be used to specify the same underlying binary information according to embodiments of the present disclosure. In this regard, topological boxes can be used to create a cryptic molecular tag or code that is resistant to copying or hacking by creating multiple boxes that specify the same 2-bit binary code. Billions of topological boxes representing the same binary information can be constructed. Moreover, instead of any single base change altering the underlying binary code represented, a single base damage will also change the binary code represented. For example, for a 10-box storage string (or nucleic acid data packet), the number of possible molecular structures = 4^10 = 1,048,576, which can represent the same 20-bit binary code, with 4 unique boxes for each 2-bit binary code pair. Figure 22 Also shown is a start strand (or start string) (SS) or acceptor strand of DNA attached at one end to a substrate, and an end cap (EC) DNA strand at the end of the DNA string or nucleic acid data packet, with multiple data boxes (which can be topological boxes) between the SS and EC.
[0210] Writing digital data in synthetic DNA can be viewed as single base synthesis and two-bit encoding per base, where any one of the four DNA bases can represent the two-bit combinations of 00, 01, 10, and 11. Using such an encoding scheme, accidentally replacing any single base with another during synthesis will change the underlying binary code represented. A similar situation occurs when any given single base is damaged. Thus, such a scheme is not suitable for data storage or secure code generation or authentication.
[0211] In one embodiment of the topological boxes described herein, each two-bit group is represented by a multi-base double-stranded DNA box. In this case, based on error checking and error correction, any single base damage does not interfere with accurate reading of the underlying binary code, especially for longer boxes. Moreover, in a system of 20 bp boxes, there are theoretically about 1.1 trillion (4^20) topological boxes, which can be evenly distributed among the four 2-bit combinations shown. If a 4-bit encoding structure is chosen, there will be about 1.1 trillion / (2^4 = 16) 4-bit permutations. Each group of bits can also be represented by topological boxes of different sizes - for example, 00 = 20 bp box, 01 = 18 bp box, 10 = 16 bp box, 11 = 19 bp box. Any other box sizes can be used.
[0212] Figure 23This diagram illustrates a comparison between homogeneous box data writing according to an embodiment of this disclosure and heterogeneous box data writing using multiple topology boxes combined with a predetermined formulation or mixture. Topology boxes can also be used to create encrypted molecular tags or encoded data resistant to replication or attacks by combining multiple unique boxes with different formulations or mixtures, but the underlying binary information remains the same. For example, each two-bit combination can be represented by Y different boxes simultaneously with a specific formulation / mixture. Therefore, sequences can be formulated in different proportions to increase additional combinatorial complexity; for example, for a 2-bit binary encoding scheme, assuming the percentage of each potential sequence in the formulation or mixture is an integer, the number of possible formulations is (100^Y)^4.
[0213] Therefore, another feasible step in encrypting basic binary information is to write, instead of using any single topology box, any set of binary information (e.g., 00, 01, 10, 11) during production operation, multiple boxes mixed in a fixed proportion for each set of binary bits. Figure 22 An example is shown. In the example shown, each group of two bits is represented by four different topology boxes, which are used in a fixed ratio of mixture / recipe. The ability to combine boxes in mixture or recipe configuration further expands the permutation space for cryptographic writing of binary sequences and enhances authentication capabilities (detailed below). Furthermore, the mixture or recipe can vary with each production run. In some implementations, the batch number can be directly associated with the mixing ratio used in that batch.
[0214] Figure 23 For the purposes of this disclosure implementation scheme, such as Figure 23 The diagram illustrates the loading of a heterogeneous mixture / formulation of an exemplary topology box into printheads 830, 832, 834, and 836 of a laser DNA printer. Specifically, Figure 30A A side view of a silicon wafer 10 is shown, the silicon wafer 10 having a patterned (or unpatterned) SiO2 layer 202 on top to form site pillars (or sites) 14 with an attached top coating 204 (e.g., HfO2), and fluid channels 15 between the sites 14. Figure 24A side view of printhead set 822 is also shown, having four nozzles 830A, 832A, 834A, 836A, respectively for the binary code 2-bit pairs (00, 01, 10, 11) used for 2-bit binary encoding, for adding the cassette associated therewith, and a fifth nozzle 814 for writing the deblock / adaptor; also shown in the figure is the possibility of using a cleaning fluid 820 for a cleaning cycle: the cleaning fluid can be spread, flowed, applied or sprayed horizontally on the surface of the silicon wafer, or operated vertically on the individual printheads 816 and corresponding nozzles 816A that are part of the printhead set 822, similarly to what described above in the co-owned patent application for printing DNA with inkjet. The printhead set 822 can be controlled by a printhead controller (discussed below in connection with Figure 24 The printhead or printhead set 822 has four chambers 830, 832, 834, 836, respectively with associated nozzles 830A, 832A, 834A, 836A, with reagents for adding a code by droplets to the starting DNA strand (or starting strand or starting string or SS) 210 in the vesicle 802 shown on top of each point 14, for example adding "00" head 830, adding "01" head 832, adding "10" head 834, adding "11" head 836, and deblock / adaptor head 814. The adding of the 00, 01, 10, 11 reagents can add the "cassettes" described herein, which contain a plurality of double-stranded DNA bases as discussed herein, and the adding reaction chemistry is the same as described above and in the co-owned US patents and patent applications.
[0215] More specifically, each chamber has a predetermined mixture 830B, 832B, 834B, 836B of a plurality of cassettes CI - CI 6 associated with each 2-bit pair, for example CI - C4 for the "00" bits, C5 - C8 for the "01" bits, C9 - CI 2 for the "10" bits, CI 2 - CI 8 for the "11" bits, and each mixture is loaded into the corresponding chamber 830, 832, 834, 836 of the printhead set 822, respectively, before the writing process starts.
[0216] In particular, in some embodiments, the addition chemistry used to write to the polymer can be the chemistry described herein and in the above-mentioned commonly owned U.S. patents, which includes a "deblocking" step. Further, in some embodiments, the addition chemistry used to write to the polymer can be the chemistry described in the above-mentioned commonly owned pending U.S. patent applications, in which a "adapter" is used in place of a deblocking enzyme. Thus, the operation of preparing the DNA strand for another addition reaction can be referred to herein as a "deblock / adapter" or "adapter BA" operation.
[0217] In some embodiments, a washing solution is flowed through the array after an addition reaction to prepare the DNA for the next addition reaction or deblocking reaction. In some embodiments, as an alternative or in addition to the illustrated side-flow wash, the printhead can have additional chambers or nozzles (shown in dashed lines) in which a washing solution is housed that is dispensed during a washing cycle. Further, in some embodiments, a deblock / adapter printhead can be applied or flowed over the wafer at an appropriate time during the writing.
[0218] Figure 31A A schematic of a process for writing a two-bit binary code on a surface of a substrate or matrix according to embodiments of the present disclosure. In particular, Figure 6A cross-sectional side view of a silicon wafer 10 (as an example substrate) with patterned layers 202, 204 showing an initial polymer or DNA strand (SS) 210 in liquid 802 attached to a site post (or site) 14 and showing a side view of a data write (or print) process 930 for adding bits or codes to the free end of the initial polymer DNA strand on the wafer, similar to that described above in the co-owned patent application for inkjet printing of DNA. In particular, the write adds start with the execution of a wash cycle 820 to prepare the DNA strand 210 for the first write add reaction. Next, the printhead dispenses droplets of "00", "01", "10" or "11" containing the cartridge or cartridge mix / formulation associated with the 2-bit code being written to the target site, as shown in blocks 912A, 912B, 912C. After the add reaction is complete, a wash cycle 802 is executed to prepare the DNA strand for the de-blocking / adaptor reaction. Next, the printhead dispenses de-blocking / adaptor droplets to the target site that just completed the add reaction, as shown in blocks 904A, 904B, 904C. After the de-blocking / adaptor reaction is complete, a wash cycle 802 is executed to prepare the DNA strand for the next add reaction. Next, the printhead dispenses droplets of "00", "01", "10" or "11" according to the target cartridge or cartridge mix / formulation associated with the 2-bit code being written to the target site, as shown in blocks 916A, 916B, 916C. After the add reaction is complete, a wash cycle 820 is executed to prepare the DNA strand for the de-blocking / adaptor reaction. Next, the printhead dispenses de-blocking / adaptor droplets to the target site that just completed the add reaction, as shown in blocks 908A, 908B, 908C. The above process is repeated until all of the target cartridges or 2-bit codes have been written to the DNA strand. The following discussion of the write add process is further discussed in connection with Figure 7-13 and 31B Further discussion of the write add process.
[0219] In some embodiments, when using the AB / BA write method, if the cartridges are designed with both AB and BA for each binary code, it can be possible to eliminate the AB adaptor, thus replacing the de-blocking / adaptor steps 904A, 908A (as shown in the code write example of Figure 31A and Figure 53A In this case, in the write (print) logic 3100 of Figure 25A and 5300 of Figure 25B Block 3120 is not executed.
[0220] Figures 25C-25J , 25BFigs. 25C, 25D, 25E, 25F, 25G, 25H, 25I, and 25J are schematic illustrations of the process of writing a storage string using a pre-determined cartridge formulation or cartridge mixture for each 2-bit pair on one point on a substrate according to embodiments of the disclosure. Specifically, it shows each writing cycle and how the cartridges are added to the storage string. For each write, the cartridges in the droplet will attach randomly to the loose chain. In this example, 10 independent DNA chains or storage strings are synthesized, each representing the same 20-bit binary code (11010010011111001001) shown.
[0221] If the first pair of two bits is 11 as shown, using the current formulation, 50% of the molecules on the surface will get the cartridge labeled C13 and 50% will get the cartridge labeled C16. Which of the 10 molecules on the surface will get which of the two "11" cartridges is completely random. This non-algorithmic randomness will also ensure that data written using such formulation (or encryption) method will likely be resistant to quantum copying or attacks, since all algorithmic random number generators have a slight bias that quantum computers can break.
[0222] The process continues, with each of the 10 DNA chains synthesized getting a random cartridge based on the formulation used to represent the relevant two-bit pair, in this case, 01. However, the molecules of all the chains on the surface represent, statistically, the specific mixture / formulation of cartridges as shown by cartridge labels C1-C16, each cartridge being a unique cartridge. Figure 25J A similar process occurs. Figure 29B A similar process occurs.
[0223] Referring to Figure 26 At the end of the synthesis run in this example, the molecules produced are indeed random, but they all represent the same underlying 20-bit binary information (11010010011111001001) shown above. In this example, 10-cartridge strings were constructed. If four different topological cartridges are used for each pair of two-bit binary codes, then over 1 million (4A0) different molecules are possible, as described above. The possible permutations of data chains or storage strings are: (number of cartridges per bit pair)A(number of cartridges in the chain).
[0224] If 128-cartridge storage strings are constructed, and four different cartridges are used for each pair of two-bit binary codes, then over 1.16e77 (4A128) different molecules are possible. If each pair of two-bit binary codes is represented by, for example, 10 cartridges, then the permutation space will increase dramatically (see below in connection with Figure 27A and 29C(Further discussion follows). The resulting permutation space is so vast and random that, from an economic perspective, it is not feasible to synthesize a series of molecules to forge NFT tokens, smart contracts, or other secure data files encoded in this way.
[0225] Figure 27A This is a schematic diagram of a box along a storage string and a box assigned to each 2-bit code in the storage string according to an embodiment of this disclosure.
[0226] Figure 27B , 27B Figures 27C and 27D are schematic diagrams illustrating a process for verifying a storage string or nucleic acid data packet using a predetermined box mixture associated with a given 2-bit binary code, according to an embodiment of this disclosure. Specifically, in Figure 28 In this process, all boxes associated with the 11-bit code are collected and analyzed, and their overall distribution should roughly match the mixtures or formulations of the relevant batch numbers. Similarly, in Figure 34 , 27C In 27D, boxes associated with 01, 00, and 10 bit codes are collected and analyzed separately. The overall distribution of each bit code should roughly match the mixture or recipe of the relevant batch number. Thus, using the box mixture or recipe assigned to a certain bit code can provide another dimension of randomness and authenticity.
[0227] Specifically, Figure 29A This illustration demonstrates two dimensions of randomness and validation for a storage string (or nucleic acid data packet) according to an embodiment of this disclosure. The first dimension runs along the storage string, where validation can be based on multiple box assignments for each 2-bit code under a given batch number. The second dimension spans all bits of the storage string across the surface, where validation is based on a mixture or recipe associated with a 2-bit code under a given batch number. This also... Figure 29A The decoding and mixture verification logic flowchart is shown above. Alternatively, the two randomness dimensions mentioned above can be described as production fingerprints and molecular fingerprints. In this case, the production fingerprint contains the basic information encoded within the nucleic acid data packet, where the mutation potential provides a high-entropy mutation space and optional randomness. The molecular fingerprint contains the physical molecular structure of the nucleic acid data packet (e.g., the DNA nucleotide sequence), which provides an independent and orthogonal (relative to the production fingerprint) high-entropy mutation space and randomness, where each molecule encoding information may be unique.
[0228] Figure 22 , 29B 29C is a binary code-box mapping table showing various allocation relationships between binary codes and boxes, and related box mixtures / formulations, based on batch numbers, according to embodiments of this disclosure. Furthermore, as described herein, the information in the binary code-box mapping table may be stored on and retrieved from a blockchain. Specifically, Figure 23A binary code-box correspondence table for a 2-bit coding scheme is shown ordered by batch number, where each 2-bit binary code is assigned 4 boxes, and each 2-bit code has a predetermined mixture / formulation. In this example, batch number 1 shows that boxes C1-C16 are assigned to the 2-bit codes as described herein in connection with Figure 22 and Figure 23 shown by the example. Where C1-C4 are assigned to the "00" bit code, C5-C8 are assigned to the "01" bit code, C9-C12 are assigned to the "10" bit code, and C13-C16 are assigned to the "11" bit code. Also, the mixture percentages are the same as described herein in connection with Figure 29A and Figure 29A the example. For batch number 2, the assignment of boxes (C's) is shifted down or rolled over by 1 position as a whole. At this point, C16 is at the top, followed by C1-C15, and the resulting 4 boxes for each 2-bit code are assigned accordingly, as shown in batch number 2 in Figure 29A . Furthermore, the mixture percentages for each group of 4 boxes are randomly shuffled with respect to that shown for batch number 1. For batch number 3, the assignment of boxes (C's) and mixture percentages (%) are randomly selected, independent of the assignment of the previous batch. For batch number 4, the assignment of boxes (C's) is rolled over or shifted with respect to that shown for batch number 1, in groups of 4 boxes. At this point, C13-C16 are at the top, followed by C1-C4, C5-C8, and C9-C12, and the resulting boxes for each 2-bit code are assigned accordingly, as shown in batch number 4 in Figure 19 . Furthermore, the mixture percentages for each group of 4 boxes are shuffled with respect to that shown for batch number 3. The above examples of different assignments and mixture percentages (or ratios) are merely illustrative, and any other values and variations can be employed.
[0229] Referring to Figure 29BIn some implementations, the binary code-box mapping table may also include a write direction for writing digital codes for a given batch, such as MSB-LSB, LSB-MSB, or random. Specifically, a memory string or data packet to be written to a given location on the chip from the substrate or wafer array surface can be written in two possible directions: MSB-LSB, i.e., from the most significant bit (MSB) to the least significant bit (LSB), i.e., from left to right; or LSB-MSB, i.e., from the least significant bit (LSB) to the most significant bit (MSB), i.e., from right to left. The number of bits in the LSB or MSB depends on the encoding type used, as detailed below. Furthermore, setting the write direction for a given batch to "random" means that the write logic can determine which write direction to use for any given location within the given batch. In this scenario, the write direction can be indicated by a flag or code (e.g., MSB-LSB / LSB-MSB flag or code, where 1 = MSB-LSB, 0 = LSB-MSB), which can be written into the end cap (EC) of all memory strings or nucleic acid data packets written at a given site. Therefore, a given batch can have randomly distributed write directions on the same chip or array. As a result, multiple sites written with the same code for redundancy and error detection / correction can have their codes randomly selected to be written into memory strings or data packets in one of two different directions. This adds another layer of randomness to the final encoded DNA / polymer memory strings or nucleic acid data packets.
[0230] For example, when the code 10101100 is written in MSB-LSB mode (using single-bit encoding), LSB(0) is closest to the end cap. However, when the same code 10101100 is written in LSB-MSB mode, MSB(1) is closest to the end cap. In the case of 2-bit binary encoding, both LSB and MSB contain two binary bits. Therefore, for the code 10101100, when written in MSB-LSB mode, LSB(00) is closest to the end cap; when the same code is written in LSB-MSB mode, MSB(10) is closest to the end cap.
[0231] Furthermore, although some examples in this paper show 1-bit binary code and some show 2-bit binary code, it should be understood that binary data can be encoded into boxes using any number of bits, as in this paper (e.g., Figure 22The encoding scheme can change for a given batch number in some embodiments, and is saved in a binary code-to-cassette correspondence table. For example, batch number 1 can employ a 2-bit encoding, batch number 2 can employ a 3-bit encoding, batch number 3 can employ a 4-bit encoding, and so on for other batch numbers. In this case, a 3-bit encoding has 8 different binary codes 000-111, each of which is assigned one or more cassettes. If each code is assigned 4 cassettes, then this configuration uses 32 unique cassettes. Similarly, if a 4-bit encoding is employed, then there are 64 codes 0000-1111, each of which is assigned one or more cassettes. If each code is assigned 4 cassettes, then this configuration uses 64 unique cassettes. In some embodiments, the data to be written can be padded with a predetermined number of additional bits, such that the total number of bits is divisible by the number of bits of the encoding scheme.
[0232] Figure 23 A binary code-to-cassette correspondence table is shown, ordered by batch number, for a 2-bit encoding scheme, in which each 2-bit binary code is assigned a variable number of cassettes based on batch number, and each 2-bit code has a predetermined mixture / ratio. In this case, batch number 1 shows that the 2-bit codes are assigned 4 cassettes, as described herein in connection with Figure 29C and Figure 29C The example shown. Batch number 2 shows that each 2-bit code is assigned 5 unique cassettes, for a total of 20 cassettes (C1-C20). Batch number 3 shows that each 2-bit code is assigned 6 unique cassettes, for a total of 24 cassettes (C1-C24). Batch number 4 shows that each 2-bit code is assigned 7 unique cassettes, for a total of 28 cassettes (C1-C28). Batch number N shows that each 2-bit code is assigned 10 unique cassettes, for a total of 40 cassettes (C1-C40). As the number of cassettes in the mixture increases, the proportion decreases, which can be balanced against the percentage threshold or tolerance of the detection system to optimize verification accuracy. X indicates not applicable.
[0233] Figure 30A A binary code-to-cassette correspondence table is shown, ordered by batch number, for a 2-bit encoding scheme, in which each 2-bit binary code is selected from a total of 40 cassettes (10 candidate cassettes per 2-bit code) based on batch number, and each 2-bit code has a predetermined mixture / ratio. In this case, batch number 1 shows that 4 unique cassettes are selected from the available 10 cassettes assigned to the 2-bit codes. Batch number 2 shows that a different 4 cassettes are selected from the 10 cassettes assigned to each 2-bit code. Batch 3 shows that a different 4 cassettes are selected from the 10 cassettes assigned to each 2-bit code. Batch 4 shows that a different 4 cassettes are selected from the 10 cassettes assigned to each 2-bit code. Batch N shows that a different 4 cassettes are selected from the 10 cassettes assigned to each 2-bit code. X indicates cassettes not used for a given batch number. Figures 29A-29CAn advantage of the illustrated scheme is that for any given batch number, only 4 boxes need to be combined in the mixture, thereby increasing the percentage ratio while maintaining randomness by providing 10 available boxes for any given 2-bit code. In some embodiments, each 2-bit code can use all 40 boxes when selecting boxes for a given 2-bit code. In addition, the same approach can be taken for any number of boxes for a given 2-bit code, e.g., 5 out of 40 boxes are selected. If desired, the total number of boxes can also be increased to increase randomness.
[0234] Figure 30B To illustrate a block diagram of an inkjet printing system 1900, which includes an inkjet printing instrument 1902 and a computer system 1904 interfaced with the instrument 1902, similar to that described in the aforementioned U.S. patent application for DNA inkjet printing. The inkjet printing instrument 1902 can include a piezoelectric inkjet printhead 1906 that delivers droplets of reagents described herein to a target write site on a wafer array 10 mounted on an XY stage 1907. The printhead and XY stage can be controlled by a printhead and array stage controller and detection logic 1908, which communicates with local control logic 1910 to write target reagents and codes to DNA strands in the manner described herein. For example, one or more of read / write address and / or data input, output, and / or control lines 1912 can receive or provide to a serial bus containing instructions for writing what code or data to the array. The computer system 1904 can receive instructions from a user 1903 and provide information to a display 1905 for use by the user 1903, and can also provide instructions to the local control logic 1919, which provides specific write requests to the printhead / printhead group 1906 as well as the array stage controller and detection logic 1908. The printhead and array stage controller and detection logic 1908 control XYZ position of the printhead and wafer array XY stage 1907, and also receive data from a droplet observer (or sensor) 1911 to determine quality control of the droplets, and provide results and error feedback to the local control logic 1910 and computer system 1904, which stores droplet error information in a DNA data server 1915 or other storage device for use in subsequent reading of data. Such information can be used to correct or ignore certain data that is known to have specific data errors due to droplet errors.
[0235] The inkjet printing instrument 1902 may include instrument (fluid / reagent) control logic 1914, which controls the reagent supply 1916 to the printhead 1906 and the fluid flow 1920 through the inflow manifold 1921 and through the wafer array 10, for example, controlling cleaning fluid 1922, lysis buffer 1924, preparation fluid 1926, etc., via valves 1920A, 1920B, 1920C and control line 1919, respectively; and controls the outflow fluid 1930 through the outflow manifold 1931, for example, controlling waste liquid 1932 via valve 1930A and control line 1933, and controlling fluid 1934 containing coding DNA dissociated from the wafer array 10 and collected (e.g., collected in collection tank 1936) for subsequent reading via valve 1930B and control line 1933. The reagent / supply loading assembly may be controlled by the instrument 1902 and may include necessary known valves and fluid systems, in accordance with the present invention. Figure 30A The data in the binary code-box mapping table is for the distribution of target boxes and mixtures / formulas associated with the binary code of a given batch number for the printhead / group loading. This data may be stored in the DNA data server 1915 or other storage devices and provided to the instrument 1902 by the computer system or local control logic 1910, or may be obtained directly from the server.
[0236] In some implementations, the printhead and array platform controller 1908 can be configured to replace (remove / load) a printhead or a set of printheads (a group of printheads) between DNA writing operations for each production batch. In this case, the printhead and array platform controller 1908 can remove an existing printhead / set 1906 and acquire a corresponding printhead / set 3004 with the target mixture / formulation (C1-Cm), and load the printhead / set into an inkjet printer to write DNA / polymer into the wafer array 10. This operation can be performed by a robotic arm 3002 or other controllable device or system, which can be part of or separate from the printhead and array platform controller 1902.
[0237] Figure 30B For the purposes of this public implementation scheme, Figure 30A Block diagram of Computer System 1904. Computer System ( Figure 30B Device 1904 can interact with inkjet printing instrument 1902, and also with instrument controller 1914, which in turn interacts with independent fluid supply 1916, etc., all of which interact with one or more CPUs / processors 1952 or logic to perform the functions described herein. Furthermore, Figure 30A and Figure 31A The computer system in the middle can be connected to the user 1903 and the display screen 1905 ( Figure 30A (Connection)
[0238] The local control logic 1910, the fluid instrument controller 1914, and the printhead and array platform controller 1908 possess the necessary electronics, computer processing power, interfaces, memory, hardware, software, firmware, logic / state machine, database, microprocessor, communication link, display or other visual or audio user interface, printing device, and any other input / output interface, including sufficient fluid and / or pneumatic control, supply, and measurement capabilities to achieve the functions described herein or to obtain the intended results.
[0239] Figure 23 The flowchart 3100, according to an embodiment of this disclosure, is for writing (printing) and unloading an coded polymer storage string in an inkjet writing system. This logic 3100 can be derived from... Figure 29A The system executes this process. In some implementations, the above writing process can be repeated for each new set of DNA strands to be written. Specifically, logic 3100 begins at box 3102: loading or printing the starting (or recipient) DNA strand (or starter strand or SS) onto the wafer array site 14. Figure 31B Next, block 3104 receives the batch number and binary code to print / write the first storage string or nucleic acid data packet. Next, logic block 3106 retrieves from the DNA data server four inkjet cartridges or printheads corresponding to each 2-bit code (00, 01, 10, 11) for that batch number, each 2-bit code assigning a different DNA cartridge mixture, for example from the corresponding binary code-cartridge correspondence tables 2900, 2920, 2904. Figure 31B , 29B (29C). Next, logic block 3108 retrieves the write direction for a given batch number from the binary code-box mapping table stored in the DNA data server; if random, it randomly selects the write direction for one or more sites to be written and saves it in the binary code-box mapping table. Next, logic block 3110 retrieves the first 2-bit binary code to be written to the substrate or wafer from the target binary code based on the write direction obtained from the binary code-box mapping table. Next, logic block 3112 performs a cleaning cycle on the wafer array to remove any excess reagent from the wafer surface. Next, logic block 3114 follows the instructions in conjunction with this document. Figure 23 The writing process involves using a corresponding box to write / print the 2-bit code to the storage string / nucleic acid data packet at the target site. After the 2-bit code is written, logic block 3116 determines whether there are more sites that need to be written before applying de-blocking / adaptor to that site. If so, the logic returns to block 3114 and proceeds according to the provisions of this document. Figure 31AThe write process writes / prints the 2-bit code to the corresponding bin at the target site until all target sites corresponding to the 2-bit code are written. Then, when the result of block 3116 is no, block 3118 waits for the addition reaction to complete. After the reaction is complete, block 3120 prints the de-blocking / adaptor to the target site. In some embodiments, the de-blocking / adaptor can be flowed over the array surface rather than applied using an inkjet cartridge or printhead (as shown, for example, in Figure 30A Next, block 3122 determines whether all 2-bit codes for the current storage string or nucleic acid data packet have been written. If not, block 3124 obtains the next 2-bit code in the storage string and returns to block 3112 to perform the cleaning cycle, repeating the process for the next 2-bit code (as shown, for example, in Figure 33 Next, block 3122 determines whether all 2-bit codes for the current storage string or nucleic acid data packet have been written. If not, block 3124 obtains the next 2-bit code in the storage string and returns to block 3112 to perform the cleaning cycle, repeating the process for the next 2-bit code (as shown, for example, in Figure 31B Next, block 3122 determines whether all 2-bit codes for the current storage string or nucleic acid data packet have been written. If not, block 3124 obtains the next 2-bit code in the storage string and returns to block 3112 to perform the cleaning cycle, repeating the process for the next 2-bit code (as shown, for example, in Figure 30A Next, block 3122 determines whether all 2-bit codes for the current storage string or nucleic acid data packet have been written. If not, block 3124 obtains the next 2-bit code in the storage string and returns to block 3112 to perform the cleaning cycle, repeating the process for the next 2-bit code (as shown, for example, in
[0240] Figure 30A For a flowchart for writing 2-bit codes to DNA / polymer storage strings in an inkjet writing system according to embodiments of the present disclosure, the logic can be as shown in Figure 30AThe system executes the following logic. Specifically, this logic checks each 2-bit code to be written and causes the corresponding inkjet cartridge or printhead, equipped with the corresponding DNA box (or topology box) mixture, to print the corresponding 2-bit code at the target site / position on the wafer array or chip. Once the corresponding 2-bit code has been written at the target number of sites, the logic determines whether a droplet observer, which may be part of the printhead and array platform controller and detection logic, has detected any droplet errors. If any error is detected, the logic saves the error location and bit number for subsequent reading, and the logic ends. Specifically, logic 3150 begins at block 3152, which determines whether the 2-bit code to be written is "00". If yes, block 3154 uses a "00" cartridge equipped with a "00" DNA box (or topology box) mixture to print the "00" bit at the target site / position on the wafer array or chip. Next, or if the result of block 3152 is no, block 3156 determines whether the 2-bit code to be written is "01". If yes, then box 3158 uses a "01" cartridge equipped with a "01" DNA box mixture to print a "01" bit at the target site / position on the wafer array or chip. Next, or if the result of box 3156 is no, then box 3160 determines whether the 2-bit code to be written is "10". If yes, then box 3162 uses a "10" cartridge equipped with a "10" DNA box mixture to print a "10" bit at the target site / position on the wafer array or chip. Next, or if the result of box 3160 is no, then box 3164 determines whether the 2-bit code to be written is "11". If yes, then box 3166 uses an "11" cartridge equipped with an "11" DNA box mixture to print an "11" bit at the target site / position on the wafer array or chip. Next, or if the result of box 3164 is no, then box 3168 determines whether the bit writing of a single site or group of sites on the chip is complete. If completed, box 3170 determines the droplet observer (or sensor) 1911 ( Figure 32A Whether any droplet errors are detected, the droplet observer may belong to the printhead and array platform controller and detection logic 1908 ( Figure 32A If the result of box 3170 is negative, an error is detected, box 3172 saves the error location and bit number for subsequent reading, and the logic ends. If the result of box 3170 is negative, no droplet error was found in this write cycle, and the logic ends.
[0241] Figure 33 This is a side view showing a plurality of sites 14 (using the code-writing method described herein) encoding DNA strands 1002, 1004, 1006 according to an embodiment of the present disclosure, and a lysis buffer 1008 for removing the encoding DNA strands from the substrate surface. Specifically, Figure 32BFIG. 10 is a side view of a silicon wafer 10 (as an example substrate or wafer) with patterned layers 202, 204 showing one end of a starting (or acceptor) DNA strand 210 attached to a site post (or site) 14, the other end attached to an encoded DNA, and the figure also shows how the encoded DNA strands 1002, 1004, 1006 are removed from the wafer 10 using a cleaving solution 1008. In some embodiments, the wafer 10 can be non-patterned or partially patterned as described in the aforementioned co-owned patent applications with respect to DNA inkjet printing. Each site post or site 14 has a plurality of encoded polymers or DNA strands (or nucleic acid data packets). When all bits or bins or codes have been written or printed, a cleaving solution 1008 can be flowed through the wafer array (or chip) releasing the encoded DNA 1002, 1004, 1006 so that it is removed from the solid substrate 204 or flowed (as shown by arrow 1010) and placed in a storage container Figure 32B
[0242] FIG. 11 is a schematic diagram of a site array with encoded DNA according to embodiments of the present disclosure, where the columns (X) are redundant sites that write the same encoded DNA data, and the rows (Y) are sites that write different encoded DNA. In some embodiments, each site on the substrate or wafer surface can write unique encoded data, which can include an address or ID associated with the storage string or nucleic acid data packet or strand written to that site, such as a storage string (or nucleic acid data packet) address or ID, such as NID1, NID2, NID3, NID4, through NIDY. In some embodiments, the same unique encoded data can be written to multiple sites on the substrate or wafer surface to provide redundancy and enhance error detection and verification capabilities. In this case, the redundancy and verification described herein can be performed for all storage strings (or nucleic acid data packets) with the same address or ID, regardless of which sites or how many sites they originated from. This increases the number of strings that participate in the longitudinal and lateral redundancy and verification described herein.
[0243] In particular, This illustrates multiple sites with the same storage string or nucleic acid data packet ID or address. For example, the first row contains multiple sites (shown as X sites) with the same nucleic acid data packet ID, NID1; the second row contains multiple sites (X sites) with the same nucleic acid data packet ID, NID2; and subsequent rows have similar redundancy settings, where the same data is written to multiple sites across the chip, thus providing redundancy and tamper-proof and error detection capabilities. In this case, all storage strings or nucleic acid data packets with the same nucleic acid data packet ID (or storage string address) can be analyzed as a group to perform verification testing. Although the same code can be written to multiple sites, the actual DNA / polymer sequence of the bases or base sets (cassettes) will differ between sites and even within the same site due to the cassette mixture / formulation and writing direction described herein.
[0244] FIG. 33 The schematic diagram, according to an embodiment of this disclosure, illustrates the process of transferring site-specific DNA storage strings or nucleic acid data packets from a substrate surface to a collection tank and reading and decoding the collected DNA. Specifically, FIG. 11 The schematic diagram illustrates an example of multiple sites 1142-1148 encoding DNA 1002-1008 (after coding) attached to a flat surface 1101 of a wafer (or other substrate), and the process of removing, storing, and reading the written data at each site (site 1-site N). See details. FIG. 33 According to some embodiments of the present invention, the figures illustrate an example of multiple sites having encoded DNA (after being written) attached to a wafer, and illustrate the process of removing, storing, and retrieving the data written at each site. After the target code is written to the DNA storage strings (or strands or nucleic acid data packets) 1002-1008 corresponding to each site 1142-1148 with the encoded DNA storage strings 1002-1008 attached, it can be unloaded as described herein and in the aforementioned patent applications, and the encoded DNA storage strings can be detached or removed from their corresponding sites. In some embodiments, multiple encoded DNA storage strings can be attached to a single given site (as described above). Subsequently, the detached encoded DNA storage strings are transferred along the output channel via fluid transport (as indicated by arrow 1110) to a collection tank or container 1112, which receives the encoded DNA strings from all sites in a given wafer array outside the wafer (or separately from the wafer). When the stored data needs to be read, any known commercial DNA reader / sequencer 1114 (such as Illumina, Oxford Nanopore or other manufacturers' DNA sequencers) can be used to read the encoded DNA storage strings in collection slot 1112. The sequencer's accuracy is sufficient to meet the target application requirements and can determine the DNA sequence written on each DNA storage string.
[0245] DNA reader / sequencing instrument 1114 can provide code data values from the stored string to a computer-based system 1126, which performs decoding and mixture verification logic 1127 (which can be combined below). FIG. 34 The flowchart 3400 is implemented, and this logic analyzes and decodes data from the DNA sequencer, based on the cassette mixture / formulation for a given batch number (according to the binary code-cassette mapping table described herein). FIGS. 29A-29C The authenticity of this document can be verified. A computer system can be used to verify the authenticity of this document. FIG. 30B The system or a similar system. Computer system 1126 can be connected to DNA data server 1124 (which can be connected to...). FIG. 30A The DNA data server 1124 communicates with a DNA data server 1915 (which is the same as or similar to the DNA data server 1124), and the DNA data server 1124 may store a binary code-box mapping table sorted by batch number for use by the decoding and mixture verification logic 1127. In some embodiments, the DNA sequencer may save the code data directly to the DNA data server 1124 for retrieval by the decoding and mixture verification logic 1127. In some embodiments, the computer system 1126 may communicate with a display 1125, which may display or report to the user the data results obtained by reading the DNA-encoded data storage string 1100.
[0246] In some implementations, data can be written into the DNA string in an address / data format, similar to... FIG. 35B As shown, the address or number of the site to be written is encoded, followed by data associated with that address (or site number). Other formats may also be used if necessary. Each site contains multiple DNA start (or acceptor) strings (as described herein), and they can all be written simultaneously. The number of DNA strings or strands at each site depends on the droplet site size, and each site can contain thousands to billions of DNA strings or strands, or other numbers of DNA strings if needed. Furthermore, in some embodiments, for applications where site addresses are not critical (e.g., where the encoding DNA is retained on the array), the site address need not be part of the code.
[0247] According to the implementation scheme of this disclosure, with reference to FIG. 34 It shows the logic 1127 used to implement decoding and mixture confirmation. FIG. 33 Flowchart 3400 shows the logic used to decode and acknowledge polymer storage string data. Logic 3400 begins at block 3402, which obtains the batch number of the wafer array or chip, which may be printed on the wafer or otherwise associated with wafer 10. FIG. 23It also retrieves DNA base data from readings of all storage strings on the chip by the DNA sequencer, for example, from DNA data server 1124. This logic also retrieves a binary code-box mapping table from DNA data server 1124. Next, logic block 3404 separates storage strings or nucleic acid data packets by address or ID and identifies boxes on each string using topological spacing (as described above). Then, logic blocks 3406, 3408, 3410, 3412, 3414, 3416, 3418, and 3420 identify boxes in a given string and, according to the binary code-box mapping table (e.g., ... FIG. 29A , 29B (As shown in 29C) The boxes assigned to the bit codes are analyzed. If a match is found, the count for that bit code is incremented, as shown in boxes 3408, 3412, 3416, and 3418. This process is repeated through boxes 3422 and 3424 until all boxes for a given storage string or nucleic acid data packet have been checked. Then, logic box 3428 arranges the 2-bit codes according to the write direction determined by reading the binary code-box mapping table or by the endcap flag of the storage string or nucleic acid data packet. Next, logic box 3430 determines whether all storage strings with the current address or data packet ID have been processed. If not, box 3432 retrieves the next storage string / nucleic acid data packet and repeats the process from box 3406 for all storage strings with the same address or ID until all are processed. After completion, logic box 3436 determines whether the count for each 2-bit code matches the expected distribution (or proportion) of boxes for that code based on the batch number of the given storage string or nucleic acid data packet address or ID. If a match is found, logic block 3440 sets the confirmation flag to "Pass," confirming that the data for the given storage string or nucleic acid packet address or ID is authentic. If a mismatch is found, logic block 3438 sets the confirmation flag to "Fail," or marks it as failed, indicating that the data is incorrect or forged. Next, logic block 3442 checks whether all storage string / nucleic acid packet addresses or IDs have been decoded and verified. If not, logic block 3444 retrieves the next storage string or nucleic acid packet address / ID and repeats the process starting from block 3406 for the next address until all addresses are decoded and verified and the result of block 3442 is "Yes." Then the logic ends.
[0248] FIG. 35A and FIG. 35B The diagram illustrates an embodiment of this disclosure, showing an example of the address, data, and error detection boxes that constitute a DNA / polymer storage string. Specifically, refer to... FIG. 35A and FIG. 35BThe data format written to the storage string can vary based on a variety of factors and design standards. Specifically, the "storage string" (or storage chain, DNA, polymer, data packet, chain) 1802 can be represented as a line with a series of ellipses 1804, which represent individual boxes written (or added) to the storage string in a given storage cell, wherein, as described herein, a box represents or signifies one or more binary (or other base) bits, depending on the target encoding scheme. In some embodiments, boxes (or bits) 1802 can be written one after another to construct a "storage word". A first exemplary data format shows three components of a storage word: an address segment, a data segment, and an error detection segment. The address segment can be a tag or pointer used by the storage system to locate the target data. Unlike the hardware address lines on a computer memory bus in conventional semiconductor memories that address unique storage locations on the physical storage chip, the storage string of this disclosure can use an address (or tag) as part of the stored data and indicate the location of the data to be retrieved. FIG. 35A and FIG. 35B In the illustrated example, the address of the data written to each site on the substrate or wafer is located near or contiguous with the data, and also includes error detection data, such as parity checks, checksums, error-correcting codes (ECC), cyclic redundancy checks (CRC), or any other form of error detection and / or security information, including encrypted information. In a storage word, the address, data, and error detection components are each located sequentially in a storage string. Since each component has a known length (number of bits), for example, address = 32 bits, data = 16 bits, and error detection = 8 bits, each storage word and its components can be determined by counting the number of bits. Furthermore, as described herein and in the aforementioned jointly owned patents and patent applications, a given bit can be represented by one or more DNA bases or oligomers (e.g., boxes). When multiple bases are used to represent one or more bits (e.g., 0, 1 or 00, 01, 10, 11, etc. for a binary system, or G, C, A, T for a quaternary system), it can be referred to as a “box,” as described herein. Therefore, as used herein, the terms bit and box are used interchangeably. In some implementations, multiple numeric words (address, data, error detection) can be stored on a given DNA storage string, depending on the length of the DNA string that can be written.
[0249] refer to FIG. 35AThe exemplary data format shows the same three components: address segment, data segment, and error detection segment. However, in some embodiments, there can be a "special bit or sequence" segment S1, S2, S3 between each segment, as shown by the storage string 1812. These special bits S1, S2, S3 can be a predetermined bit sequence or code that indicates the type of the next segment, for example, 1001001001 can indicate that the next segment is an address segment, while 10101010 can indicate that the next segment is a data segment, and 1100110011 can indicate that the next segment is an error detection segment. In some embodiments, the special bits can be different molecular bits or bit structures attached to the storage string, such as dumbbell, flower, or other "large" molecular structures that are easily recognizable when the DNA storage string is read off-line outside the nanowriting chip described herein. The special bits can also not take large size structures, but have other molecular properties that provide unique variations to the polymer structure corresponding to the bit value. The storage string can take any other desired data formatting.
[0250] FIG. 36A A schematic of a method for creating unique encrypted DNA fingerprints according to embodiments of the present disclosure. Specifically, the individual variability and uniqueness of the original synthesized (or written) DNA 3602 can be further enhanced by taking a complete set of synthesized molecules 3602 and amplifying the set (or sample or group) in independent PCR reactions. Each PCR amplification reaction introduces inherent bias. And, each PCR reaction preferentially amplifies a different subset of molecules 3603, 3609, 3615 in the original mixture 3602, as shown by PCR reactions 1-3, 3604, 3610, 3616, respectively, resulting in different molecular fingerprints 3606, 3612, 3618, as shown by FIG. 36A These molecular fingerprints 3606, 3612, 3618 can then be used to create custom molecular codes for integration into a single item or group of items or for other secure data purposes.
[0251] FIG. 36BAn illustration of the three layers of data derived from a common DNA sequence according to embodiments of the present disclosure. Specifically, this illustration shows that in some embodiments, each read of the molecular code 3650 can generate three layers of data: a binary layer 3652, a production log fingerprint layer 3654, and an item fingerprint layer 3656. The binary layer 3652 is immutable and can be permanently linked to a blockchain or NFT hash or any other secure traceable database. The production lot fingerprint layer 3654 is determined by measuring the percentage (or ratio) of different DNA cassette variants used to write the bits. The original fingerprint 3650 can be stored in the blockchain. The item fingerprint layer 3656 can be viewed as a list of random numbers from each read, where the decoded sequence has a unique value. A certain number of values must match the original detected values during verification to achieve authentication. In some embodiments, all three layers 3652, 3654, 3656 are located in the same DNA sequence and are inseparable. In some embodiments, the top layer 3652 enables the sequence to be bound to a blockchain, where the blockchain contains encrypted information to verify the other two layers. If a public blockchain is used, these layers will remain even if the code manufacturer goes out of business or no longer exists, because the DNA will always be readable for a very long time in the future.
[0252] In some embodiments, in addition to performing the PCR fingerprint generation shown FIG. 36A In addition to the PCR fingerprint generation shown, additional steps can be performed to provide additional protection against unauthorized copying of the code. For example, a small amount of the original DNA sample or "seed" can be mixed into the final batch. Specifically, in some embodiments, the original synthesized (or written) DNA (or a portion thereof) can be collected in collection containers or vials, all of which are unique, and the sample is extracted and PCR amplified to create a unique fingerprint as described herein in connection with FIG. 36A At the same time, a small sample or "seed" of this unique batch is not amplified and is added to the final output mixture. The final output mixture will have a unique fingerprint but also a unique seed sequence that should not exist in any copy (as opposed to the PCR amplified sample). This method can detect if a third party attempts to copy the process by amplifying the entire sample (including the seed), which would result in a failed verification.
[0253] For example, in some embodiments, the output mixture that can be further integrated to the article comprises directly-written molecular samples. The directly-written molecular samples can be further amplified, e.g., by PCR, where the amplification process can introduce bias artifacts in the relative proportions of the original molecules, resulting in a unique mixture with an associated fingerprint. In some embodiments, the output mixture comprising amplified sequences can further comprise seeds of the original DNA mixture, e.g., where the sequences of the original DNA mixture are present only in single copy. In such cases, the output mixture can be resistant to amplification attacks, where informed analysis of the output mixture, e.g., from a sequence sampling integrated into a suspected counterfeit article and subsequently extracted therefrom, will detect and provide evidence whether an unauthorized third party sampled and amplified the output mixture, e.g., to integrate into a non-genuine or counterfeit article; such unauthorized interaction with the output mixture will be manifested in the verification analysis, e.g., where the counterfeit output mixture comprises multiple copies of the original seed molecules. In some embodiments, the directly-written molecular samples (“sample molecules”), optionally amplified, and the seed molecules can have different lengths or different numbers of nucleotides. In some embodiments, the sample molecules, optionally amplified, and the seed molecules can comprise different end cap moieties, e.g., such that primers / probes can index the molecules to be read. In some embodiments, the sample molecules and the seed molecules are generated in the same, different, or multiple production batches or reactions. In some embodiments, the sample molecules and the seed molecules comprise the same or different number or composition of cassettes. In some embodiments, the sample molecules and the seed molecules comprise the same or different chemical moieties at the end of the molecule, and / or are integrated into the same or different chemical moieties within the nucleotide backbone, and / or modify the same or different chemical moieties on the nucleotides within the cassette. In some embodiments, the sample molecules and the seed molecules are co-integrated into a microbead, e.g., a silica microbead. In some embodiments, the sample molecules and the seed molecules are separately integrated into a microbead, e.g., a silica microbead, e.g., in different microbead populations, or in the same microbead population but in different sub-regions of the microbead, e.g., where the sample molecules are located inside the silica microbead, and the seed molecules are adsorbed to the outer surface of the silica microbead, or vice versa.
[0254] FIG. 37 FIG. 38A is a schematic diagram showing an encoding / decoding system method for encoding and decoding digital files to / from DNA, according to embodiments of the present disclosure.
[0255] FIG. 38A , 38B , 38C, 38D, 38E, 38F are schematic diagrams showing an encoding / decoding system method for encoding and decoding digital files to / from DNA, according to embodiments of the present disclosure. FIG. 37This is a schematic diagram of a system that encodes digital files into DNA for writing. Specifically, after pre-lengthening the data of the original digital file and padding it to the next block size, the data is split into multiple blocks (B). Each block (B) is then divided into "nackets" or "nucleic acid packets," because each DNA storage string or nucleic acid packet can only hold a certain number of bases, corresponding to a certain number of bytes of data. For example, a storage string or nucleic acid packet can hold approximately 650-2000 DNA bases, with a box length of approximately 20-22 bases, meaning the storage string or nucleic acid packet length can range from approximately 32-100 boxes, with other values possible based on the chemical system. Therefore, with 32 boxes, each 2 bits, a storage string or nucleic acid packet can represent only 64 bits or 8 bytes (assuming 8 bits / byte). A block-level CRC (Cyclic Redundancy Check), such as CRC32 on the block, can be pre-applied to each block (B) before being divided into data payloads (W). Next, the parity check payload (Z) of the nucleic acid data packet is calculated, for example using the ZEFC standard library, which is added to the nucleic acid data packet; other CRC methods can be used if necessary. The result is the output nucleic acid data packet payload (Y) to the next stage, where Y = W + Z. Finally, a nucleic acid data packet ID (NID) is assigned to each nucleic acid data packet, where the total number of nucleic acid data packets is N = B(W + Z). Furthermore, the nucleic acid data packet CRC is calculated based on the combination of the nucleic acid data packet ID and the nucleic acid data packet payload. Therefore, the final binary code written for a given storage string or nucleic acid data packet is [nucleic acid data packet ID][CRC][nucleic acid data packet payload], such as... FIG. 38E As shown. Next, the system can use the method described herein, through the inkjet DNA writing system described herein or other DNA synthesis systems, to convert the required binary code into a storage string or nucleic acid data packet to be written onto the wafer array or chip surface.
[0256] FIG. 39A , 39B 39C, 39D, 39E, 39F, and 39G are for displaying according to the publicly available implementation scheme. FIG. 37 This diagram illustrates the method by which the system writes DNA and decodes it back into the original digital file. Specifically, the encoding process is reversed, data is extracted, and the validity of the nucleic acid data packets is determined, verified using CRC. The output nucleic acid data packets can be grouped into two sets or classified by quality score. Low-quality nucleic acid data packets may have multiple payloads. Next, the nucleic acid data packets are assembled into blocks (B), and the original blocks are identified using error correction methods such as ZEFC, SHA, or MD5 reconstruction. The blocks are then verified using CRC, and this process is repeated on all nucleic acid data packets to obtain all blocks read. These blocks are then reassembled to obtain the original raw data file.
[0257] FIG. 40A , 40B , 40C is a data graph showing results data from an encoding / decoding system using FIG. 37 the physical mix approach. Specifically, a nucleic acid data packet read count versus nucleic acid data packet address (or nucleic acid data packet ID) is shown. This data shows that a large number of full length, CRC correct and consistent nucleic acid data packets were found in the 200 byte test. FIG. 40A A pie chart 4050, 4052 is shown for the 3.5K byte test, the left pie chart 4050 shows the distribution of all reads to valid reads (9.94%), the right pie chart 4052 shows the distribution of full length nucleic acid data packet ID, unsigned reads, full length nucleic acid data packets, full length and CRC correct nucleic acid data packets, and full length, CRC correct and consistent nucleic acid data packets. FIG. 40B A bar graph (or histogram) 4060 is shown for the 3.5K byte test, where the Y axis is the nucleic acid data packet read count (log scale), and the X axis is the nucleic acid data packet address (or ID) (similar to FIG. 40C FIG. 40A
[0258] FIG. 50 An alternative embodiment is shown where computer (or CPU) generated randomness is used instead of the physical mix of cassettes to write to a randomly selected cassette mix. Specifically, according to embodiments of the disclosure, FIG. 50 is a schematic diagram showing the printhead groups 5010, 5012, 5014, 5016 of a laser inkjet DNA printer, with the individual topological cassette nozzles 5010A, 5012A, 5014A, 5016A corresponding to each head group. The printhead groups 5010, 5012, 5014, 5016 are controlled by printhead control logic (or controller) 5020, which selects the appropriate nozzle (within the printhead) to write to a site 14 on a chip or wafer. In this case, each 2-bit code is associated with 4 cassettes (“00” for C1-C4, “01” for C5-C8, “10” for C9-C12, “11” for C13-C16). As described below, the controller 5020 uses a random selection process performed by control logic (such as a QRNG (quantum random number generator) or any other desired random number generator that provides sufficiently random output) to randomly select which cassette to write to from the assigned cassettes. This logic also records each C# number selected and printed in the writing process. After the chip writing process is complete, this logic stores the total number of writes for each C# associated with each 2-bit pair, and calculates the percentage of use for each C# within each 2-bit pair, and stores a code-cassette correspondence table, which is later used in the authentication process, similar to the process performed with the physical mix approach.
[0259] In this case, each site has one-dimensional bin randomness along the length of the storage string, rather than FIG. 28 the two-dimensional randomness of the physical mixture approach. Thus, if the number of bins along the storage string or nucleic acid data packet is insufficient to provide authentication, authentication can be performed across multiple sites (e.g., one or more rows) of the chip 5100, as FIG. 51A shown. In this case, the logic calculates the percentage of use of each C# within each 2-bit pair in a given row (or group of rows) and stores the results in a row-based code-bin correspondence table 5102, each row having a computer-generated, random proportion of bins (Cs) associated with each two-bit code.
[0260] In some embodiments, the logic can calculate the percentage of use of each C# within each 2-bit pair across the entire chip or array, as FIG. 51B shown. In this case, the logic calculates the percentage of use of each C# within each 2-bit pair across the entire chip 5110 and stores the results in a chip-based code-bin correspondence table 5112, the entire chip having a computer-generated, random proportion of bins (Cs) associated with each two-bit code across the entire chip, and each chip or lot can be a different set of proportions.
[0261] Referring to FIG. 52 , in some embodiments, the printhead group 5200 can have all of the bins of the entire chip and independent bin nozzles, e.g., C1-C16, each of which can be individually addressable by the controller 5020. In this case, the control logic 5020 determines the required bin C1-C16 to be written based on the bin assignment for each 2-bit code and selects that bin for writing and performs the writing. This logic can be similar to the above FIG. 50 described, except that in some embodiments, only a single control line is required rather than multiple control lines and multiple printhead groups.
[0262] Referring to FIG. 53A , which shows a flowchart of writing (printing) and unloading an encoded polymer storage string in an inkjet writing system using computer-based randomness for bin writing selection, according to embodiments of the present disclosure. In particular, this logic is similar to the logic of FIG. 31A including blocks 3102-3132, except that rather than acquiring four cartridges with different mixtures, a printhead with a single assigned bin group is acquired, as shown in block 5302 (instead of block 3106). Further, block 5304 is provided for writing / printing the 2-bit code, referring to the writing process in FIG. 53B instead of block 3114 referring to the writing process in FIG. 31B .
[0263] Referring toFIG. 53B which shows a flowchart of writing (printing) 2-bit codes to DNA / polymer storage strands in an inkjet writing system using computer-based randomness for cartridge write selection according to embodiments of the present disclosure. Specifically, the logic is similar to that of FIG. 31B with the difference that instead of printing bits using pre-set mixture cartridges, the logic obtains the cartridges containing the corresponding Cs of the 2-bit code to be written from corresponding blocks 5354, 5358, 5362, 5366 in the logic, and then randomly selects a C# from the Cs assigned to the 2-bit code. Next, the corresponding blocks 5354, 5358, 5362, 5366 print the corresponding 2-bit code at the target site / position on the array or chip with the randomly selected C#. Then the corresponding blocks 5354, 5358, 5362, 5366 increment the corresponding C# count value for the chip and / or row being written. The process continues until block 5368 determines that the bit writing for the site or group of sites is complete. When complete, the determination of block 5368 is “Yes” and block 5370 saves the C# count value in the code-cartridge correspondence table. Next, block 5372 determines whether the drop observer (or sensor) 1911 (shown in FIG. 30A ) detects any drop errors, which can be part of the printhead and array platform controller and detection logic 1908 (shown in FIG. 30A ). If the result is “Yes”, an error is detected and block 5374 saves the error location and bit number for later reading and the logic exits. If the determination of block 5372 is “No”, the write cycle did not find a drop error and the logic exits. The process can use a pre-set nucleic acid data packet ID to row conversion rule. Specifically, if a pre-determined number of points per row and a pre-determined number of redundant points for error protection are provided, the system can have a pre-determined number of rows (or corresponding nucleic acid data packets) to achieve sufficient authentication through cartridge C# proportion verification. In addition, FIG. 53B the logic 5350 also stores and updates the C# count value in the code-cartridge correspondence table after each site write is complete for later use in the authentication process.
[0264] Reference is made to FIG. 54 which shows a flowchart 5400 of decoding and confirming polymer storage strand data when using computer-based randomness for cartridge write selection according to embodiments of the present disclosure. Specifically, the logic 5400 is similar to that of FIG. 34The logic 3400 (including blocks 3402-3440) is similar, except that instead of performing an authentication check after each nucleic acid data packet ID is decoded, the logic waits until all of the storage string / nucleic acid data packet IDs (or at least up to the number of rows or nucleic acid data packet IDs used for verification) are decoded, then checks the proportion of C# count values to determine whether the ID, row, or chip passed authentication, as shown, with block 3442 performed after block 3430.
[0265] It will be appreciated that the surface of the substrate to be written can be flat (unpatterned) or patterned.
[0266] In some embodiments, the present disclosure can be used in conjunction with non-fungible tokens (NFTs), tokens, contract addresses, PKI components, digital certificates, private database identifiers, ERP database identifiers, new device UDIs, global trade numbers, GTINs, UPC codes, QR codes, EANs, ISBNs, Library of Congress numbers, FNSKUs, ITF-14s, contract IDs (e.g., DOD, Dod CIC credentials), patient identifiers, EMR records (e.g., Epic system patient IDs), contractor license numbers, professional certification numbers, notary identification numbers, building construction permit numbers, building construction or quality control inspector numbers. The present disclosure can also be used in conjunction with physical currency (paper, metal, etc.) and digital currency, including cryptocurrencies such as payment cryptocurrencies, tokens, stable coins, central bank digital currencies, etc., specifically including Bitcoin, Ethereum, Tether, Ripple, Binance Coin, USDCoin, Cardano, Solana, Dogecoin, Polygon, etc., including but not limited to other cryptocurrencies that are currently known or subsequently discovered, developed, that can employ independent blockchains. Further, the systems and methods of the present disclosure can authenticate items by retaining, querying, or verifying production fingerprints and / or molecular fingerprints, which can be performed in public databases or in separate authentication databases, which can employ hashing, other / additional encryption, or plaintext. Additionally, as described herein, in some embodiments, the data encoded by the present disclosure can be a non-fungible token (NFT), and the authentication data and / or encoded data can be stored on a blockchain.
[0267] In some embodiments, the present disclosure provides an item authentication method according to Method 1 (Method 1A), wherein a nucleic acid data package is synthesized by sequentially adding cassettes to a DNA acceptor strand using an inkjet printhead (e.g., a piezoelectric printhead); wherein each cassette comprises a plurality of nucleotides; wherein, at each sequential addition step, the cassette comprises a population of heterogenous cassettes comprising cassettes having at least two different sequences that encode the same data in a machine-readable code (e.g., binary code or ternary code); and wherein the cassettes are dispensed by an inkjet writing printhead to at least one writing site of a wafer array, the printhead or nozzle writing the same code to a plurality of polymer storage strands dispensed to the at least one site, e.g., the method comprises the steps of: a) loading a starting polymer or DNA at a target site to be written to, one end of which is anchored to the target site; b) washing the surface of the site; c) positioning an inkjet nozzle loaded with a population of heterogenous cassettes comprising cassettes having at least two different sequences that encode the same information in one or more bits (e.g., 1 or 0 in binary code, or 00, 01, 10, 11, etc.) corresponding to a unique code, over a target site to be written to; d) causing the inkjet nozzle to release a droplet comprising the population of heterogenous cassettes to the site, thereby writing one bit or portion of the unique code to a DNA or polymer storage string (or strand) associated with the site; and e) washing the surface of the site.
[0268] Optionally, the method further comprises steps f) - i): f) causing the inkjet nozzle to release a droplet of a deblocking / adapter reagent to the site; g) washing the surface of the site; and h) repeating steps c) to g) until the unique code is fully written to the storage string of the site; i) removing the storage string from the site and introducing it into a collection or storage vessel for subsequent integration into or onto an item.
[0269] For example, in the above method, the cassettes can be added by a topoisomerase-mediated ligation reaction, e.g., by: (i) reacting a double-stranded acceptor DNA strand with a topoisomerase loaded with double-stranded DNA cassettes from a population of heterogenous cassettes covalently bound to the topoisomerase; wherein one strand of the acceptor DNA has a 5’ overhang; wherein each cassette comprises an information sequence, a topoisomerase recognition sequence, and a 5’ overhang on both strands; wherein the 5' overhang of the oligomer strand not carrying the topoisomerase ("bottom strand") is complementary to the 5' overhang of the acceptor DNA, but not to the 5' overhang of the strand carrying the topoisomerase ("top strand") in the cassette; wherein the 5' end of the strand carrying the topoisomerase ("top strand") in the cassette is not protected, e.g. not phosphorylated (i.e. 5'-OH), from the 5' end of the acceptor DNA; and wherein the topoisomerase loaded with the double-stranded DNA cassette is delivered to the location of the acceptor strand by a piezoelectric inkjet nozzle; (ii) reacting the extended acceptor DNA of step (i) with a topoisomerase loaded with another double-stranded DNA cassette; wherein the other cassette comprises an information sequence identical or different to the information sequence of the cassette of step (i), a topoisomerase recognition sequence, and 5' overhangs on both strands; wherein the 5' overhang of the strand not carrying the topoisomerase ("bottom strand") in the other cassette is complementary to the 5' overhang of the extended acceptor DNA, but not to the 5' overhang of the strand carrying the topoisomerase ("top strand") in the other cassette; and wherein the 5' end of the strand carrying the topoisomerase ("top strand") in the other cassette is not protected, e.g. not phosphorylated (i.e. 5'-OH); and (iii) repeating steps (i) and (ii) until the target nucleotide sequence is obtained; wherein, optionally, a washing step is performed after step (i) and / or after step (ii).
[0270] For example, in some embodiments, the present disclosure provides a method of writing target binary codes employing DNA or polymer strands or storage string, the target binary codes comprising a plurality of 2-bit binary codes, the method comprising: providing a plurality of unique DNA cassettes for writing four different 2-bit binary codes, which is a predetermined unique set of the plurality of DNA cassettes associated with each of the four 2-bit binary codes, each DNA cassette having a same length defined by a predetermined number of sites, each site comprising one of four DNA or polymer bases; providing four inkjet cartridges, each cartridge associated with a different 2-bit binary code, the liquid within each cartridge comprising a different predetermined DNA cassette mixture of the set of DNA cassettes associated with the given 2-bit binary code; wherein the predetermined DNA cassette mixture is associated with a current batch or date code; obtaining a first 2-bit binary code from the target binary codes to be written on a substrate surface; writing the first 2-bit binary code by applying a droplet of the liquid in the cartridge associated with the first 2-bit binary code to a storage write site on the substrate surface, the droplet comprising the DNA cassettes associated with the first 2-bit binary code; wherein the DNA cassettes in the droplet are connected to the existing DNA cassettes on the surface in a random arrangement based at least on DNA cassette connection kinetics and the DNA cassettes in the droplet; repeating the obtaining and writing steps for subsequent 2-bit binary codes until the target binary codes for a given storage site on the substrate surface are written; wherein each writing step generates a random arrangement of the DNA cassettes associated with the current 2-bit binary code that are connected to the existing DNA cassettes on the surface, thereby forming a plurality of storage strings at the given storage site; wherein the total distribution of all DNA cassettes associated with a given 2-bit binary code across all storage strings is substantially consistent with the predetermined DNA cassette mixture within a predetermined tolerance; and wherein the DNA cassettes associated with a given 2-bit binary code are randomly distributed along a given storage string.
[0271] Additionally, in some embodiments, the unique set of the plurality of DNA cassettes associated with each of the four 2-bit binary codes varies with batch number or time code. Further, in some embodiments, the predetermined unique set of the plurality of DNA cassettes associated with each of the four 2-bit binary codes comprises a unique set of four. Further, in some embodiments, the plurality of unique DNA cassettes for writing four different 2-bit binary codes comprises 16 unique DNA cassettes. Further, in some embodiments, the number of the plurality of unique DNA cassettes for writing four different 2-bit binary codes is an integer greater than 2. Further, in some embodiments, the unique set of the plurality of DNA cassettes is associated with each of the four 2-bit binary codes.
[0272] Further, in some embodiments, the number of sites that make up the length of the DNA cassette comprises an integer greater than 3. Further, in some embodiments, each site comprises one of four DNA bases plus an additional polymer unit, wherein each site comprises one of at least five unique polymer units. Further, in some embodiments, the first DNA cassette comprises a start cassette or a target sequence that is not part of the target binary code to be written. Further, in some embodiments, the DNA cassette comprises a topology cassette having a topoisomerase portion and a cassette binary code portion. Further, in some embodiments, the 2-bit binary code comprises an n-bit binary code. In some embodiments, the predetermined DNA cassette mix associated with each 2-bit binary code is derived from a batch number or date code. Further, in some embodiments, the 2-bit binary code can be an n-bit binary code.
[0273] In another aspect, the present disclosure provides a method of synthesizing DNA, such as DNA 2 and any of its subsequent sequences, wherein the DNA comprises a transition between non-identical nucleotides corresponding to a series of bits in a machine-readable code (e.g., a ternary code), the method comprising: stepwise addition of nucleotides (dNTPs) in a kinetically controlled reaction mixture comprising one or more transferases (e.g., terminal deoxynucleotidyl transferase, TdT) and one or more dNTP-degrading enzymes (e.g., apyrase), wherein each step addition employs a different nucleotide. In this method, during each step addition, an indeterminate number of nucleotides (e.g., about 5-15, depending on the optimal ratio of TdT to apyrase) are added to each strand before the dNTPs are consumed by the apyrase, followed by addition of a different dNTP, so that the resulting strands have different lengths, and the data is encoded in the transition between non-identical nucleotides, which is the same for each strand, thereby providing a population of heterogenous nucleic acid data packets, wherein each nucleic acid data packet comprises a plurality of DNA molecules encoding the same data (here, the data is at the junction of non-identical nucleotides), wherein the sequence of the DNA molecules has heterogeneity (here, the length of the contiguous sequence of identical nucleotides is variable). With four natural dNTPs, there are three possible transitions for each nucleotide, e.g., AT / AC / AG, TA / TC / TG, CA / CG / CT, and GC / GA / GT. This possibility enables more synonymous heterogenous sequences, e.g., with a ternary code comprising 0, 1, 2, each of 0, 1, 2 can be represented by any one of the four different transitions (see, e.g., FIG. 41 the set of possible permutations shown).
[0274] The present disclosure also provides methods of decoding a population of DNA molecules; for example, sequencing the population of DNA molecules, filtering out identical nucleotide sequences to compress the DNA molecule sequences into compressed representative sequences, and decoding the compressed representative sequences back to the original data string using the scheme used in encoding the data. Alternatively or additionally, statistical inference methods and / or models can be used to further analyze the sequences of the population of DNA molecules, for example, as disclosed in Lee H.H. et al. “DNA data storage using synthetic homogenous DNA.” Nat. Commun. (2019) 10:2383, the contents of which are incorporated herein by reference. Terminator-free template-independent enzymatic DNA synthesis for digital information storage
[0275] For example, the present disclosure provides a method (Method 2) of writing a target code (e.g., ternary code) using DNA strands, the method comprising: i. providing a reaction mixture comprising one or more transferases (e.g., terminal deoxynucleotidyl transferase, TdT) and one or more dNTP-degrading enzymes (e.g., apyrase); ii. adding deoxyribonucleotide triphosphates (dNTPs) to the reaction mixture; iii. waiting until the dNTPs in step ii) are added to DNA strands or degraded; iv. repeating steps ii) and iii) until the target bit sequence is obtained; wherein a different dNTP species is used in any two consecutive additions; thereby obtaining a population of DNA molecules encoding the target data string.
[0276] For example, in specific embodiments, the present disclosure provides: 2.1 According to Method 2, further comprising the steps of: v. optionally, storing the reaction mixture for subsequent addition, purification, or processing; vi. purifying the synthetic DNA or polymer strands or storage strings comprising the data string; and vii. optionally, storing the purified DNA or polymer strands or storage strings for subsequent use, analysis, addition, purification, or processing.
[0277] 2.2 Any preceding method, wherein the reaction mixture comprises terminal deoxynucleotidyl transferase (TdT).
[0278] 2.3 Any preceding method, wherein the reaction mixture further comprises apyrase.
[0279] 2.4 Any preceding method, wherein the reaction mixture is an aqueous system, e.g., a buffer.
[0280] 2.5 Any of the preceding methods, wherein the reaction mixture further comprises other additives, such as ions, such as cations, such as divalent cations, such as cobalt.
[0281] 2.6 Any of the preceding methods, wherein the reaction mixture comprises a mixture of TdT and an adenosine triphosphatase, such as both in stoichiometric ratios to achieve kinetic control stepwise addition of dNTPs.
[0282] 2.7 Any of the preceding methods, wherein the dNTPs comprise adenosine triphosphate (ATP), guanosine triphosphate (GTP), cytosine triphosphate (CTP), thymidine triphosphate (TTP); optionally, further comprising uridine triphosphate (UTP).
[0283] 2.8 Any of the preceding methods, wherein the 3-bit ternary code comprises an n-bit ternary code.
[0284] 2.9 Any of the preceding methods, wherein the synthetic DNA or polymer chain or storage string, or population of DNA molecules synthesized comprises any of DNA 2 and its subsequent sequences.
[0285] 2.10 Any of the preceding methods, for use in combination with any of method 1 and its subsequent methods, method 3 and its subsequent methods, method 4 and its subsequent methods, method 5 and its subsequent methods, method 6 and its subsequent methods, and / or method 7 and its subsequent methods.
[0286] Accordingly, the present disclosure provides an item authentication method (method 3) comprising: i. synthesizing a DNA sequence comprising nucleic acid packets (nackets), wherein each nucleic acid packet comprises a plurality of DNA molecules encoding the same data, wherein the sequence of the DNA molecules is synthesized using one or more transferases (such as terminal deoxynucleotidyl transferase, TdT); ii. integrating the DNA sequence into or onto an item; iii. extracting the DNA sequence from the item; and iv. analyzing the extracted DNA sequence; v. optionally, aligning the analyzed DNA sequence to a database of DNA sequences; vi. optionally, confirming item authenticity.
[0287] For example, in particular embodiments, the present disclosure provides: 3.1 Method 3, wherein the DNA sequence encodes data used as an identification code for the item.
[0288] 3.2 Method 3.1, wherein the data used as an identification code is randomly generated.
[0289] 3.3 Any preceding method, wherein the DNA sequence comprises any one of DNA 2 and its subsequent sequences.
[0290] 3.4 Any preceding method, wherein the DNA sequence is synthesized by sequentially adding homopolymer stretches, wherein each subsequent homopolymer stretch comprises a different nucleotide than the adjacent homopolymer stretch.
[0291] 3.5 Method 3.4, wherein the homopolymer stretches are synthesized using a transferase (e.g., terminal deoxynucleotidyl transferase, TdT).
[0292] 3.6 Method 3.4, wherein the homopolymer stretches are synthesized using TdT.
[0293] 3.7 Any preceding method, wherein the DNA sequence is integrated into the article by direct surface conjugation to the article.
[0294] 3.8 Any preceding method, wherein the DNA sequence is integrated into a component part or material used to manufacture the article, optionally into a textile, fabric, leather, biomaterial article, polymer, plastic, wood, metal, ink, paint, solution, suspension, and raw material.
[0295] 3.9 Any preceding method, wherein the DNA sequence is encapsulated in a microcontainer prior to integration into an article, optionally a microsphere, optionally a silica microsphere.
[0296] 3.10 Any preceding method, wherein the DNA sequence is encapsulated in a molecular assembly, such as a lipid nanoparticle, protein complex or aggregate, or crystal lattice.
[0297] 3.11 Any preceding method, wherein the DNA sequence is inserted into one or more cells, optionally into a larger DNA construct and / or genome, optionally into a yeast, bacterial, fungal, plant, or animal cell, optionally wherein the cell is used to produce a food, drink, biomaterial, or material, such as cheese, beer, wine, vegan leather, a pharmaceutical.
[0298] 3.12 Any preceding method, wherein the integrated DNA sequence is extracted from the article by physical means, optionally by cutting, grinding, scoring, slicing, shredding, or pulverizing one or more portions of the article.
[0299] 3.13 Any preceding method, wherein the integrated DNA sequence is extracted from the article by chemical means, optionally by dissolving or lysing one or more portions of the DNA sequence and / or the article.
[0300] 3.14 Any of the preceding methods, wherein the extracted DNA sequence is isolated and / or purified by chromatography, such as ion exchange chromatography, size exclusion chromatography, normal or reverse phase high performance liquid chromatography (HPLC), antibody affinity chromatography, or combinations thereof.
[0301] 3.15 Any of the preceding methods, wherein the extracted DNA sequence is isolated and / or purified by immobilization, such as solid phase reversible immobilization (SPRI), immunoprecipitation (or antibody pulldown), or combinations thereof; further, optionally in solution, resin, slurry, magnetic beads, filter, or combinations thereof.
[0302] 3.16 Any of the preceding methods, wherein the extracted DNA sequence is isolated and / or purified by electrophoresis, such as polyacrylamide gel electrophoresis, two-dimensional electrophoresis, pulse field electrophoresis, Southern blotting, or combinations thereof.
[0303] 3.17 Any of the preceding methods, wherein the extracted DNA sequence is isolated and / or purified by centrifugation, further optionally by filtration (e.g., centrifugal column).
[0304] 3.18 Any of the preceding methods, wherein the extracted DNA sequence is analyzed by mass spectrometry and / or high-throughput DNA sequencing.
[0305] 3.19 Any of the preceding methods, wherein the extracted DNA sequence is compared to a database comprising original synthetic identification codes of the item.
[0306] 3.20 Any of the preceding methods, wherein the extracted DNA sequence is compared to one or more previous analysis results of DNA sequences extracted from the same or similar item.
[0307] 3.21 Any of the preceding methods, for use in combination with any of Method 1 and its subsequent methods, Method 2 and its subsequent methods, Method 4 and its subsequent methods, Method 5 and its subsequent methods, Method 6 and its subsequent methods, and / or Method 7 and its subsequent methods.
[0308] 3.22 Any of the preceding methods, wherein the DNA sequence comprises any of DNA 1 and its subsequent sequences and / or DNA 2 and its subsequent sequences.
[0309] Accordingly, the present disclosure provides a method of writing anti-attack digital codes using DNA (Method 4), comprising: i. receiving a target digital code to be written, the target code being divided into four two-bit binary codes to be written (e.g., 00, 01, 10, 11); ii. providing a predetermined mixture of four predetermined numbers of unique DNA cassette strings, each mixture corresponding to a different predetermined binary code value, each mixture having a predetermined proportion of unique DNA cassettes within the mixture, and each mixture having unique DNA cassette strings different from the DNA cassette strings of the other mixtures; iii. depositing a mixture droplet associated with a given binary code to be written to the substrate to add a DNA cassette string to the encoding DNA string being written, the droplet containing a predetermined mixture of unique cassettes associated with the given binary code; and iv. repeating the depositing step until the target code is written to the encoding DNA string.
[0310] For example, in particular embodiments, the present disclosure provides: 4.1 According to method 4, further comprising adding an end cap to the encoding DNA string after the target code is written.
[0311] 4.2 According to method 4.1, wherein the end cap contains information about the target digital code or how to read the code.
[0312] 4.3 Any preceding method, wherein the substrate has a recipient DNA strand, one end of which is attached to the substrate and the other end of which can be attached to one unique DNA cassette to be added.
[0313] 4.4 Any preceding method, wherein the predetermined number of unique DNA cassettes of one mixture is different from at least one other mixture.
[0314] 4.5 Any preceding method, wherein the target digital code is encoded with authentication data in the NFT and stored in the blockchain.
[0315] 4.6 Any preceding method, wherein the encoding DNA string is embedded in a physical item to be authenticated.
[0316] 4.7 Any preceding method, for use in combination with any one of method 1 and its subsequent methods, method 2 and its subsequent methods, method 3 and its subsequent methods, method 5 and its subsequent methods, method 6 and its subsequent methods, and / or method 7 and its subsequent methods.
[0317] 4.8 Any preceding method, wherein the DNA sequence comprises any one of DNA 1 and its subsequent sequences and / or DNA 2 and its subsequent sequences.
[0318] Accordingly, the present disclosure provides a method of writing an attack-resistant digital code using DNA (method 5), comprising: i. receiving a target digital code to be written, the target code divided into a plurality of n-bit binary codes to be written, where n is greater than 1; ii. providing a predetermined mixture of at least two, a predetermined number of unique DNA cassette strings, each mixture corresponding to a different predetermined n-bit binary code value, each mixture having a predetermined proportion of unique DNA cassettes in the mixture, and each mixture having a different unique DNA cassette string than the other mixtures; iii. depositing a mixture droplet associated with a given n-bit binary code to be written to the substrate to add a DNA cassette string to the encoding DNA string being written, the droplet containing a predetermined mixture of unique cassettes associated with the given n-bit binary code; and iv. repeating the depositing step until the target code is written to the encoding DNA string.
[0319] For example, in particular embodiments, the present disclosure provides: 5.1 According to method 5, further comprising adding an end cap to the encoding DNA string after the target code is written.
[0320] 5.2 According to method 5.1, wherein the end cap contains information about the target digital code or how to read the code.
[0321] 5.3 Any preceding method, wherein the substrate has a recipient DNA strand, one end of which is attached to the substrate and the other end of which can be attached to one unique DNA cassette to be added.
[0322] 5.4 Any preceding method, wherein the predetermined number of unique DNA cassettes of one mixture is different from at least one other mixture.
[0323] 5.5 Any preceding method, wherein the target digital code is encoded with authentication data in the NFT and stored in the blockchain.
[0324] 5.6 Any preceding method, wherein the encoding DNA string is embedded in a physical item to be authenticated.
[0325] 5.7 Any preceding method, used in combination with any of method 1 and its subsequent methods, method 2 and its subsequent methods, method 3 and its subsequent methods, method 4 and its subsequent methods, method 6 and its subsequent methods, and / or method 7 and its subsequent methods.
[0326] 5.8 Any preceding method, wherein the DNA sequence comprises any of DNA 1 and its subsequent sequences and / or DNA 2 and its subsequent sequences.
[0327] Accordingly, the present disclosure provides a method of writing attack-resistant digital codes using DNA (method 6), comprising: i. receiving a target digital code to be written, the target digital code being divided into four two-bit binary codes (e.g., 00, 01, 10, 11) to be written; ii. providing a collection of four unique DNA cassette strings, each collection containing a predetermined number of unique DNA cassettes, and each collection corresponding to a different predetermined two-bit binary code value, such that each collection of unique cassettes corresponds to a different two-bit binary code, and each collection of unique cassette strings is different from the other DNA cassette strings; iii. randomly selecting one of the unique cassettes corresponding to a given two-bit binary code to be written as a selected unique cassette; iv. depositing droplets of the selected unique cassette associated with a given two-bit binary code to be written to the substrate to add the selected unique cassette to the encoded DNA string being written at a given writing site on the substrate; v. repeating the selecting and depositing steps until the target code is written to the encoded DNA string at the given writing site on the substrate; and vi. counting the number of times each unique cassette is used in writing each two-bit binary code.
[0328] For example, in particular embodiments, the present disclosure provides: 6.1 According to method 6, further comprising adding an end cap to the encoded DNA string after the target code is written.
[0329] 6.2 According to method 6.1, wherein the end cap contains information about the target digital code or how to read the code.
[0330] 6.3 Any of the preceding methods, wherein the substrate has a recipient DNA strand that is attached to the substrate at one end and is available to attach to one of the unique DNA cassettes to be added.
[0331] 6.4 Any of the preceding methods, wherein the predetermined number of unique DNA cassettes of one mixture is different from at least one other mixture.
[0332] 6.5 Any of the preceding methods, wherein the target digital code is encoded with authentication data on the NFT and stored on a blockchain.
[0333] 6.6 Any of the preceding methods, wherein the encoded DNA string is embedded in a physical item to be authenticated.
[0334] 6.7 Any of the preceding methods, used in combination with any of method 1 and its subsequent methods, method 2 and its subsequent methods, method 3 and its subsequent methods, method 4 and its subsequent methods, method 5 and its subsequent methods, and / or method 7 and its methods.
[0335] 6.8 Any of the preceding methods, wherein the DNA sequence comprises any of DNA 1 and its subsequent sequences and / or DNA 2 and its subsequent sequences.
[0336] Accordingly, the present disclosure provides a method of writing anti-attack digital code using DNA (Method 7), comprising: i. receiving a target digital code to be written, the target digital code being divided into a plurality of n-bit binary codes to be written, wherein n is greater than 1 ; ii. providing a set of at least two, unique DNA cassette strings, each set comprising a predetermined number of unique DNA cassettes, and each set corresponding to a different predetermined n-bit binary code value, such that each set of unique cassettes corresponds to a different n-bit binary code, and each set of unique cassette strings is different from the other DNA cassette strings; iii. randomly selecting one of the unique cassettes corresponding to a given n-bit binary code to be written as a selected unique cassette; iv. depositing droplets of the selected unique cassette associated with a given n-bit binary code to be written to a substrate to add the selected unique cassette to the encoded DNA string being written at a given writing site on the substrate; v. repeating the selecting and depositing steps until the target code is written to the encoded DNA string at the given writing site on the substrate; and vi. counting the number of times each unique cassette is used in writing each n-bit binary code.
[0337] For example, in specific embodiments, the present disclosure provides: 7.1 According to Method 7, further comprising adding an end cap to the encoded DNA string after the target code is written.
[0338] 7.2 According to Method 7.1, wherein the end cap comprises information about the target digital code or how to read the code.
[0339] 7.3 Any of the preceding methods, wherein the substrate has a recipient DNA strand, one end of which is attached to the substrate and the other end of which can be attached to one unique DNA cassette to be added.
[0340] 7.4 Any of the preceding methods, wherein the predetermined number of unique DNA cassettes of one mixture is different from at least one other mixture.
[0341] 7.5 Any of the preceding methods, wherein the target digital code is encoded with authentication data in an NFT and stored in a blockchain.
[0342] 7.6 Any of the preceding methods, wherein the encoded DNA string is embedded in a physical item to be authenticated.
[0343] 7.7 Any of the foregoing methods for use in combination with any of Method 1 and its subsequent methods, Method 2 and its subsequent methods, Method 3 and its subsequent methods, Method 4 and its subsequent methods, Method 5 and its subsequent methods, and / or Method 6 and its subsequent methods.
[0344] 7.8 Any of the foregoing methods, wherein the DNA sequence comprises any of DNA 1 and its subsequent sequences and / or DNA 2 and its subsequent sequences.
[0345] The systems, computers, servers, devices, etc. described herein have the electronic devices, computing processing capabilities, interfaces, memory, hardware, software, firmware, logic / state machines, databases, microprocessors, communication links (wired or wireless), displays or other visual / audio user interfaces, printing devices, and any other input / output interfaces necessary to implement the functions described herein or to achieve the results described herein. Unless otherwise explicitly or implicitly stated herein, the process or method steps described herein can be implemented in software modules (or computer programs) that are executed on one or more general purpose computers. Specialized hardware can be substituted for performing certain operations. Thus, any of the methods described herein can be performed by hardware, software, or any combination thereof. Moreover, a computer readable storage medium can store instructions which, when executed by a machine (e.g., a computer), implement the operations described in any of the embodiments described herein.
[0346] Additionally, the computers or computer-based devices described herein can include any number of computing devices capable of implementing the functions described herein, including but not limited to: tablet computers, notebook computers, desktop computers, smartphones, mobile communication devices, smart televisions, set-top boxes, e-readers / players, etc.
[0347] Although the present disclosure is described herein using exemplary techniques, algorithms, or processes, those skilled in the art will appreciate that other techniques, algorithms, processes, or other combinations and sequences of the techniques, algorithms, processes described herein can be employed or executed, as long as the same functions and / or results described herein are achieved, and fall within the scope of the present disclosure.
[0348] Any process description, step, or logic flow diagram block herein represents a possible implementation and does not imply a fixed order or sequence of steps; alternative implementations are within the scope of the systems and methods disclosed herein, and can be performed in different orders or concurrently, including substantially concurrently or in reverse order, depending on the functionality involved, as would be understood by those reasonably skilled in the art.
[0349] It should be understood that any feature, function, property, alternative or modification described herein in relation to one embodiment can also be applicable to any other embodiment unless otherwise specified. Further, the drawings are not drawn to scale unless otherwise specified.
[0350] Conditional language, such as, among others, “can,” “could,” “might,” or “may,” unless specifically stated otherwise, generally are intended to convey that certain embodiments include, while other embodiments do not include, certain features, elements, or steps. Thus, such conditional language is not generally intended to imply that features, elements, or steps are in any way required for one or more embodiments or that one or more embodiments necessarily include logic for deciding, with or without user input or prompting, whether these features, elements, or steps are included or are to be performed in a particular embodiment.
[0351] Embodiments There are several aspects to consider when using DNA for item authentication and provenance, for example: Encoding: converting a machine-readable code (e.g. binary code, identification code, NFT) into DNA (e.g. nucleic acid data package).
[0352] Availability: either directly integrating the free DNA strand into the item, or alternatively encapsulating the DNA in, for example, silica microbeads or microspheres.
[0353] Formulation: a method of physically mixing the DNA (either free or encapsulated) into the target item or material, for example a material used for subsequent item manufacturing.
[0354] Application: a method of using the formulated item or material, for example applying ink to paper, applying paint to canvas or plasterboard, etc.
[0355] Sampling: a method of extracting the DNA (either free or encapsulated) from the item or material; optionally, further removing the encapsulated DNA from the encapsulating material (e.g. silica microbeads or microspheres).
[0356] Reading: a method of DNA analysis, for example DNA sequencing; optionally, further comprising one or more amplification steps, for example PCR amplification.
[0357] Decoding: a method of optionally converting the DNA sequence into the original machine-readable code, i.e. reconstructing the original data file.
[0358] Example 1: item authentication using a fountain pen ink To illustrate one embodiment of this disclosure, six commercially available fountain pen inks of different colors were selected. Each ink was labeled Ink No. 1 through Ink No. 6, and each ink was serially diluted four times in a 10-fold series. Additionally, a topoisomerase-mediated heterogeneous DNA cassette data writing method was used to encode a 32-byte NFT along with relevant metadata and error correction features into a synthesized DNA strand, each strand of which contained 51 nucleic acid data packets. DNA was added to each ink sample at a concentration of 0.3 ng / µL (i.e., inks No. 1 through No. 6, four dilutions for each ink). As a preliminary assessment, the DNA was added to the ink samples, thoroughly mixed, and immediately aliquoted for DNA analysis. Subsequently, the DNA was isolated and amplified to verify that its introduction into the ink during the article (i.e., ink) certification process did not have any adverse effects.
[0359] Next, Ink No. 4 and Ink No. 5 were selected for further evaluation, as both are black inks and the above experiments showed that color did not appear to affect DNA. The DNA encoding the NFT was integrated into the pen ink as described above. The ink was then filled into a pen and used to write on commercially available printing paper. Analysis was performed after 7 days to assess the stability of the DNA in the liquid ink and on the paper after writing / drying. The DNA was then isolated and amplified. When sampling directly from the ink solution, the ink solution was diluted aliquots and amplified directly by PCR. When sampling from dried ink on paper, the dried ink area was gently wiped with a moistened cotton swab, the swab was immersed in a small amount of water, and then amplified by PCR. Alternatively, dried ink on paper can be sampled by adding a small amount of water (e.g., 10 µL) to the dried ink area on the paper, dissolving some of the dried ink, aspirating it with a pipette, and then amplifying by PCR. In these examples, the resulting liquid is typically significantly diluted (e.g., >1 / 1000) before PCR.
[0360] Gel electrophoresis of the amplified DNA showed that the DNA in Ink No. 4 remained stable after 7 days. Surprisingly, the concentration of DNA amplified from the liquid ink of Ink No. 5 was significantly lower than that of Ink No. 4. In contrast to the liquid ink samples, the DNA in both Ink No. 4 and Ink No. 5, written on paper on day 0, showed stability at day 7. Notably, further observation revealed that the DNA recovery rates from the liquid ink samples and the paper-written samples in Ink No. 4 were similar. Based on the observed stability, Ink No. 4 was used for subsequent evaluation.
[0361] Subsequently, Ink Sample 4 was used for deep sequencing analysis of the DNA encoding the NFT, and the results are as follows: FIG. 42 As shown. More specifically, the NFT uses 51 nucleic acid data packets encoded into DNA, and the heterologous DNA box writing method used in synthesizing the DNA strand encoding the NFT can provide approximately 10 9a collection of unique DNA sequences. Aliquots taken directly from this synthetic collection of DNA sequences were subjected to PCR analysis, and approximately 10 6 a collection of unique DNA sequences (i.e., 1,623,092 unique DNA sequences). This collection of DNA encoding NFTs was incorporated into Ink #4 and used to write on paper as described previously. Two dried ink samples written on paper were analyzed using PCR and deep sequencing, and were labeled as Ink Sample #1 and Ink Sample #2. During analysis, it was observed that Ink Sample #1 had 5,160 unique DNA sequences (1,311 of which were common to the original DNA sequences identified from the previously analyzed collection of DNA sequences); and Ink Sample #2 had 6,218 unique DNA sequences (2,615 of which were common to the original DNA sequences identified from the previously analyzed collection of DNA sequences). In addition, there were 442 unique DNA sequences common between Ink Sample #1 and Ink Sample #2. This demonstrates that the heterologous DNA cassette data writing method produces significant heterogeneity among DNA sequences, even though each DNA strand is ultimately synonymous with every other DNA strand from the same original collection of DNA strands.
[0362] The above protocol was repeated to further assess the stability of the DNA encoding NFTs written in ink on paper over time. More specifically, DNA written in ink on paper was extracted and analyzed at 2 and 6 weeks after writing on paper. Notably, as shown in FIG. 43 , the DNA stability at 2 and 6 weeks was very similar, with no significant difference between time points. In addition, the full-length nucleic acid data packets for each of the 51 nucleic acid data packet sites were readily identified upon deep sequencing of the recovered DNA strands, indicating that there was no significant breakage of the DNA strands. Moreover, the sequenced nucleic acid data packets produced consensus sequences at each data packet site that were useful in decoding the DNA sequences back to the original NFT code. By decoding in the manner described herein (e.g., by using the consensus sequences), the original NFT code could be reliably recovered, and the item (i.e., the ink) was easily authenticated. FIG. 34
[0363] The above protocol was further repeated to assess the stability of the DNA encoding NFTs written in ink on paper after 8 weeks, and the results are shown in FIG. 44 . In this example, 3 replicate writing samples (labeled Replicate #1-3) were assessed at 8 weeks after writing. Comparison of the nucleic acid data packet analysis showed that the DNA samples recovered from the 3 replicate samples were highly similar to each other and were significantly consistent with the nucleic acid data packet analysis results at 2 and 6 weeks. Upon deep sequencing of the 3 replicate writing samples at 8 weeks, each sample was compared to the other two, and the results are shown in FIG. 45 As shown. In the first analysis, replicate 1 was observed to have 8,033 unique data packets, and replicate 2 had 9,965 unique nucleic acid data packets; between replicate 1 and replicate 2, there were 36 nucleic acid data packets in common. Comparing the number of recoveries for each data packet ID, the comparison between replicate 1 and 2 produced a linear trend, R2= 0.954. In the second analysis, replicate 2 had 9,690 unique data packets, and replicate 3 had 10,160; between replicate 2 and replicate 3, there were 311 nucleic acid data packets in common. Comparing the number of recoveries for each data packet ID, the comparison between replicate 2 and 3 produced a linear trend, R2= 0.969. In the third analysis, replicate 1 had 8,045 unique data packets, and replicate 3 had 10,447; between replicate 1 and replicate 3, there were 24 nucleic acid data packets in common. Comparing the number of recoveries for each data packet ID, the comparison between replicate 1 and 3 produced a linear trend, R2= 0.954. These data demonstrate, among other things, that the majority of synonymous nucleic acid data packet sequences are unique sequences between nucleic acid data packet populations of different samples, but that there can still be a small amount of overlap.
[0364] After completing the preliminary evaluation of DNA stability, heat was used to simulate accelerated aging of the DNA samples. In these experiments, 1 / 4 inch paper punches containing 1 μΐ of ink were placed in sealed microfuge tubes. It was estimated that 1 μΐ of ink contained approximately 4 x 1010 8 DNA molecules encoding NFT and 1 x 1010 8ddPCR tracer molecule. Sealed microfuge tubes containing ink-marked paper punches were placed in a 75 °C oven for different lengths of time, then transferred to a 4 °C freezer for storage until analysis. It was estimated that ink-marked paper stored at 75 °C for 0, 1, 2, 3, 4, 5, 6, 7, 8, 9 days would be roughly equivalent to storage at room temperature for about 0, 2.3, 4.6, 6.8, 9.1, 11.4, 13.7, 16.1, 18.3, 20.5 years, respectively. As a control, ink-marked paper was stored at -20 °C throughout the experiment. After 9 days, one sample was transferred from the 75 °C oven to the 4 °C freezer each day, and each sample was analyzed using digital PCR. In this case, the 0thand 1stday samples had roughly the same concentration, the 2ndthrough 6thday samples had steadily decreasing DNA concentrations at the same number of PCR cycles, and the 7ththrough 9thday samples had lower DNA concentrations. This can indicate that DNA degrades over time under 75 °C accelerated aging conditions, but the extent of degradation is not clear. Subsequently, the aged samples were amplified with PCR at different cycle numbers to obtain enough material for sequencing. In this case, the ddPCR tracer molecule in the NFT-encoding DNA added to the ink on the paper punches was used to amplify a 700 bp length of DNA. When the amplified DNA was quantified, it was observed that about 6.5% of the DNA was recovered in the 0thday sample. Subsequently, about 3% of the DNA was recovered in the 1stday (about equivalent to 2.3 years), about 1% of the DNA was recovered in the 2ndday (about equivalent to 4.6 years), and the recovery of DNA from subsequent aged samples decreased in turn. The results are shown in FIG. 46
[0365] PCR analysis of the accelerated aged DNA was continued, and the amplicons flanking the 700 bp DNA fragment targeted by the ddPCR tracer molecule were used to analyze double-stranded DNA breaks in the aged samples. Surprisingly, the DNA exhibited strong resistance to breaks overall at the time points evaluated: the DNA break rate was less than 10% for the 0th, 1std, and 2ndday (about equivalent to 0, 2.3, and 4.6 years) samples, and the DNA break rate was between 10 and 25% for the 3rd, 4th, and 5thday (about equivalent to 6.8, 9.1, and 11.4 years) samples. However, the DNA breaks were more pronounced for the 6ththrough 9thday samples, with a break rate between 40 and 65%. The results are shown in FIG. 47
[0366] Finally, by directly comparing the sequenced DNA samples after accelerated aging, it was observed that the DNA error rate only increased slightly over time, while the sequence efficiency (i.e., the proportion of DNA that was “correct” reads or consensus sequences) decreased over time. The results are shown in FIG. 48 FIG. 49 This is further supported by the sequence length distribution, which shows a shift from a single dominant length to a range of shorter DNA strands over time. Thus, these results show that while DNA can be damaged over time, the error rate of the sequence remains relatively stable, and even after 20 years of equivalent accelerated aging, DNA can still be decoded and the consensus sequence recovered.
[0367] Example 2: DNA encapsulation and extraction in silica microbeads It is known that DNA can be encapsulated in nanoscale silica microbeads, which can be melted into a variety of materials for printing or casting arbitrary shaped objects, and later recovered. See, e.g., Koch J et al. “ A DNA-of-things storage architecture to create materials with embedded memory. ” Nat. Biotechnol. (2020) 38(1): 39-43; U.S. Patent No. 9,850,531 “Molecular code systems”; Bossert et al. “ A hydrofluoric acid-free method to dissolve and quantify silica nanoparticles in aqueous and solid matrices ” Sci. Rep. (2019) 9:7938, the contents of each of which are incorporated herein by reference.
[0368] For example, using the heterogenous DNA cartridge data writing method described in Example 1, a machine readable code is converted into a set of DNA strands. The DNA is encapsulated in silica microbeads, such as silica microspheres, after synthesis, before integration into a material or object.
[0369] Silica seed particles are mixed with a solution of free DNA encoding the NFT, allowing the DNA strands to coat the seed particles. Optionally, the silica seed particles can be modified with amine functional groups to enhance the interaction with the DNA polymer. The DNA-coated seed particles are then mixed with a solution of tetraethoxysilane (TEOS) and base in ethanol, which grows a layer of SiO2 around the DNA, resulting in DNA-encapsulated silica microbeads. More specifically, 5 pL of free DNA (at a concentration of 28 ng / pL) is mixed with 10 pL of silica seed particles (at a concentration of 60 mg / mL) in 500 pL of TE buffer. The resulting mixture is centrifuged at 21,500 g for 1 minute, the supernatant is removed, and the pellet is dispersed in 1 mL of ethanol. To this suspension, 2 pL of APTES, 20 pL of TEOS, and 20 pL of TE buffer are added. The solution is allowed to react overnight with shaking at room temperature, and then centrifuged again, and the pellet is resuspended after washing with ethanol and TE buffer.
[0370] The first extraction protocol was used after encapsulation. In the first extraction protocol, the silica microbeads encapsulating the DNA were dissolved in a buffered oxide etch solution comprising an aqueous mixture of ammonium fluoride and hydrofluoric acid, which can be performed at 0-50 °C, but can be performed readily at room temperature. The microbeads dissolved rapidly in the etch solution within seconds, releasing the original free DNA in a high-salt solution (e.g., F - , NH4 + , SiF6² - ) and it is believed that the relatively high pKa of the hydrofluoric acid avoids damage to the DNA. More specifically, 5 μΐ of silica microbeads encapsulating the DNA were added to 10 μΐ of a buffered oxide etch solution (0.34 g of NH4F and 10 g of 1% HF in TE buffer) and shaken for 1 minute. The mixture changed from turbid to clear and the resulting solution was dialyzed against 10 mL of water for 30 minutes. After dialysis, the free DNA was analyzed by PCR as described in Example 1.
[0371] A second extraction protocol can also be used as an alternative, particularly because the use of hydrofluoric acid is generally undesirable. In this alternative extraction protocol, the etch solution used to dissolve the silica microbeads consists of an aqueous solution of potassium hydroxide. In this case, 10 μg / mL of silica microbeads were mixed with 1 M KOH in an aqueous solution at pH 12 and the silica microbeads dissolved overnight at room temperature. Alternatively, 10 μg / mL of silica microbeads were mixed with 0.1 M KOH in an aqueous solution at pH 12 and the silica microbeads dissolved in 15 minutes under microwave irradiation at 1500 W. After the silica microbeads dissolved and the encapsulated DNA was extracted, the free DNA was dialyzed and analyzed as described above.
[0372] While the above examples implement the disclosure using exemplary steps, materials, items, concentrations, or processes, one skilled in the art will appreciate that alternative steps, materials, items, concentrations, or processes, or other combinations and orders of the steps, materials, items, concentrations, and processes described herein can be used or performed, so long as the same function and / or result described herein is achieved, and fall within the scope of the disclosure. For example, in addition to the above examples, other exemplary implementations have been successfully developed, including the use of the present disclosure in latex paint (free and encapsulated DNA), acrylic paint (free and encapsulated DNA), industrial inkjet printer ink (free DNA), perfume (free DNA), oil painting pigment (encapsulated DNA), permanent marker pen ink (free and encapsulated DNA), stamp pad ink (free DNA), watercolor pigment (free DNA), and 3D printing plastic (encapsulated DNA).
Claims
1. A set of deoxyribonucleic acid (DNA) sequences (e.g., selected from DNA 1 and its subsequent sequences and / or DNA 2 and its subsequent sequences) encoding data that can be used for article authentication and anti-counterfeiting protection, comprising a nucleic acid packet ("nacket"), wherein each nucleic acid packet is encoded by multiple DNA molecules encoding the same data, wherein the sequences of said DNA molecules are heterogeneous.
2. The DNA sequence group according to claim 1, wherein the DNA sequence is prepared by heterogeneous box data writing, wherein two or more box sequences are provided for a single bit or multiple bit combination in the machine-readable code, such that all or almost all DNA molecules in the nucleic acid data packet encode the same data, but the sequences of each molecule exhibit extremely high diversity, wherein the nucleic acid data packet contains multiple heterogeneous boxes.
3. The DNA sequence group according to claim 1 or 2, wherein the data is encoded in n-bit code, where n is greater than 1, for example, binary or ternary code.
4. The DNA sequence group according to any one of the preceding claims, wherein the DNA sequence is prepared from heterologous boxes encoding the same single-bit or multi-bit data, wherein the abundance percentage of different box variants used in the DNA writing process constitutes a unique and distinguishable feature of the DNA.
5. The DNA sequence group according to any of the preceding claims, wherein the data encoded in the DNA is a non-fungible token (NFT).
6. The DNA sequence set according to any one of the preceding claims, wherein the one or more DNA sequences and / or boxes contain one or more topoisomerase recognition sequences, for example, wherein the topoisomerase recognition sequence is 5'-CCCTT-3', 5'-TCCTT-3', 5'-CCCTG-3', or 5'-TGACT-3'.
7. The DNA sequence set according to any one of the preceding claims, wherein the DNA comprises cassettes, each cassette comprising: (i) An information structure field having a sequence of one or more bits corresponding to machine-readable code; and (ii) Topoisomerase recognition sequence; The length of the box is 18 to 25 nucleotides.
8. The DNA sequence group according to any of the preceding claims, wherein the DNA is integrated into or associated with the goods to identify and authenticate the goods.
9. The DNA sequence group according to any of the preceding claims, wherein the DNA is adsorbed, integrated into, or encapsulated in silica microbeads or particles.
10. An article authentication method (e.g., according to method 1 above and any of its subsequent methods), said method comprising: i. Synthesizing a DNA sequence group, such as according to claim 1, wherein the DNA sequence group comprises a nucleic acid packet, wherein each nucleic acid packet contains a plurality of DNA molecules encoding the same data, wherein the sequences of said DNA molecules are heterogeneous; ii. Integrating the DNA sequence into or onto an article; iii. Extracting the DNA sequence from the article; and iv. Analyze the extracted DNA sequence; v. Optionally, compare the analyzed DNA sequence with a DNA sequence database; vi. To verify the authenticity of an item, optionally.
11. The method of claim 10, wherein the DNA sequence is synthesized by sequentially adding a DNA cassette to a DNA acceptor strand; wherein, In each step of the sequential addition process, the box contains a group of heterologous synonymous boxes, such that the box has at least two different sequences, but it encodes the same data in machine-readable code (e.g., binary or ternary code).
12. The method of claim 10 or 11, wherein the boxes are joined together using a ligase.
13. The method of claim 10 or 11, wherein the boxes are joined together using a topoisomerase.
14. The method according to any one of claims 10-13, wherein the DNA sequence comprises a DNA sequence synthesized using transferase-based synthesis and data encoding.
15. The method according to any one of the preceding claims, wherein the nucleic acid data packet is synthesized by sequentially adding cassettes to a DNA receptor strand using an inkjet printhead (e.g., a piezoelectric printhead), wherein each cassette contains a plurality of nucleotides, wherein in each sequential addition step, the cassette contains a heterologous cassette population having at least two different sequences but encoding the same data in machine-readable code (e.g., binary or ternary code), and wherein the cassette is dispensed to at least one write site on a wafer array by an inkjet writing printhead, the printhead or nozzle writing the same code to a plurality of polymer storage strands dispensed to the at least one site; for example, the method includes the following steps: a) Load a starting polymer or DNA onto the target site to be written, with one end of the polymer or DNA attached to the target site; b) Wash the surface of the site; c) Place an inkjet nozzle loaded with a heterogeneous box group over the target site to be written, the heterogeneous box group comprising boxes having at least two different sequences but each of which encodes the same information with one or more bits (e.g., 1 or 0 in binary code, or 00, 01, 10, 11, etc.), the information corresponding to a unique code; d) Causing the inkjet nozzle to release a droplet containing the heterogeneous box community to the site, thereby writing a bit or a portion of the unique code into the DNA or polymer storage string (or strand) associated with the site; and e) Clean the surface of the site. Optionally, the method further includes steps f)–i): f) The inkjet nozzle releases a deblocking / connector reagent droplet to the site; g) Clean the surface of the site; as well as h) Repeat steps c) to g) until the unique code has been written into the storage string of the site; i) Remove the storage string from the site and import it into a collection or storage container for subsequent integration into or onto an item.
16. The method of claim 15, wherein the box is added via a topoisomerase-mediated ligation reaction; for example, by: (i) Reacting a double-stranded acceptor DNA strand with a topoisomerase loaded with a double-stranded DNA cassette, the double-stranded DNA cassette being derived from a population of heterologous cassettes covalently bound to the topoisomerase. The receptor DNA has a 5' overhang on one strand. Each box contains an information sequence, a topoisomerase recognition sequence, and 5' overhangs located on both strands; The 5' overhang of the oligomeric chain ("bottom chain") that does not carry topoisomerase is complementary to the 5' overhang of the receptor DNA, but not complementary to the 5' overhang of the chain ("top chain") that carries topoisomerase in the cassette. The 5' end of the chain carrying the topoisomerase ("top chain") in the cassette and the 5' end of the recipient DNA are both unprotected, for example, unphosphorylated (i.e., 5'-OH); and The topoisomerase loaded with a double-stranded DNA box is delivered to the location of the acceptor strand by a piezoelectric inkjet nozzle. (ii) React the extended recipient DNA from step (i) with a topoisomerase loaded with another double-stranded DNA cassette. The other box contains an information sequence that is the same as or different from any information sequence in the box in step (i), a topoisomerase recognition sequence, and 5' protrusions on both strands; The 5' overhang of the chain in the other cassette that does not carry topoisomerase ("bottom chain") is complementary to the 5' overhang of the extended receptor DNA, but not complementary to the 5' overhang of the chain in the other cassette that carries topoisomerase ("top chain"); and The 5' end of the chain ("top chain") carrying the topoisomerase in the other box is not protected, for example, it is not phosphorylated (i.e., 5'-OH); and (iii) Repeat steps (i) and (ii) until the target nucleotide sequence is obtained; wherein, Optionally, a cleaning step is provided after step (i) and / or after step (ii); and, Optionally, the target nucleotide sequence thus obtained is further reacted with a terminal sequence containing one or more replication primers (e.g., one or more PCR primer sequences).
17. The method according to any of the preceding claims (e.g., according to method 2 above and any subsequent method), wherein the nucleic acid data packet or the box for preparing the nucleic acid data packet contains target code, such as ternary code, using DNA or polymer chains or storage strings, wherein the data is encoded at a series of transitions between non-identical nucleotides, one bit corresponding to one transition, the method comprising: i. Provide a reaction mixture comprising one or more transferases (e.g., terminal deoxynucleotidyl transferase, TdT) and one or more dNTP degrading enzymes (e.g., adenosine triphosphate diphosphatase); ii. Add deoxyribonucleoside triphosphates (dNTPs) to the reaction mixture, for example, the dNTPs are selected from dATP, dCTP, dGTP and dTTP; iii. Wait until the dNTPs in step ii) are added or degraded; iv. Repeat steps ii) and iii) until the target bit sequence is obtained; different dNTP types are used in any two consecutive additions; This yields the DNA molecule population that encodes the target data string.
18. An article authentication method (e.g., according to method 3 above and any of its subsequent methods), which includes: i. Synthesize one or more DNA sequences comprising a nacket, such as the DNA sequence group of claim 1, wherein each nacket comprises multiple DNA molecules encoding the same data, wherein the sequence of said DNA molecules is synthesized using one or more transferases, such as terminal deoxynucleotidyl transferase, such as according to any one of the DNA 2 and subsequent sequences above. ii. Integrating one or more of the DNA sequences into or onto an article; iii. Extracting one or more DNA sequences from the article; and iv. Analyze one or more extracted DNA sequences; v. Optionally, compare one or more DNA sequences being analyzed with a DNA sequence database; vi. To verify the authenticity of an item, optionally.
19. A method for writing resistant digital code using DNA, comprising: i. Receive the target digital code to be written, which is divided into four types of two-bit binary codes to be written (e.g., 00, 01, 10, 11). ii. Provide a predetermined mixture of four predetermined quantities of unique DNA box strings, each mixture corresponding to a different predetermined two-bit binary code value, each mixture having a predetermined proportion of unique DNA boxes within the mixture, and each mixture having unique DNA box strings different from the DNA box strings of the other mixtures; iii. Depositing droplets of a mixture associated with a given two-bit binary code to be written onto a substrate to add a DNA box string to the encoded DNA string being written, the droplets containing a predetermined mixture of unique boxes associated with the given two-bit binary code; and iv. Repeat the deposition steps until the target code is written into the encoded DNA string.
20. The method of claim 19, further comprising: After the target code is written, end caps are added to the encoded DNA string.
21. The method of claim 20, wherein the end cap contains information about the target numeric code or how to read the code.
22. The method according to any one of claims 19-21, wherein the substrate has a receptor DNA strand, one end of which is attached to the substrate and the other end of which is connectable to a unique DNA cassette to be added.
23. The method according to any one of claims 19-22, wherein the predetermined number of unique DNA boxes in one mixture differs from that in at least one other mixture.
24. The method according to any one of claims 19-23, wherein the target digital code is encoded together with the authentication data in the NFT and stored in the blockchain.
25. The method according to any one of claims 19-24, wherein the encoded DNA string is embedded in the physical article to be authenticated.
26. A method for writing resistant digital code using DNA, including: i. Receive the target digital code to be written, the target code being divided into multiple n-bit binary codes to be written, where n is greater than 1; ii. Provide a predetermined mixture of at least two types and a predetermined number of unique DNA box strings, each mixture corresponding to a different predetermined n-bit binary code value, each mixture having a predetermined proportion of unique DNA boxes in the mixture, and the unique DNA box strings of each mixture being different from the DNA box strings in other mixtures; iii. Depositing droplets of a mixture associated with a given n-bit binary code to be written onto a substrate to add a DNA box string to the encoded DNA string being written, the droplets containing a predetermined mixture of unique boxes associated with the given n-bit binary code; and iv. Repeat the deposition steps until the target code is written into the encoded DNA string.
27. The method of claim 26, further comprising: After the target code is written, end caps are added to the encoded DNA string.
28. The method of claim 27, wherein the end cap contains information about the target numeric code or how to read the code.
29. The method according to any one of claims 26-28, wherein the substrate has a receptor DNA strand, one end of which is attached to the substrate and the other end of which is connectable to a unique DNA cassette to be added.
30. The method according to any one of claims 26-29, wherein the predetermined number of unique DNA boxes in one mixture differs from that in at least one other mixture.
31. The method according to any one of claims 26-30, wherein the target digital code is encoded together with the authentication data in the NFT and stored in the blockchain.
32. The method according to any one of claims 26-31, wherein the encoded DNA string is embedded in the physical article to be authenticated.
33. A method of using DNA to write resistant digital code, which includes: i. Receive the target digital code to be written, which is divided into four types of two-bit binary codes to be written (e.g., 00, 01, 10, 11). ii. Provide a set of four unique DNA box strings, each set containing a predetermined number of unique DNA boxes, and each set corresponding to a different predetermined two-bit binary code value, such that each unique box set corresponds to a different two-bit binary code, and each unique box string set is different from the other DNA box strings; iii. Randomly select one from the unique boxes corresponding to the given two-bit binary code to be written, and use it as the selected unique box; iv. Deposit droplets of the selected unique box associated with the given two-bit binary code to be written onto the substrate to add the selected unique box to the encoded DNA string being written; v. Repeat the selection and deposition steps until the target code is written into the coding DNA string at the given writing site on the substrate; as well as vi. Count the number of times each unique box is used when writing each two-bit binary code.
34. Methods for writing resistant digital codes using DNA, including: i. Receive the target digital code to be written, the target digital code being divided into multiple n-bit binary codes to be written, where n is greater than 1; ii. Provide at least two sets of unique DNA box strings, each set containing a predetermined number of unique DNA boxes, and each set corresponding to a different predetermined n-bit binary code value, such that each set of unique boxes corresponds to a different n-bit binary code, and each set of unique box strings is different from other DNA box strings; iii. Randomly select one from the unique boxes corresponding to the given n-bit binary code to be written, and use it as the selected unique box; iv. Deposit droplets of the selected unique boxes associated with the given n-bit binary code to be written onto the substrate to add the selected unique boxes to the encoded DNA string being written; v. Repeat the selection and deposition steps until the target code is written into the coding DNA string at the given writing site on the substrate; as well as vi. Count the number of times each unique box is used when writing each n-bit binary code.
Citation Information
Patent Citations
Methods, compositions, and devices for information storage
US10438662B2
Systems and methods for writing, reading, and controlling data stored in a polymer
US10640822B2
Methods of synthesizing DNA
US11505825B2
Enzymes and systems for synthesizing DNA
US11655465B1
Systems and methods for writing data stored in a polymer using inkjet droplets
US20240308259A1