Methods and compositions

The error-correcting scheme using Reed-Solomon and BCH codes with taboo and wish list sequences addresses security, accuracy, and compliance issues in DNA barcodes, enhancing their reliability and adaptability in synthetic biology.

GB2642439APending Publication Date: 2026-01-14GITLIFE BIOTECH LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
GB2024009820
Authority / Receiving Office
GB · GB
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-05
Publication Date
2026-01-14

AI Technical Summary

Technical Problem

Existing DNA barcode systems face challenges in maintaining security, accuracy, flexibility, compliance with regulatory requirements, and ensuring fair benefit sharing in synthetic biology applications, particularly due to errors, biological constraints, and adversarial attacks.

Method used

A robust error-correcting scheme using Reed-Solomon and BCH codes, combined with taboo and wish list sequences, ensures secure and accurate distribution of DNA barcodes across multiple locations, incorporating redundancy and customization for enhanced security and compliance.

Benefits of technology

The solution enhances the security, accuracy, and flexibility of DNA barcodes, ensuring compliance with regulatory standards and facilitating fair benefit sharing in synthetic biology applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method for generating a distributed polymeric information tag or barcode comprising partitioning, separating, fragmenting, scattering, binning, or splitting information into at least two discrete or
Need to check novelty before this filing date? Find Prior Art

Description

Field The invention is in the field of biotechnology and asset management. Background In the rapidly evolving field of synthetic biology, the insertion of relatively short (< 1KB) DNA snippets, e.g., barcodes or watermarks or information tags, into the genomes of biological assets (bioassets) is gaining traction. For example, barcodes serve two primary purposes: (1) to physically link the biological sample to its digital footprint within, e.g., a specialized version control system or other database that can be centralised or decentralised (e.g. based on distributed ledger technology), and (2) to assert biological and digital rights (intellectual property - IP) over a bioasset. More generally, the integration of digital snippets into biological assets enhances their utility, security, and reliability, supporting the advancement of synthetic biology and biotechnology in a transparent, business and regulatory friendly manner. Barcodes are short, unique sequences of DNA that can be embedded in genetic material to serve as identifiers. They have been used in various areas of life sciences: • Mutant Identification: Barcodes help in identifying and tracking mutants in mixed populations, such as microbial studies and cancer research. • Cell Lineage Tracking: DNA barcodes track cell lineage by progressively editing the barcode sequence using CRISPR technology. • Gene Synthesis: Barcodes are used to isolate and assemble synthetic genes and large pathways in synthetic biology. • Drug Discovery: Barcodes tag and identify chemicals that bind to specific target molecules, aiding in drug discovery processes. Watermarks are sequences embedded in DNA to assert authorship or ownership of genetic constructs. They have been used for: • Authorship Claims: Asserting the origin and ownership of synthetic genetic material. • Infectious Agents: Watermarking DNA sequences of pathogens for tracking and authentication purposes. However, several challenges need to be addressed to ensure the robustness and reliability of this approach. Summary of present invention There are various challenges associated with integrating DNA barcodes into bioassets. For example, errors can arise during DNA sequencing, amplification, synthesis, or natural evolution / mutation. These errors can lead to incorrect barcodes, compromising the integrity and utility of the digital snippets. There are various biological constraints across different bioassets, for example GC content, genome architecture, and codon usage differences. These constraints can affect the stability and function of inserted barcodes. Barcodes can be susceptible to tampering or removal by adversaries, threatening the security and reliability of the system. Finally, it is important that the inserted barcodes do not interfere with the bioasset's normal biological functions is crucial. This includes avoiding detrimental impacts on gene regulatory control and expression, protein function, and overall cellular health. The inventors have designed methods and compositions that address these problems with known barcodes and barcoding strategies. The methods described herein are suitable for generating identifier tags, or barcodes, that can be implemented in any type of polymeric material wherein at least two different versions of the monomer that forms the polymer are available. In some preferred embodiments the polymers are biological polymers that can be encoded in DNA or RNA and be inserted into the genetic material of for example a cell or a plasmid. The types of polymer and monomer are elaborated on elsewhere herein, but it should be appreciated that whilst a focus of the invention is on the application and marking of biological assets such as cells, genetic components and viruses, the polymers that are designed by the methods of the invention may be applied directly to non-biological assets. A feature of the present invention is the design of, and then in some embodiments, the production or and application of a collection of discrete polymers that collectively make up a distributed polymeric information tag, or dPIT. The dPIT is similar to known barcodes, but comprises a number of discrete parts that may be distributed in multiple locations within an asset such as a cell. In some instances one or more of the polymers that form the dPIT are not incorporated into a cell and are instead retained digitally, for example digitally and confidentially, so that the entire dPIT can only be decoded if access to the digital polymer sequence is provided. The polymers of the dPIT may themselves simply be polymers of known sequences that the interested party can retrieve from the asset, e.g., via sequencing, so that a level of identification can be obtained - presence of the sequences indicating that the asset is what you expect. In other instances, the polymers of the dPIT are linked to digitally stored information. For example, in some embodiments, information - such as binary information that details certain features of the asset, or ownership or licencing terms - is encoded in the polymers of the dPIT, such that upon decoding ( / .e., identifying the polymers in the asset and reading the polymer sequence) the information is obtained. A key feature of the invention is the partitioning of the total sum information, or otherwise producing discrete polymer sequences that are for distributed application to an asset with or without retention of at least one of the polymer sequences in a separate asset or digitally. There are various ways in which the collection of polymers that comprise the dPIT can be generated, for example via the portioning of information such as binary information into at least two information subsets, with subsequent encoding into a polymeric sequence; encoding information such as binary information into a polymeric sequence followed by partitioning of the polymeric sequence; or simply generating at least two polymeric sequences via any means. Figure 1 shows various means of putting the invention into effect, and is shown for exemplary purposes only as there are various other permutations of the various method steps and features that are encompassed by the invention. Accordingly, the invention provides a method of generating a distributed polymeric information tag (dPIT) comprising a plurality of sub-informative polymer sequences (sub-IPS) said method comprising: a) partitioning information into at least two discrete information subsets and encoding each discrete information subset in a plurality of information symbols, wherein the plurality of information symbols represents a plurality of information monomers, to form the at least two discrete sub-informative polymer sequences (sub-IPS); b) providing an informative polymer sequence (IPS), wherein the informative polymer sequence comprises a plurality of information symbols representing a plurality of information monomers and partitioning the IPS into at least two discrete sub-IPSs sequences; and / or c) providing a plurality of discrete sub-informative polymer sequences (sub-IPS) wherein each discrete sub-IPS comprises a plurality of information symbols, wherein the plurality of information symbols represents a plurality of information monomers. By generating we also include the mean of designing, since the steps of arriving at the various polymeric sequences can all be done in silico or via computer implementation - i.e., it is not necessary to generate physical polymers of the various sequences during the design process. The skilled person will appreciate that the use of the term "symbol" is intended to reflect the computational nature of the design process, and a symbol simply represents a particular monomer, e.g., a U may represent uracil. The information symbols are the symbols (or monomers) that actually provide the information that the asset is to be marked with. In some embodiments described below the polymer sequences also include redundant symbols (or monomers). By discrete we include the meaning of "separate", i.e., the sub-informative polymeric sequences are separate individual sequences (which correlate to separate individual monomers when produced physically). Each discrete polymer can comprise additional monomer elements, e.g., monomers that are nucleotides that are homologous to a target site for insertion via homologous recombination. As discussed above, in some embodiments binary information is encoded in the polymeric sequences. Accordingly in some embodiments: a) the information that is partitioned in (a) is binary information; and / or b) the informative polymer sequence of (b) was generated by encoding information in a plurality of information symbols to generate the informative polymer sequence, optionally wherein the information that is encoded is binary information. It will be appreciated that the discrete sub-informative polymer sequences may be of any length. In some applications of the invention there is no length restriction and so any length of polymer sequence is appropriate. Similarly, any number of discrete sub-informative polymer sequences is possible, and in some instances, there are no restrictions on the number of sequences that can be accommodated in an asset. A long piece of information or informative polymer sequence (IPS) may be partitioned into a few larger polymeric sequences, or multiple shorter polymeric sequences. The skilled person is well able to decide on the appropriate design choices. However, in some applications there are restraints on the length and / or number of polymeric sequences that could be accommodated within an asset. The skilled person will understand that when aiming to mark a biological asset with a polymer, there are certain restrictions and intolerances. For example, trying to mark a plasmid with a long polymeric sequence is unlikely to be successful; and the incorporation of a polymer into an essential gene of a cell is likewise something to avoid. Various such design features will be well apparent to the skilled person. Accordingly, in some embodiments where the dPIT is for incorporation into a biological asset, the method may comprise the step of obtaining a partition spectrum for said biological asset wherein said partition profile describes the frequency of available positions for insertion of discrete polymer fragments of different lengths into one or more corresponding polymers of said biological asset. For example, a partition spectrum may detail how many positions there in a genome that would tolerate a polymer insertion of 100 nucleotides. An exemplary partition spectrum is shown in Figure 2. In some instances the partition spectrum may describe the frequency of available positions for insertion of discrete polymer fragments of between: 0-49, 50-99, 100-149, 150-199, 200-249, 250-299, 300-349, 350-399, 400-449, 450-499, 500-549, 550-599, 600-649, 650-699, 700-749, 750-799, 800-849, 850-899, 900-949, 950-999, 1000-1049, 1050-1099, 1100-1149, 1150-1199, 1200-1249, 1250-1299, 1300-1349, 1350-1399, 1400-1449, 1450-1499, 1500-1549, 1550-1599, 1600-1649,1650-1699,1700-1749, 1750-1799, 1800-1849, 1850-1899, 1900-1949, 1950-1999, 2000-2049, 2050-2099, 2100-2149, 2150-2199, 2200-2249, 2250-2299, 2300-2349, 2350-2399,2400-2449, 2450-2499, 2500-2549, 2550-2599, 2600-2649, 2650-2699, 2700-2749, 2750-2799, 2800-2849, 2850-2899, 2900-2949, 2950-2999, 3000-3049, 3050-3099,3100-3149, 3150-3199, 3200-3249, 3250-3299, 3300-3349, 3350-3399, 3400-3449, 3450-3499, 3500-3549, 3550-3599, 3600-3649, 3650-3699, 3700-3749, 3750-3799, 3800-3849, 3850-3899, 3900-3949, 3950-3999, 4000-4049, 4050-4099, 4100-4149,4150-4199, 4200-4249, 4250-4299, 4300-4349, 4350-4399, 4400-4449, 4450- 4499,4500-4549, 4550-4599, 4600-4649, 4650-4699, 4700-4749, 4750-4799, 4800-4849,4850-4899, 4900-4949, 4950-4999, 5000-5049, 5050-5099, 5100-5149, 5150-5199, 5200-5249, 5250-5299, 5300-5349, 5350-5399, 5400-5449, 5450-5499, 5500-5549, 5550-5599, 5600-5649 monomers in length. The bin size of such a partition spectrum is a parameter that can be adjusted by a skilled person. This allows the skilled person to determine the appropriate length and number of separate polymeric sequences that the dPIT should comprise (also taking into consideration that not all of the polymers that comprise the dPIT have to be inserted into the asset). Accordingly in some embodiments: a) where the method comprises partitioning information into at least two discrete information subsets, the length of and number of discrete information subsets is chosen based on a partition spectrum; b) where the method comprises providing an informative polymer sequence (IPS) and partitioning the IPS into at least two discrete sub-IPSs sequences, the length of and number of discrete sub-IPSs is chosen based on a partition spectrum; and / or c) where the method comprises providing a plurality of discrete sub-informative polymer sequences (sub-IPS), the length of and number of discrete sub-IPS is chosen based on a partition spectrum such that polymeric fragments corresponding to said discrete sub-IPS would be accommodated within the biological asset. Where the asset is a biological asset, it may be any biological asset. In preferred embodiments the biological asset has a polymer of the type of the dPIT and into which the polymers of the dPIT can be incorporated. In some embodiments the asset biological asset and is selected from the group comprising or consisting of: a) a cell or one or more chromosomes of a cell, optionally: i) a prokaryotic cell, optionally a bacterial cell; ii) a eukaryotic cell, optionally a yeast cell, a fungal cell, a mammalian cell, a plant cell; b) A circular double-stranded DNA structure, optionally a plasmid; c) A circular single-stranded DNA structure; d) A nanostructure, wherein the nanostructure comprises at least one DNA or RNA or protein or peptide molecule, optionally wherein the at least one DNA molecule is single-stranded or double-stranded and / or wherein the at least one RIMA molecule is single-stranded or double-stranded; e) a phagemid; f) a phage; g) a virus, optionally a viral vector, optionally an AAV or lentiviral vector; h) a population of cells; and / or i) a nanoparticle, for example a lipid nanoparticle, vesicle, micelle, etc, for example a nanoparticle that is for use as a vaccine. As stated elsewhere herein, the asset may also be a non-biological asset. In these instances the asset will typically comprise no suitable polymer into which the dPIT can be introduced. However, one or more of the polymers that comprise the dPIT can be otherwise applied to a non-biological asset. An example of such an approach that has been commercially exploited is the marking of personal property using DNA. However, these approaches do not use polymers that comprise a dPIT as described herein. As set out above, the method of the invention is a method of designing or generating sequence information that comprises the dPIT. Accordingly, in some embodiments said generating and / or partitioning and / or providing is performed using computational processes and said sequences are digital sequences. In additional embodiments the method does comprise the additional steps of producing the physical polymers, also set out elsewhere herein. An advantage of the present invention is that allows for error correction of polymeric sequences that may have diverged from the original sequence used to mark the asset. For example, where the asset is a bioasset, nucleic acids will overtime incorporate various mutations, such as substitutions and deletions. Where such mutations occur in one of the distributed polymers used to mark the asset, upon sequence retrieval and decoding the method allows for the correction of these errors (digitally) and the original sequence to be identified. The method of the invention accomplishes this by encoding into the collections of polymeric sequences that make up the dPIT a number of redundant symbols or monomers. These redundant symbols or monomers are used for error correction of the informative symbols or monomers. Accordingly, the method of generating or designing the dPIT includes methods of generating and including the redundant symbols or monomers. For example, in some embodiments of the method of generating the dPIT, the method comprises generating a plurality of redundant symbols by processing the information symbols present in the at least two discrete sub-informative polymer sequences (sub- IPS) with an error correction coding process. In some embodiments the method further comprises combining the plurality of redundant symbols with the informative symbols, optionally wherein at least one or more or all of the at least two discrete sub-informative polymer sequences (sub-IPS) comprises redundant symbols to form at least one or more error-correcting discrete sub-IPS. There are various different types of error correction that can be applied, singly or in combination, to the sub-informative polymer sequences. For example, in some embodiments the error correction coding process is an intra-partition error correcting code process (for providing error correcting capability within a sub-IPS (or partition), and is applied to the at least two discrete sub-informative polymer sequences (sub-IPS) to generate for each discrete sub-informative polymer sequences (sub-IPS) an additional sequence that is a redundant sequence comprising a plurality of redundant symbols that represent a plurality of redundant monomers for use in error correction of the information monomers. The redundant sequence comprising the plurality of redundant symbols can be combined with the corresponding discrete sub-informative polymer sequences (sub-IPS) to form error-correcting discrete sub-IPSs. It is also possible to conduct an error correction coding process between the separate discrete sub-informative polymer sequences (sub-IPS) and / or error-correcting discrete sub-IPSs. This inter-partition error correction coding process generates entire new partition sequences, which can be used in combination with the discrete sub-informative polymer sequences (sub-IPS) and / or error-correcting discrete sub-IPSs for error correction. Accordingly in the same or different embodiments the method comprises the step of applying an inter-partition error correction coding process to the at least two discrete sub-informative polymer sequences (sub-IPS) or to the at least one of the error-correcting discrete sub-IPSs to generate at least one discrete redundant polymer sequence comprising a plurality of monomers that are redundant. The redundant polymer sequences have no informative symbols or monomers. The intra- and inter-partition error correction coding process can be any error correction coding process. For example, the intra-partition error correction coding process may comprise a Bose-Chaudhuri-Hocquenghem (BCH) error correction coding process. A BCH process can provide an efficient error correction process for binary data (as may be encoded within a sub-IPS) and has good performance for handling random errors. In some examples, the inter-partition error correction process may comprise a Reed Solomon coding process. Reed Solomon error correction can provide for highly efficient burst error correction (as may be found between sub-IPS) and is particularly robust in noisy environments. Table 1 describes a number of exemplary error correction coding processes, along with their various properties, advantages, disadvantages and known use cases. The table lists alternatives to BCH error correction coding processes that may be particularly suitable for intra-partition 5 error correction coding and a list of alternatives to Reed Solomon error correction that may be particularly suitable for inter-partition error correction coding. Table 1 Aspect BCH Codes Alternatives to BCH ECC RS Codes Alternatives to RS ECC Code Structure and Basis Cyclic code over binary fields (GF(2)) Hamming Codes: Linear block codes over GF(2) Non-binary cyclic code over finite fields (GF(q), typically 2m) Turbo Codes: Parallel concatenated convolutional codes LDPC Codes: Linear block codes over GF(2) LDPC Codes: Linear block codes over GF(2m) Convolutional Codes: Codes generated by convolution of input bits Polar Codes: Uses channel polarization to achieve capacity Error Correction Capability Corrects multiple random bit errors Corrects singlebit errors and can detect two-bit errors Corrects both random and burst errors Corrects random errors with near Shannon-limit performance Corrects random and burst errors with sparse matrices Corrects random errors, close to capacity achieving Corrects random errors using trellis structure Corrects errors using polarization technique Field of Operation Binary fields Binary fields Larger fields (GF(2m)) Binary or larger fields depending on implementation Binary fields Binary or larger fields depending on implementation Binary fields Binary fields Complexity Generally lower complexity for encoding and decoding Very low complexity, simple to implement Higher complexity due to operations over larger fields Higher complexity, iterative decoding process Moderate complexity, efficient Iterative decoding Moderate complexity, efficient decoding Moderate complexity, efficient Moderate complexity, sequential decoding efficient iterative decoding Best Use Cases Digital Communication Systems: Satellite, mobile communication Digital systems with low error rates, such as RAM Optical and Magnetic Storage: CDs, DVDs, Blu-ray discs, hard drives Deep space communication, wireless systems, data storage Data Storage: Solid-state drives (SSDs), memory devices Modern communication systems, wireless communications Data Transmission: DSL, Ethernet, wireless communication Data Transmission: High-speed communication systems QR Codes: Error correction for readable damaged codes Data storage, optical and magnetic storage Digital Broadcasting: Television and radio broadcasting Broadcast systems, high-definition video, and multimedia applications Simple error correction in low-cost systems QR Codes: Robust error correction for larger data regions Modern communication systems requiring high reliability and efficiency Deep space communication, data transmission Wireless communications, 5G, and beyond Advantages Efficient for binary data, good for random errors Simple and easy to implement, fast encoding and decoding Highly efficient for burst error correction, robust in noisy environments Near Shannonlimit performance, highly efficient for random errors Highly efficient, good for sparse error patterns Highly efficient, close to capacity achieving Efficient for sequential data processing Efficient decoding, especially for long block lengths Disadvantages Limited to random errors, higher complexity than simpler codes Limited to single-bit error correction, less efficient for multiple errors More complex and computationally intensive, larger overhead More complex decoding, higher latency due to iterative processes Complexity in encoding and decoding for large block lengths Complexity in encoding and decoding, especially for shorter block lengths Performance depends on code length and sparsity Requires careful design for optimal performance For example any one or more of the intra-partition error correction coding process and / or the inter-partition error correction coding process may be selected from any one or more of: Bose-Chaudhuri-Hocquenghem; Hamming Codes: Linear block codes over GF(2); LDPC Codes: Linear block codes over GF(2); and or Convolutional Codes: Codes generated by convolution of input bits; Reed-Solomon codes; non-binary cyclic code over finite fields (GF(q), typically 2m). In some preferred embodiments the intra-partition error correction coding process may be selected from any of a Bose-Chaudhuri-Hocquenghem error correction coding process; Hamming Codes: Linear block codes over GF(2); LDPC Codes: Linear block codes over GF(2); and or Convolutional Codes: Codes generated by convolution of input bits. In some preferred embodiments the inter-partition error correction coding process comprises a Reed-Solomon error correction coding process or non-binary cyclic code over finite fields (GF(q), typically 2m). In preferred embodiments the intra-partition error correction coding process is a Bose-Chaudhuri-Hocquenghem error correction coding process and the inter-partition error correction coding process is a Reed-Solomon error correction coding process. The skilled person will appreciate that particularly where the asset is a bioasset and the polymers are to be inserted into the bioasset, there are certain characteristics of the polymerthat should be avoided and certain characteristics that should be designed-in. For example, where the polymer is DNA, the skilled person will appreciate that introducing a DNA polymer to the asset that has commonly used restriction endonuclease sites, or strong secondary structures, is to be avoided. Accordingly, at, before, or after the intra and / or inter error correction coding processes, one or more monomers of each of the polymeric sequences may be modified, for example by adding new monomers, in such a way so as to impart desirable characteristics or disrupt undesirable characteristics. Accordingly in some embodiments the method may comprise (at any point) the introduction of one or more characteristic symbols into one or more of the at least two discrete sub-informative polymer sequences (sub-IPS) and / or to at least one of the error-correcting discrete sub-IPSs, and / or to the at least one discrete redundant polymer sequences so as to: a) impart desirable sequence characteristics on the at least one or more of the at least two discrete part-informative biological polymer sequence; and / or b) remove undesirable sequence characteristics from the at least one or more of the at least two discrete part-informative biological polymer sequence. The addition / disruption of characteristics is applicable to any polymer for which the invention is suitable. Exemplary polymers are described elsewhere herein. For example, in the context of a biological asset where the polymers are DNA polymers, the desirable sequence characteristics may be characteristics selected from: characteristics that camouflage the discrete part-informative sequence with respect to the target site in the one or more target biological polymers, balanced GC content, a secondary structure that is a hairpin, a secondary structure that is a kissing loop, or a sequence that supports genetic engineering; and / or the undesirable sequence characteristics may be selected from: characteristics that are readily distinguishable from the surrounding sequence of the target site in the one or more target biological polymers, endonuclease recognition sites such as restriction endonuclease sites, greater than X% GC content, sequences prone to secondary structure formation, a sequence that may interfere with nearby gene expression. The collection of the various polymeric sequences collectively make up the distributed polymeric information tag (dPIT). For example in some embodiments: a) the at least one error-correcting discrete sub-IPS and the at least one discrete redundant polymer sequence together collectively form the distributed polymeric information tag (dPIT); b) the at least two discrete sub-informative polymer sequences (sub-IPS) together collectively form the distributed polymeric information tag (dPIT); c) the error-correcting discrete sub-IPSs together collectively form the distributed polymeric information tag (dPIT); or d) the at least two discrete sub-informative polymer sequences (sub-IPS) and the at least one discrete redundant polymer sequence together collectively form the distributed polymeric information tag (dPIT). As set out elsewhere, the polymer may be any type of polymer, providing that there are least two versions of the relevant monomers available to encode the relevant information. For example some exemplary polymers suitable for use with the invention are: Deoxyribonucleic acid (DNA), Ribonucleic acid (RNA), Protein or peptide or polypeptide, Peptide Nucleic Acids (PNAs), Polypeptides, Polysaccharides, Polyethylene Glycol (PEG), Polylactic Acid (PLA), Poly(ethylene terephthalate) (PET), Poly(methyl methacrylate) (PMMA), Polyaniline (PANI), Polycarbonate (PC), Polystyrene (PS), Poly(vinyl alcohol) (PVA), Polylactide-co-glycolide (PLGA), Polydimethylsiloxane (PDMS), Polyurethanes, Polycaprolactone (PCL), Polyhydroxyalkanoates (PHAs). The skilled person will appreciate which monomers are used in which polymers, however, for example: Where the polymer is DNA the monomer is selected from the group comprising or consisting of: Adenine (A), Thymine (T), Cytosine (C), Guanine (G), or a non-natural nucleotide, optionally wherein the nucleotide is a modified nucleotide; Where the polymer is RIMA the monomer is selected from the group comprising or consisting of: Adenine (A), Uracil (U), Cytosine (C), Guanine (G), or a non-natural nucleotide, optionally wherein the nucleotide is a modified nucleotide; Wherein the polymer is a protein or peptide or polypeptide the monomer is selected from the group comprising or consisting of: an Amino acids optionally alanine (A), arginine (R), asparagine (N), aspartic acid (D), cysteine (C), glutamic acid (E), glutamine (Q), glycine (G), histidine (H), isoleucine (I), leucine (L), lysine (K), methionine (M), phenylalanine (F), proline (P), serine (S), threonine (T), tryptophan (W), tyrosine (Y), and valine (V), or a non-natural amino acid, optionally wherein the amino acid is a modified amino acid. Where the polymer is a PNA the monomer is selected from the group comprising or consisting of: Adenine (A), Thymine (T), Cytosine (C), Guanine (G) or a non-natural nucleotide, optionally wherein the nucleotide is a modified nucleotide; Where the polymer is a polysaccharide the monomer is selected from the group comprising or consisting of: Glucose, Fructose, Galactose, Mannose; Where the polymer is PEG the monomer is selected from the group comprising or consisting of: Ethylene glycol, Propylene glycol, Butylene glycol, Diethylene glycol; Where the polymer is PLA the monomer is selected from the group comprising or consisting of: L-lactic acid, D-lactic acid, Mesolactic acid, 2-Hydroxypropanoic acid; Where the polymer is PET the monomer is selected from the group comprising or consisting of: Ethylene glycol, Terephthalic acid, Isophthalic acid, Naphthalene dicarboxylate; Where the polymer is PMMA the monomer is selected from the group comprising or consisting of: Methyl methacrylate, Ethyl methacrylate, Butyl methacrylate, 2-Hydroxyethyl methacrylate; Where the polymer is PANI the monomer is selected from the group comprising or consisting of: Aniline, o-Toluidine, m-Toluidine, p-Toluidine; Where the polymer is PC the monomer is selected from the group comprising or consisting of: Bisphenol A, Bisphenol S, Cyclohexane dimethanol, Terephthalic acid; Where the polymer is PS the monomer is selected from the group comprising or consisting of: Styrene, a-Methylstyrene, Butadiene, Acrylonitrile; Where the polymer is PVA the monomer is selected from the group comprising or consisting of: Vinyl acetate, Vinyl alcohol, Ethylene, Acrylamide; Where the polymer is PLGA the monomer is selected from the group comprising or consisting of: L-lactic acid, D-lactic acid, Glycolic acid, 2-Hydroxypropanoic acid; Where the polymer is PDMS the monomer is selected from the group comprising or consisting of: Dimethylsiloxane, Diphenylsiloxane, Methylphenylsiloxane, Vinylsiloxane; Where the polymer is polyurethane the monomer is selected from the group comprising or consisting of: Toluene diisocyanate, Methylene diphenyl diisocyanate, Hexamethylene diisocyanate, Polycaprolactone polyol; Where the polymer is PCL the monomer is selected from the group comprising or consisting of: E-Caprolactone, 5-Valerolactone, y-Butyrolactone, L-Lactide; Where the polymer is PHA the monomer is selected from the group comprising or consisting of: 3-Hydroxybutyrate, 3-Hydroxyvalerate, 3-Hydroxyhexanoate, 4-Hydroxybutyrate. Any of the above monomers are non-exhaustive examples and may be a derivative or analogue thereof. Exemplary polymers, the relevant chemistry and exemplary means of retrieving sequence data are shown in Table 2. Table 2 Polymer Monomers Chemistry Sequencing Process Deoxyribon ucleic acid (DNA) Adenine (A), Thymine (T), Cytosine (C), Guanine (G), or a non-natural nucleotide, optionally wherein the nucleotide is a modified nucleotide Phosphodiester bond Nanopore, Hybridization assays, Mass spectrometry Ribonucleic acid (RNA) Adenine (A), Uracil (U), Cytosine (C), Guanine (G), or a non-natural nucleotide, optionally wherein the nucleotide is a modified nucleotide Phosphodiester bond Protein or peptide or polypeptide Amino acids - alanine (A), arginine (R), asparagine (N), aspartic acid (D), cysteine (C), Peptide bond Polymer Monomers Chemistry Sequencing Process glutamic acid (E), glutamine (Q), glycine (G), histidine (H), isoleucine (I), leucine (L), lysine (K), methionine (M), phenylalanine (F), proline (P), serine (S), threonine (T), tryptophan (W), tyrosine (Y), and valine (V), or a nonnatural amino acid, optionally wherein the amino acid is a modified amino acid. Peptide Nucleic Acids (PNAs) Adenine (A), Thymine (T), Cytosine (C), Guanine (G), or a non-natural nucleotide, optionally wherein the nucleotide is a modified nucleotide Peptide bond formation Nanopore, Hybridization assays, Mass spectrometry Polypeptide s Alanine (A), Glycine (G), Valine (V), Leucine (L) Peptide bond formation Nanopore, Mass spectrometry, Edman degradation Polysaccha rides Glucose, Fructose, Galactose, Mannose Glycosidic bond formation Nanopore, Mass spectrometry, NMR spectroscopy Polyethylen e Glycol (PEG) Ethylene glycol, Propylene glycol, Butylene glycol, Diethylene glycol Polymerization, click-chemistry Nanopore, Mass spectrometry, NMR spectroscopy Polylactic Acid (PLA) L-lactic acid, D-lactic acid, Mesolactic acid, 2- Hydroxypropanoic acid Polymerization Nanopore, Mass spectrometry Poly(ethyle ne terephthala te) (PET) Ethylene glycol, Terephthalic acid, Isophthalic acid, Naphthalene dicarboxylate Polycondensatio n Nanopore, Mass spectrometry, NMR spectroscopy Poly(methy 1 methacryla te) (PMMA) Methyl methacrylate, Ethyl methacrylate, Butyl methacrylate, 2-Hydroxyethyl methacrylate Radical polymerization Nanopore, Mass spectrometry, NMR spectroscopy Polyaniline (PANI) Aniline, o-Toluidine, m- Toluidine, p-Toluidine Oxidative polymerization Nanopore, Mass spectrometry, Polymer Monomers Chemistry Sequencing Process NMR spectroscopy Polycarbon ate (PC) Bisphenol A, Bisphenol S, Cyclohexane dimethanol, Terephthalic acid Polycondensatio n Nanopore, Mass spectrometry, NMR spectroscopy Polystyrene (PS) Styrene, a-Methylstyrene, Butadiene, Acrylonitrile Radical polymerization Nanopore, Mass spectrometry, NMR spectroscopy Po / y( vinyl alcohol) (PVA) Vinyl acetate, Vinyl alcohol, Ethylene, Acrylamide Polymerization Nanopore, Mass spectrometry, NMR spectroscopy Polylactidecoglycolide (PLGA) L-lactic acid, D-lactic acid, Glycolic acid, 2- Hydroxypropanoic acid Polymerization Nanopore, Mass spectrometry Polydimeth ylsiloxane (PDMS) Dimethylsiloxane, Diphenylsiloxane, Methylphenylsiloxane, Vinylsiloxane Polymerization Nanopore, Mass spectrometry, NMR spectroscopy Polyuretha nes Toluene diisocyanate, Methylene diphenyl diisocyanate, Hexamethylene diisocyanate, Polycaprolactone polyol Polyaddition, polycondensatio n Nanopore, Mass spectrometry, NMR spectroscopy Polycaprola ctone (PCL) E-Caprolactone, 5- Valerolactone, y- Butyrolactone, L-Lactide Ring-opening polymerization Nanopore, Mass spectrometry Polyhydrox yalkanoate s (PHAs) 3-Hydroxybutyrate, 3- Hydroxyvalerate, 3- Hydroxyhexanoate, 4- Hydroxybutyrate Polymerization Nanopore, Mass spectrometry The skilled person will appreciate that some of the relevant polymers may take various different forms and these various different forms are contemplated by the present invention. For example, where the polymer is DNA, the DNA may be double-stranded; 5 or may be single-stranded; or may be linear; or may be circular; and where the polymer is RNA the RIMA may be double-stranded; or may be single-stranded; or may be linear; or may be circular. The skilled person will appreciate that where the polymer is an amino acid polymer 10 such as a protein or peptide, the amino acid sequence of the polymer can be encoded in DNA or RNA. The skilled person is well able to generate a suitable DNA or RNA sequence that encodes a specific protein sequence. Accordingly in some embodiments wherein where the at least two discrete sub-informative polymer sequences (sub-IPS) and / or the at least one of the error-correcting discrete sub-IPSs, and / or the at least one discrete redundant polymer sequences are protein sequences, the step of preparing a DNA or RNA sequence that encodes said at least two discrete part-informative biological polymer sequences and / or the at least two error-correcting part-informative biological polymer sequences and / or the at least one discrete redundant biological polymer sequence. It will be appreciated that much of the method of designing the necessary polymer sequences can be done in silico and by the use of a computer for example. However, the practically use the designed polymers, they must be generated in a physical form. Methods of making the relevant polymers are known, for example methods of DNA, RNA and protein synthesis are well known. In some embodiments then the method comprises producing one or more polymers corresponding to the sequences set out in the at least two discrete sub-informative polymer sequences (sub-IPS) and / or the at least one of the error-correcting discrete sub-IPSs, and / or the at least one discrete redundant polymer sequences so as to produce one or more or all of: a) at least one error-correcting discrete sub-IPS polymer corresponding to the at least one error-correcting discrete sub-IPS sequence; b) at least one discrete redundant polymer corresponding to the at least one discrete redundant polymer sequence; c) at least one sub-informative polymer corresponding to at least one the discrete sub-informative polymer sequence to produce a collection of polymeric information tag polymers. In some embodiments the collection of polymeric information tag polymers comprises at least one errorcorrecting discrete sub-IPS polymer and at least one discrete redundant polymer. The methods described herein may also include a physical step of introducing the polymers of the dPIT to the asset. For example, in some embodiments the method comprises the step of applying, inserting or otherwise causing at least one of the polymers of the collection of the polymeric information tag polymers to be physically located in or on one or more assets. As set out elsewhere, the asset may be a biological asset and in preferred instances the biological asset comprises one or more target polymers corresponding to the type of polymer of the collection of polymeric information tag polymers. However in other embodiments the asset is not a biological asset and does not comprise one or more target polymers. However, the invention also encompasses methods of introducing or applying the polymers of the dPIT to non-biological assets. Preferences for the target polymers are described elsewhere herein and include: a) one or more chromosomes of a cell (i.e. different polymers of the collection may target different chromosomes within the same cell), optionally: i) a prokaryotic cell, optionally a bacterial cell; ii) a eukaryotic cell, optionally a yeast cell, a fungal cell, a mammalian cell, a plant cell; b) A circular double-stranded DNA structure, optionally a plasmid c) A circular single-stranded DNA structure d) A nanostructure, wherein the nanostructure comprises at least one DNA or RNA or protein or peptide molecule, optionally wherein the at least one DNA molecule is single-stranded or double-stranded and / or wherein the at least one RNA molecule is single-stranded or double-stranded; e) a phagemid f) a phage g) a virus, optionally a viral vector, optionally an AAV or lentiviral vector; h) one or more chromosomes or plasmids of separate cells in a cell culture; i) a nanoparticle, for example a lipid nanoparticle, vesicle, micelle, for example a nanoparticle that is for use as a vaccine. In some embodiments the method comprises the step of inserting or otherwise causing more than one polymer of the collection of polymeric information tag polymers to be physically located in one or more of the target biological polymers, for example to be physically located in: a) The genome of the same cell b) The same circular double-stranded DNA structure, optionally a plasmid c) The same circular single-stranded DNA structure d) The same nanostructure, wherein the nanostructure comprises at least one DNA or RNA or amino acid polymer molecule, optionally wherein the at least one DNA molecule is single-stranded or double-stranded and / or wherein the at least one RNA molecule is single-stranded or double-stranded; e) The same phagemid f) The same phage g) The same virus, optionally a viral vector, optionally an AAV or lentiviral vector h) members of a cell population, optionally one or more chromosomes or plasmids of separate cells in a cell population; i) the same nanoparticle, for example a lipid nanoparticle, vesicle, micelle etc. In some embodiments all of the polymers of the collection of polymeric information tag polymers are physically located in one or more of the target biological polymers of the asset (where the asset if a population of cells we include the meaning whereby collectively across the population of cells all of the polymers of the collection are present). In some embodiments at least one of the polymers of the collection of polymeric information tag polymers is not embedded in the same asset, for example is not embedded in the same one or more target biological polymers as the rest of the polymers of the collection of polymeric information tag polymers. This embodiment allows part of the information tag to be withheld and maintained confidentially, such that only when all pieces of the tag, i.e. sequences of each of the polymers of the collection are obtained and decoded together can the initial information be retrieved. In some embodiments then at least one of the: a) at least one error-correcting discrete sub-IPS polymer sequences; b) at least one discrete redundant polymer sequence; c) at least one sub-informative polymer sequence; remains as a digital sequence and a corresponding polymer is not produced. The invention also provides a computer implemented method for generating a distributed polymeric information tag (dPIT) according to the invention for incorporation into or application to one or more assets, wherein the method comprises the method of generating the dPIT of the invention. The invention also provides a computer implemented method for generating a distributed polymeric information tag (dPIT)for incorporation into or application to one or more assets, the method comprising: receiving an informative polymer sequence (IPS), wherein the informative polymer sequence comprises a plurality of information symbols representing a plurality of information monomers; partitioning the informative polymer sequence (IPS)to generate at least two discrete sub-IPSs sequences each comprising a plurality of information symbols; and generating a plurality of redundant symbols by processing the plurality of information symbols of the at least two discrete sub-IPSs sequences with an error correction coding process; combining the redundant symbols with the at least two discrete sub-IPSs sequences to generate at least one error-correcting discrete sub-IPS polymer sequences that each comprises the plurality of information symbols of the discrete sub-IPSs sequences and a plurality of redundant symbols. In some embodiments of the computer implemented method, generating the plurality of redundant symbols comprises: for each of the at least two discrete sub-IPSs sequences, applying an intra-partition error correction coding process to the plurality of information symbols within the discrete sub-IPSs sequences to generate a plurality of partition redundant symbols; and including the partition redundant symbols in the discrete sub-IPSs sequences. In some embodiments of the computer implemented method the intra-partition error correction coding process comprises a Bose-Chaudhuri-Hocquenghem error correction coding process. In some embodiments, generating the plurality of redundant symbols comprises: applying an inter-partition error correction coding process to the plurality of information symbols within each of the at least two discrete sub-IPSs sequences to generate one or more discrete redundant polymer sequences comprising a respective plurality of redundant symbols. In some embodiments the inter-partition error correction coding process comprises a Reed-Solomon error correction coding process. The invention also provides a digital distributed polymeric information tag (dPIT) produced according to any of the methods of the invention described herein. In addition to providing methods of generating the dPIT, and digital versions of the dPIT, the invention also provides the physical polymers and collection of polymers that make up the dPIT. Accordingly the invention provides a physical distributed polymeric information tag (dPIT) comprising the collection of polymeric information tag polymers of the invention and as designed and generated by the methods of the invention. The invention provides a physical distributed polymeric information tag (dPIT) comprising a collection of polymeric information tag polymers, wherein the collection of polymeric information tag polymers comprises: a) at least one error-correcting discrete sub-IPS polymer produced according to the method of the invention and at least one discrete redundant polymer produced according to any method provided herein; b) at least one error-correcting discrete sub-IPS polymer produced according to the method of the invention; c) at least one discrete redundant polymer produced according to the method of the invention; and / or d) at least one sub-informative polymer. In some embodiments the physical distributed polymeric information tag (dPIT)also comprises a non-physical digital polymer, for example that that is stored on a data storage facility. The invention also provides method of producing one or more polymers corresponding to the sequences set out in any of the at least two discrete sub-informative polymer sequences (sub-IPS) and / or the at least one of the error-correcting discrete sub-IPSs, and / or the at least one discrete redundant polymer sequences generated according to the method of any of the preceding paragraphs so as to produce one or more or all of: a) at least one error-correcting discrete sub-IPS polymer corresponding to the at least one error-correcting discrete sub-IPS sequence; b) at least one discrete redundant polymer corresponding to the at least one discrete redundant polymer sequence; c) at least one sub-informative polymer corresponding to at least one the discrete sub-informative polymer sequence to produce a collection of polymeric information tag polymers for example wherein the collection of polymeric information tag polymers comprises: at least one error-correcting discrete sub-IPS polymer and at least one discrete redundant polymer. In these instances it will be appreciated that the sequences of the polymers of the dPIT may have been pre-generated according to any of the methods described herein, such that the method of producing the polymers according to those designs does not require the generation of the dPIT to be part of the same method. In addition to providing methods of designing and generating, and physically producing the collection of polymers that comprise the dPIT, and providing the polymers that comprise the dPIT per se, the invention also provides a method of labelling an asset with the dPIT, said method comprising applying one or more of the polymers of the collection of polymers that comprise the dPIT to said asset. The invention also provides a method of labelling a biological asset that comprises one or more target polymers with the dPIT, said method comprising inserting or otherwise causing at least one of the polymers of the collection of polymers that comprise the dPIT to be physically located in one or more of the target biological polymers. The invention also provides an asset, for example a biological asset comprising at least part of the dPIT as described herein. Preference for various features such as the biological asset and dPIT are as described elsewhere herein. In some embodiments the asset, for example the biological asset comprises at least one or at least two of the polymers that make up the collection of polymers that comprise the dPIT. The asset may also be a non-biological asset as described elsewhere herein. In some embodiments, the biological asset does not comprise all of the discrete part-informative biological polymers and / or error-correcting part-informative biological polymers and / or discrete redundant biological polymers that together form the multipart biological signature. In one aspect, the invention provides a collection of biological assets that collectively comprise at least part of the multi-part biological signature of any of the preceding claims, optionally wherein: a) the collection is a collection of cells wherein different cells comprise different biological polymers of the signature, optionally different discrete part-informative biological polymers and / or error-correcting part-informative biological polymers and / or discrete redundant biological polymers that together form the multi-part biological signature; b) the collection is a collection of plasmids, wherein different plasmids comprise different biological polymers of the signature, optionally different discrete part- informative biological polymers and / or error-correcting part-informative biological polymers and / or discrete redundant biological polymers that together form the multipart biological signature. The biological asset or collection of biological asset may be any suitable biological asset. For example, in some embodiments the biological asset is: a) The genome of a cell b) A circular double-stranded DNA structure, optionally a plasmid c) A circular single-stranded DNA structure d) A nanostructure, wherein the nanostructure comprises at least one DNA or RNA or amino acid polymer molecule, optionally wherein the at least one DNA molecule is single-stranded or double-stranded and / or wherein the at least one RNA molecule is single-stranded or double-stranded; e) A phagemid f) A phage g) A virus, optionally a viral vector, optionally an AAV or lentiviral vector h) a cell population, optionally one or more chromosomes or plasmids of separate cells in a cell population; i) a cell, such as a bacterial cell, an archaeal cell, or a eukaryotic cell; j) a nanoparticle, for example a lipid nanoparticle, vesicle, micelle, for example a nanoparticle that is for use as a vaccine. In one aspect, the invention provides a method of identifying whether an asset comprises at least part of the multi-part biological signature of any of the preceding claims, wherein the method comprises retrieving the sequence of the biological polymer at the relevant target sites to generate a plurality of identification sequences. Any suitable method of retrieving the sequence may be employed. For example, where the biological polymer is a nucleic acid such as DNA or RNA, the sequence may be retrieved by nucleic acid sequencing using, for example, Sanger sequencing or nextgeneration sequencing. Where the biological polymer is a polypeptide, the sequence may be retrieved by protein sequencing. In some embodiments, the method further comprises comparing the retrieved identification sequences to the sequence that would be expected should the asset comprise at least part of the multi-part biological signature. The comparison may be performed with any suitable bioinformatic comparison tool, for example by using multiple sequence alignment. Suitable multiple sequence alignment tools are known to the person skilled in the art, and include for example NCBI-BLAST and MAFFT. In one aspect, the invention provides a method of identifying an asset, optionally a biological asset, comprising at least part of the multi-part biological signature of any of the preceding claims, wherein the method comprises: obtaining a digital polymer sequence at each of the target regions in the asset, wherein the digital polymer sequences comprise a plurality of symbols, comparing the retrieved sequences to a database of sequences and determining if the sequences are present or substantially present in the database (which can be centralised or decentralised, e.g., using distributed ledger technology) Reading the information symbols, thereby identifying the asset. In one aspect, the invention provides a method of identifying an asset, optionally a biological asset, comprising at least part of the multi-part biological signature of any of the preceding claims, wherein the method comprises: obtaining a digital polymer sequence at each of the target regions in the asset, wherein the digital polymer sequences comprise a plurality of symbols, Comparing the retrieved sequences to a database of sequences and determining if the sequences are present or substantially present in the database Reading the information symbols Performing error correction on the information symbols using the redundant symbols to retrieve the identity of the asset. As above, the database may be a centralised database ora decentralised database, for example using distributed ledger technology. In one aspect, the invention provides a computer implemented method for generating a biological marker sequence for insertion into a biological asset, the method comprising: receiving an information identifier string; partitioning the information identifier string to generate a plurality of identifier partitions each comprising a plurality of information symbols; and generating a plurality of redundant symbols by processing the plurality of information symbols of one or more identifier partitions with an error correction coding process; combining the redundant symbols with the plurality of identifier partitions to generate the biological marker sequence [as the plurality of identifier partitions]. In some embodiments, generating the plurality of redundant symbols comprises: for each identifier partition, applying an intra-partition error correction coding process to the plurality of information symbols within the identifier partition to generate a plurality of partition redundant symbols; and including the partition redundant symbols in the identifier partition. In some embodiments, the intra-partition error correction coding process comprises a Bose-Chaudhuri-Hocquenghem error correction coding process. In some embodiments, generating the plurality of redundant symbols comprises: applying an inter-partition error correction coding process to the plurality of identifier partitions to generate one or more redundant partitions comprising a respective plurality of redundant symbols. In some embodiments, the inter-partition error correction coding process comprises a Reed-Solomon error correction coding process. In some embodiments, the information symbols comprise a biological sequence element comprising one or more of: a nucleotide; and an amino acid. In some embodiments, the information identifier string comprises a plurality of information symbols and the plurality of information symbols of each identifier partition comprises a different subset of the plurality of information symbols of the information identifier string. In some embodiments, the information identifier string comprises a plurality of binary bits and the method comprises converting the plurality of binary bits to a plurality of information symbols. In some embodiments, the method further comprises: adding periodically spaced functionality symbols to each of the plurality of identifier partitions to satisfy a functionality requirement specification, optionally wherein the functionality requirement specification comprises: one or more prohibited features; and / or one or more desired features; Optionally wherein the one or more prohibited features comprise one or more of: a restriction enzyme recognition site; a GC content region exceeding a GC content threshold; a sequence prone to secondary structure formation; and / or wherein the one or more desired features comprise one or more of: a GC content within a GC content threshold range; secondary structure formations including one or more of: hairpins, kissing loops; and other sequences structures that do not disrupt gene function or that support genetic engineering. In one aspect, the invention provides a biological asset comprising a biological marker sequence, the biological marker sequence comprising: a plurality of identifier partitions each positioned at a different site in the biological asset, wherein each identifier partition comprises a plurality of information symbols representing a portion of an information identifier string; and a plurality of redundant symbols for providing error correction of the plurality of information symbols of the identifier partitions. In some embodiments, the plurality of redundant symbols comprises a plurality of partition redundant symbols in each of the identifier partitions. Wherein the plurality of partition redundant symbols of each identifier partition are derived from the information symbols of the identifier partition using an intra partition error correction coding process. In some embodiments, the plurality of redundant symbols comprises one or more redundant partitions each comprising a respective plurality of redundant symbols. Wherein the one or more redundant partitions are derived from the plurality of identifier partitions using an inter-partition error correction coding process. In one aspect, the invention provides a method of encoding a biological asset with a biological marker sequence, wherein the biological marker sequence comprises: a plurality of identifier partitions, wherein each identifier partition comprises a plurality of information symbols representing a portion of an information identifier string; and a plurality of redundant symbols for providing error correction of the plurality of information symbols of the identifier partitions, the method comprising: positioning at least one of the identifier partitions in the biological asset. In some embodiments, the method comprises storing at least one of the identifier partitions in a data storage facility. In one aspect, the invention provides a method of identifying a biological asset comprising a biological marker sequence, wherein the biological marker sequence comprises: a plurality of identifier partitions each positioned at a different site in the biological asset, wherein each identifier partition comprises a plurality of information symbols representing a portion of an information identifier string; and a plurality of redundant symbols for providing error correction of the plurality of information symbols of the identifier partitions, wherein the method comprises: reading the information symbols of the plurality of identifier partitions; and performing error correction on the plurality of identifier partitions using the redundant symbols to retrieve the information identifier string. In some embodiments, the method comprises: retrieving a further identifier partition from a data storage facility, the further partition comprising a plurality of further information symbols representing a further portion of the identifier string; and combining the further information symbols with the information symbols of the plurality of identifier partitions to retrieve the identifier string. In some embodiments, the biological asset may be used to mark a non-biological asset, for example may be applied to a non-biological asset. Such marking of a non-biological asset with the biological asset provided herein may, for example, allow the non-biological asset to be traced through a supply chain; may allow the ownership of the non-biological asset to be marked and subsequently determined; and / or may allow the origin of a non-biological asset to be marked and determined. As will be understood, any molecule or entity comprising a dPIT as described herein may be used to mark a non-biological entity as disclosed herein. The invention also provides embodiments set out in the following numbered paragraphs: 1. A method for generating a distributed polymeric information tag (dPIT) comprising a plurality of sub-informative polymer sequences (sub-IPS) said method comprising: a) partitioning information into at least two discrete information subsets and encoding each discrete information subset in a plurality of information symbols, wherein the plurality of information symbols represents a plurality of information monomers, to form the at least two discrete sub-informative polymer sequences (sub-IPS); b) providing an informative polymer sequence (IPS), wherein the informative polymer sequence comprises a plurality of information symbols representing a plurality of information monomers and partitioning the IPS into at least two discrete sub-IPSs sequences; and / or c) providing a plurality of discrete sub-informative polymer sequences (sub-IPS) wherein each discrete sub-IPS comprises a plurality of information symbols, wherein the plurality of information symbols represents a plurality of information monomers. 2. The method of paragraph 1 wherein: a) the information that is partitioned in (a) is binary information; and / or b) the informative polymer sequence of (b) was generated by encoding information in a plurality of information symbols to generate the informative polymer sequence, optionally wherein the information that is encoded is binary information. 3. The method of paragraph 1 or 2 wherein where the dPIT is for incorporation into a biological asset, the method comprises the step of obtaining a partition spectrum for said biological asset wherein said partition profile describes the frequency of available positions for insertion of discrete polymer fragments of different lengths into one or more corresponding polymers of said biological asset, optionally wherein the bin size of the partition spectrum is of any size, optionally wherein the partition spectrum describes the frequency of available positions for insertion of discrete polymer fragments with a bin size of between: 0-49, 50-99, 100-149, 150-199, 200-249, 250-299, 300-349, 350-399, 400-449, 450-499, 500-549, 550-599, 600-649, 650-699, 700-749, 750-799, 800-849, 850-899, 900-949, 950-999, 1000-1049, 1050-1099, 1100-1149, 1150-1199, 1200-1249, 1250-1299, 1300-1349, 1350-1399, 1400-1449, 1450-1499,1500-1549, 1550-1599, 1600-1649, 1650-1699, 1700-1749, 1750-1799, 1800-1849, 1850-1899, 1900-1949, 1950-1999, 2000-2049, 2050-2099, 2100-2149, 2150-2199, 2200-2249, 2250-2299, 2300-2349, 2350-2399, 2400-2449, 2450-2499, 2500-2549, 2550-2599, 2600-2649, 2650-2699, 2700-2749, 2750-2799, 2800-2849, 2850-2899, 2900-2949, 2950-2999, 3000-3049, 3050-3099, 3100-3149, 3150-3199,3200-3249, 3250-3299, 3300-3349, 3350-3399, 3400-3449, 3450-3499, 3500-3549, 3550-3599, 3600-3649, 3650-3699, 3700-3749, 3750-3799, 3800-3849, 3850-3899, 3900-3949, 3950-3999, 4000-4049, 4050-4099, 4100-4149, 4150-4199, 4200-4249, 4250-4299, 4300-4349, 4350-4399, 4400-4449, 4450-4499, 4500-4549, 4550-4599, 4600-4649, 4650-4699, 4700-4749, 4750-4799, 4800-4849, 4850-4899, 4900-4949, 4950-4999, 5000-5049, 5050-5099, 5100-5149, 5150-5199, 5200-5249, 5250-5299, 5300-5349, 5350-5399, 5400-5449, 5450-5499, 5500-5549, 5550-5599, 5600-5649 monomers in length, Optionally wherein the biological asset is selected from the group comprising or consisting of: a) one or more chromosomes of a cell, optionally: i) a prokaryotic cell, optionally a bacterial cell; ii) a eukaryotic cell, optionally a yeast cell, a fungal cell, a mammalian cell, a plant cell; b) A circular double-stranded DNA structure, optionally a plasmid c) A circular single-stranded DNA structure d) A nanostructure, wherein the nanostructure comprises at least one DNA or RNA or protein or peptide molecule, optionally wherein the at least one DNA molecule is singlestranded or double-stranded and / or wherein the at least one RNA molecule is singlestranded or double-stranded; e) a phagemid f) a phage g) a virus, optionally a viral vector, optionally an AAV or lentiviral vector; h) a nanoparticle, for example a lipid nanoparticle, vesicle, micelle, for example a nanoparticle that is for use as a vaccine. 4. The method of paragraph 3 wherein: a) where the method comprises partitioning information into at least two discrete information subsets, the length of and number of discrete information subsets is chosen based on the partition spectrum; b) where the method comprises providing an informative polymer sequence (IPS) and partitioning the IPS into at least two discrete sub-IPSs sequences, the length of and number of discrete sub-IPSs is chosen based on the partition spectrum; and / or c) where the method comprises providing a plurality of discrete sub-informative polymer sequences (sub-IPS), the length of and number of discrete sub-IPS is chosen based on the partition spectrum; such that polymeric fragments corresponding to said discrete sub-IPS into said biological asset would be accommodated within the biological asset. 5. The method of any of paragraphs 1-5 wherein said generating and / or partitioning and / or providing is in silico and said sequences are digital sequences. 6. The method of any of paragraphs 1-6, wherein the method comprises generating a plurality of redundant symbols by processing the information symbols present in the at least two discrete sub-informative polymer sequences (sub-IPS) with an error correction coding process. 7. The method of paragraph 6 further comprising combining the plurality of redundant symbols with the informative symbols, optionally wherein at least one or more or all of the at least two discrete sub-informative polymer sequences (sub-IPS) comprises redundant symbols to form at least one or more error-correcting discrete sub-IPS. 8. The method of any of paragraph 6 or 7 wherein: a) the error correction coding process is an intra-partition error correcting code process and is applied to the at least two discrete sub-informative polymer sequences (sub-IPS) to generate for each discrete sub-informative polymer sequences (sub-IPS) an additional sequence that is a redundant sequence comprising a plurality of symbols that represent a plurality of redundant monomers for use in error correction of the information monomers; and b) combining the redundant sequence comprising a plurality of redundant symbols with the corresponding discrete sub-informative polymer sequences (sub-IPS) to form error-correcting discrete sub-IPSs. 9. The method of any of paragraphs 1-8 wherein the method further comprises introducing of one or more characteristic symbols into one or more of the at least two discrete sub-informative polymer sequences (sub-IPS) or at least one of the errorcorrecting discrete sub-IPSs so as to: a) impart desirable sequence characteristics on the at least one or more of the at least two discrete part-informative biological polymer sequence; and / or b) remove undesirable sequence characteristics from the at least one or more of the at least two discrete part-informative biological polymer sequence; optionally wherein desirable sequence characteristics are characteristics are selected from: characteristics that camouflage the discrete part-informative sequence with respect to the target site in the one or more target biological polymers, balanced GC content, a secondary structure that is a hairpin, a secondary structure that is a kissing loop, or a sequence that supports genetic engineering; and / or wherein undesirable sequence characteristics are selected from: characteristics that are readily distinguishable from the surrounding sequence of the target site in the one or more target biological polymers, endonuclease recognition sites such as restriction endonuclease sites, greater than X% GC content, sequences prone to secondary structure formation, a sequence that may interfere with nearby gene expression. 10. The method of any of paragraphs 1-9 comprising the step of applying an interpartition error correction coding process to the at least two discrete sub-informative polymer sequences (sub-IPS) or to the at least one of the error-correcting discrete sub-IPSs to generate at least one discrete redundant polymer sequence comprising a plurality of monomers that are redundant. 11. The method of paragraph 10 further comprising introducing one or more characteristic symbols into one or more of the at least two discrete sub-informative polymer sequences (sub-IPS) and / or to at least one of the error-correcting discrete sub-IPSs, and / or to the at least one discrete redundant polymer sequences so as to: a) impart desirable sequence characteristics on the at least one or more of the at least two discrete part-informative biological polymer sequence; and / or b) remove undesirable sequence characteristics from the at least one or more of the at least two discrete part-informative biological polymer sequence; optionally wherein desirable sequence characteristics are characteristics are selected from: characteristics that camouflage the discrete part-informative sequence with respect to the target site in the one or more target biological polymers, balanced GC content, a secondary structure that is a hairpin, a secondary structure that is a kissing loop, or a sequence that supports genetic engineering; and / or wherein undesirable sequence characteristics are selected from: characteristics that are readily distinguishable from the surrounding sequence of the target site in the one or more target biological polymers, endonuclease recognition sites such as restriction endonuclease sites, greater than X% GC content, sequences prone to secondary structure formation, a sequence that may interfere with nearby gene expression. 12. The method of any of paragraphs 1-11 wherein: a) the at least one error-correcting discrete sub-IPS and the at least one discrete redundant polymer sequence together collectively form the distributed polymeric information tag (dPIT); b) the at least two discrete sub-informative polymer sequences (sub-IPS) together collectively form the distributed polymeric information tag (dPIT); c) the error-correcting discrete sub-IPSs together collectively form the distributed polymeric information tag (dPIT); or d) the at least two discrete sub-informative polymer sequences (sub-IPS) and the at least one discrete redundant polymer sequence together collectively form the distributed polymeric information tag (dPIT). 13. The method of any of paragraphs 8-12 wherein the intra-partition error correction coding process comprises a Bose-Chaudhuri-Hocquenghem error correction coding process; and / or where the inter-partition error correction coding process comprises a Reed-Solomon error correction coding process. 14 The method of any of paragraphs 1-13 wherein the polymer is selected from the group comprising or consisting of: Deoxyribonucleic acid (DNA), Ribonucleic acid (RNA), Protein or peptide or polypeptide, Peptide Nucleic Acids (PNAs), Polypeptides, Polysaccharides, Polyethylene Glycol (PEG), Polylactic Acid (PLA), Poly(ethylene terephthalate) (PET), Poly(methyl methacrylate) (PMMA), Polyaniline (PANI), Polycarbonate (PC), Polystyrene (PS), Poly(vinyl alcohol) (PVA), Polylactide-co-glycolide (PLGA), Polydimethylsiloxane (PDMS), Polyurethanes, Polycaprolactone (PCL), Polyhydroxyalkanoates (PHAs). 15. The method of paragraph 14 wherein: Where the polymer is DNA the monomer is selected from the group comprising or consisting of: Adenine (A), Thymine (T), Cytosine (C), Guanine (G), or a non-natural nucleotide, optionally wherein the nucleotide is a modified nucleotide; Where the polymer is RNA the monomer is selected from the group comprising or consisting of: Adenine (A), Uracil (U), Cytosine (C), Guanine (G), or a non-natural nucleotide, optionally wherein the nucleotide is a modified nucleotide; Wherein the polymer is a protein or peptide or polypeptide the monomer is selected from the group comprising or consisting of: an Amino acids optionally alanine (A), arginine (R), asparagine (N), aspartic acid (D), cysteine (C), glutamic acid (E), glutamine (Q), glycine (G), histidine (H), isoleucine (I), leucine (L), lysine (K), methionine (M), phenylalanine (F), proline (P), serine (S), threonine (T), tryptophan (W), tyrosine (Y), and valine (V), or a non-natural amino acid, optionally wherein the amino acid is a modified amino acid. Where the polymer is a PNA the monomer is selected from the group comprising or consisting of: Adenine (A), Thymine (T), Cytosine (C), Guanine (G), or a non-natural nucleotide, optionally wherein the nucleotide is a modified nucleotide; Where the polymer is a polysaccharide the monomer is selected from the group comprising or consisting of: Glucose, Fructose, Galactose, Mannose; Where the polymer is PEG the monomer is selected from the group comprising or consisting of: Ethylene glycol, Propylene glycol, Butylene glycol, Diethylene glycol; Where the polymer is PLA the monomer is selected from the group comprising or consisting of: L-lactic acid, D-lactic acid, Mesolactic acid, 2-Hydroxypropanoic acid; Where the polymer is PET the monomer is selected from the group comprising or consisting of: Ethylene glycol, Terephthalic acid, Isophthalic acid, Naphthalene dicarboxylate; Where the polymer is PMMA the monomer is selected from the group comprising or consisting of: Methyl methacrylate, Ethyl methacrylate, Butyl methacrylate, 2-Hydroxyethyl methacrylate; Where the polymer is PANI the monomer is selected from the group comprising or consisting of: Aniline, o-Toluidine, m-Toluidine, p-Toluidine; Where the polymer is PC the monomer is selected from the group comprising or consisting of: Bisphenol A, Bisphenol S, Cyclohexane dimethanol, Terephthalic acid; Where the polymer is PS the monomer is selected from the group comprising or consisting of: Styrene, a-Methylstyrene, Butadiene, Acrylonitrile; Where the polymer is PVA the monomer is selected from the group comprising or consisting of: Vinyl acetate, Vinyl alcohol, Ethylene, Acrylamide; Where the polymer is PLGA the monomer is selected from the group comprising or consisting of: L-lactic acid, D-lactic acid, Glycolic acid, 2-Hydroxypropanoic acid; Where the polymer is PDMS the monomer is selected from the group comprising or consisting of: Dimethylsiloxane, Diphenylsiloxane, Methylphenylsiloxane, Vinylsiloxane; Where the polymer is polyurethane the monomer is selected from the group comprising or consisting of: Toluene diisocyanate, Methylene diphenyl diisocyanate, Hexamethylene diisocyanate, Polycaprolactone polyol; Where the polymer is PCL the monomer is selected from the group comprising or consisting of: E-Caprolactone, 6-Valerolactone, y-Butyrolactone, L-Lactide; Where the polymer is PHA the monomer is selected from the group comprising or consisting of: 3-Hydroxybutyrate, 3-Hydroxyvalerate, 3-Hydroxyhexanoate, 4-Hydroxybutyrate. 16. The method of any of the preceding paragraphs further comprising, wherein where the at least two discrete sub-informative polymer sequences (sub-IPS) and / or the at least one of the error-correcting discrete sub-IPSs, and / or the at least one discrete redundant polymer sequences are protein sequences, the step of preparing a DNA or RIMA sequence that encodes said at least two discrete part-informative biological polymer sequences and / or the at least two error-correcting part-informative biological polymer sequences and / or the at least one discrete redundant biological polymer sequence. 17. The method of any of the preceding paragraphs comprising producing one or more polymers corresponding to the sequences set out in the at least two discrete sub-informative polymer sequences (sub-IPS) and / or the at least one of the errorcorrecting discrete sub-IPSs, and / or the at least one discrete redundant polymer sequences so as to produce one or more or all of: a) at least one error-correcting discrete sub-IPS polymer corresponding to the at least one error-correcting discrete sub-IPS sequence; b) at least one discrete redundant polymer corresponding to the at least one discrete redundant polymer sequence; c) at least one sub-informative polymer corresponding to at least one the discrete sub-informative polymer sequence to produce a collection of polymeric information tag polymers optionally wherein the collection of polymeric information tag polymers comprises: at least one error-correcting discrete sub-IPS polymer and at least one discrete redundant polymer. 18. The method of paragraph 17 further comprising the step of applying, inserting or otherwise causing at least one of the polymers of the collection of the polymeric information tag polymers to be physically located in or on one or more assets. 19. The method of paragraph 18 wherein the asset is a biological asset. 20. The method of paragraph 19 wherein the biological asset comprises one or more target polymers corresponding to the type of polymer of the collection of polymeric information tag polymers. 21. The method of paragraph 20 wherein the target polymers are selected from the group comprising or consisting of: a) one or more chromosomes of a cell, optionally: i) a prokaryotic cell, optionally a bacterial cell; ii) a eukaryotic cell, optionally a yeast cell, a fungal cell, a mammalian cell, a plant cell; b) A circular double-stranded DNA structure, optionally a plasmid c) A circular single-stranded DNA structure d) A nanostructure, wherein the nanostructure comprises at least one DNA or RNA or protein or peptide molecule, optionally wherein the at least one DNA molecule is singlestranded or double-stranded and / or wherein the at least one RNA molecule is singlestranded or double-stranded; e) a phagemid f) a phage g) a virus, optionally a viral vector, optionally an AAV or lentiviral vector; h) one or more chromosomes or plasmids of separate cells in a cell culture; i) a nanoparticle, for example a lipid nanoparticle, vesicle, micelle, for example a nanoparticle that is for use as a vaccine. 22. The method of any of paragraphs 18-21 wherein the method comprises the step of inserting or otherwise causing more than one polymer of the collection of polymeric information tag polymers to be physically located in one or more of the target biological polymers, optionally to be physically located in: a) The genome of the same cell b) The same circular double-stranded DNA structure, optionally a plasmid c) The same circular single-stranded DNA structure d) The same nanostructure, wherein the nanostructure comprises at least one DNA or RNA or amino acid polymer molecule, optionally wherein the at least one DNA molecule is single-stranded or double-stranded and / or wherein the at least one RNA molecule is single-stranded or double-stranded; e) The same phagemid f) The same phage g) The same virus, optionally a viral vector, optionally an AAV or lentiviral vector h) members of a cell population, optionally one or more chromosomes or plasmids of separate cells in a cell population; I) a nanoparticle, for example a lipid nanoparticle, vesicle, micelle, for example a nanoparticle that is for use as a vaccine, optionally wherein all of the polymers of the collection of polymeric information tag polymers are physically located in one or more of the target biological polymers. 23. The method of any of paragraphs 20-22 wherein at least one of the polymers of the collection of polymeric information tag polymers is not embedded in the same asset, optionally not embedded in the same one or more target biological polymers as the rest of the polymers of the collection of polymeric information tag polymers, optionally wherein: at least one of the polymers of the collection of polymeric information tag polymers is not embedded a) in the genome of the same cell; b) in the same circular double-stranded DNA structure; c) in the same circular single-stranded DNA structure; d) in the same nanostructure; e) in the same phagemid; f) in the same phage; g) in the same virus, optionally not in the same viral vector, optionally not in the same AAV or lentiviral vector; h) in members of the same cell population i) in the same nanoparticle, for example same lipid nanoparticle, vesicle, micelle etc. as the rest of the polymers of the collection of polymeric information tag polymers. 24. The method of any of paragraphs 1-23 wherein at least one of the: a) at least one error-correcting discrete sub-IPS polymer sequences; b) at least one discrete redundant polymer sequence; c) at least one sub-informative polymer sequence; remains as a digital sequence and a corresponding polymer is not produced. 25. A computer implemented method for generating a distributed polymeric information tag (dPIT)for incorporation into or application to one or more assets, the method comprising the method of any one or more of paragraphs 1-24. 26. A computer implemented method for generating a distributed polymeric information tag (dPIT)for incorporation into or application to one or more assets, the method comprising: receiving an informative polymer sequence (IPS), wherein the informative polymer sequence comprises a plurality of information symbols representing a plurality of information monomers; partitioning the informative polymer sequence (IPS)to generate at least two discrete sub-IPSs sequences each comprising a plurality of information symbols; and generating a plurality of redundant symbols by processing the plurality of information symbols of the at least two discrete sub-IPSs sequences with an error correction coding process; combining the redundant symbols with the at least two discrete sub-IPSs sequences to generate at least one error-correcting discrete sub-IPS polymer sequences that each comprises the plurality of information symbols of the discrete sub-IPSs sequences and a plurality of redundant symbols, optionally wherein the error correction coding process comprises any of the following: Bose-Chaudhuri-Hocquenghem; Hamming Codes: Linear block codes over GF(2); LDPC Codes: Linear block codes over GF(2); and or Convolutional Codes: Codes generated by convolution of input bits; Reed-Solomon codes; non-binary cyclic code over finite fields (GF(q), typically 2m). 27. The computer implemented method of paragraph 26, wherein generating the plurality of redundant symbols comprises: for each of the at least two discrete sub-IPSs sequences, applying an intra-partition error correction coding process to the plurality of information symbols within the discrete sub-IPSs sequences to generate a plurality of partition redundant symbols; and including the partition redundant symbols in the discrete sub-IPSs sequences. 28. The computer implemented method of paragraph 27, wherein the intra-partition error correction coding process comprises any of the following: a Bose-Chaudhuri-Hocquenghem error correction coding processHamming Codes: Linear block codes over GF(2); LDPC Codes: Linear block codes overGF(2); and or Convolutional Codes: Codes generated by convolution of input bits. 29. The method of any of paragraphs 26-28, wherein generating the plurality of redundant symbols comprises: applying an inter-partition error correction coding process to the plurality of information symbols within each of the at least two discrete sub-IPSs sequences to generate one or more discrete redundant b polymer sequence comprising a respective plurality of redundant symbols. 30. The method of paragraph 29, wherein the inter-partition error correction coding process comprises a Reed-Solomon error correction coding process or non-binary cyclic code over finite fields (GF(q), typically 2m). 31. A digital distributed polymeric information tag (dPIT) produced according to any of the preceding methods. 32. A physical distributed polymeric information tag (dPIT) comprising the collection of polymeric information tag polymers of any of the preceding paragraphs. 33. A physical distributed polymeric information tag (dPIT) comprising a collection of polymeric information tag polymers, wherein the collection of polymeric information tag polymers comprises: a) at least one error-correcting discrete sub-IPS polymer produced according to a method of any of the preceding paragraphs and at least one discrete redundant polymer produced according to a method of any of the preceding paragraphs; b) at least one error-correcting discrete sub-IPS polymer produced according to a method of any of the preceding paragraphs; c) at least one discrete redundant polymer produced according to a method of any of the preceding paragraphs; and / or d) at least one sub-informative polymer. 34. The physical multi-part biological signature of any of the preceding paragraphs further comprising a non-physical digital discrete biological polymer sequence as defined in any of the preceding paragraphs, optionally that is stored on a data storage facility. 35. A method of labelling an asset with a multi-part biological signature, said method comprising applying one or more of the discrete biological polymers of the physical multi-part biological signature of any of paragraphs 32-34 to said asset. 36. A method of labelling a biological asset that comprises one or more target polymers with a multi-part biological signature, said method comprising inserting or otherwise causing at least one of the discrete biological polymers of the physical multipart biological signature of any of paragraphs 32-34 to physically located in one or more of the target biological polymers. 37. A biological asset comprising at least part of the distributed polymeric information tag (dPIT) or at least one or more polymers of the collection of polymeric information tag polymers as defined in any of the preceding paragraphs. 38. The biological asset of paragraph 37 wherein the biological asset does not comprise all of the polymers of the collection of polymeric information tag polymers. 38. A collection of biological assets that collectively comprise at least part of the distributed polymeric information tag (dPIT) or at least one or more polymers of the collection of of polymeric information tag polymers as defined in any of the preceding paragraphs, optionally wherein: a) the collection of biological assets is a collection of cells wherein different cells comprise different polymers of the collection of polymeric information tag polymers; b) the collection of biological assets is a collection of plasmids, wherein different plasmids comprise different polymers of the collection of polymeric information tag polymers. 39. A method of identifying whether an asset comprises at least part of the distributed polymeric information tag (dPIT) or at least one or more polymers of the collection of polymeric information tag polymers as defined in any of the preceding paragraphs, wherein the method comprises retrieving the sequence of the biological polymer at the relevant target site(s) to generate a plurality of identification sequences. 40. The method of paragraph 36 further comprising comparing the retrieved identification sequences to the sequence that would be expected should the asset comprise at least part of the distributed polymeric information tag (dPIT) or at least one or more polymers of the collection of polymeric information tag polymers as defined in any of the preceding paragraphs. 41. A method of identifying an asset, optionally a biological asset, comprising at least part of the distributed polymeric information tag (dPIT) or at least one or more polymers of the collection of polymeric information tag polymers as defined in any of the preceding paragraphs, wherein the method comprises: obtaining a digital polymer sequence at each of the target regions in the asset, wherein the digital polymer sequences comprise a plurality of symbols, Comparing the retrieved sequences to a database of sequences and determining if the sequences are present or substantially present in the database Reading the information symbols, thereby identifying the asset, optionally wherein the database is a centralised database or a decentralised database, optionally using distributed ledger technology. 42. A method of identifying an asset, optionally a biological asset, comprising at least part of the distributed polymeric information tag (dPIT) or at least one or more polymers of the collection of polymeric information tag polymers as defined in any of the preceding paragraphs, wherein the method comprises: obtaining a digital polymer sequence at each of the target regions in the asset, wherein the digital polymer sequences comprise a plurality of symbols, Comparing the retrieved sequences to a database of sequences and determining if the sequences are present or substantially present in the database Reading the information symbols Performing error correction on the information symbols using the redundant symbols to retrieve the identity of the asset, optionally wherein the database is a centralised database or a decentralised database, optionally using distributed ledger technology. 43. A computer implemented method for generating a biological marker sequence for insertion into a biological asset, the method comprising: receiving an information identifier string; partitioning the information identifier string to generate a plurality of identifier partitions each comprising a plurality of information symbols; and generating a plurality of redundant symbols by processing the plurality of information symbols of one or more identifier partitions with an error correction coding process; combining the redundant symbols with the plurality of identifier partitions to generate the biological marker sequence [as the plurality of identifier partitions]. 44. The method of paragraph 43, wherein generating the plurality of redundant symbols comprises: for each identifier partition, applying an intra-partition error correction coding process to the plurality of information symbols within the identifier partition to generate a plurality of partition redundant symbols; and including the partition redundant symbols in the identifier partition. 45. The method of paragraph 44, wherein the intra-partition error correction coding process comprises a Bose-Chaudhuri-Hocquenghem error correction coding process. 46. The method of any of paragraphs 43-45, wherein generating the plurality of redundant symbols comprises: applying an inter-partition error correction coding process to the plurality of identifier partitions to generate one or more redundant partitions comprising a respective plurality of redundant symbols. 47. The method of paragraph 46, wherein the inter-partition error correction coding process comprises a Reed-Solomon error correction coding process. 48. The method of any of paragraphs 43-47 wherein the information symbols comprise a biological sequence element comprising one or more of: a nucleotide; and an amino acid. 49. The method of any of paragraphs 43-48 wherein the information identifier string comprises a plurality of information symbols and the plurality of information symbols of each identifier partition comprises a different subset of the plurality of information symbols of the information identifier string. 50. The method of any of paragraphs 43-49 wherein the information identifier string comprises a plurality of binary bits and the method comprises converting the plurality of binary bits to a plurality of information symbols. 51. The method of any of paragraphs 44-50 further comprising: adding periodically spaced functionality symbols to each of the plurality of identifier partitions to satisfy a functionality requirement specification, optionally wherein the functionality requirement specification comprises: one or more prohibited features; and / or one or more desired features; Optionally wherein the one or more prohibited features comprise one or more of: a restriction enzyme recognition site; a GC content region exceeding a GC content threshold; a sequence prone to secondary structure formation; and / or wherein the one or more desired features comprise one or more of: a GC content within a GC content threshold range; secondary structure formations including one or more of: hairpins, kissing loops; and other sequences structures that do not disrupt gene function or that support genetic engineering. 52. A biological asset comprising a biological marker sequence, the biological marker sequence comprising: a plurality of identifier partitions each positioned at a different site in the biological asset, wherein each identifier partition comprises a plurality of information symbols representing a portion of an information identifier string; and a plurality of redundant symbols for providing error correction of the plurality of information symbols of the identifier partitions. 53. The biological asset of paragraph 52 wherein the plurality of redundant symbols comprises a plurality of partition redundant symbols in each of the identifier partitions. Wherein the plurality of partition redundant symbols of each identifier partition are derived from the information symbols of the identifier partition using an intra partition error correction coding process. 54. The biological asset of paragraph 52 or 53 wherein the plurality of redundant symbols comprises one or more redundant partitions each comprising a respective plurality of redundant symbols. Wherein the one or more redundant partitions are derived from the plurality of identifier partitions using an inter-partition error correction coding process. 55. A method of encoding a biological asset with a biological marker sequence, wherein the biological marker sequence comprises: a plurality of identifier partitions, wherein each identifier partition comprises a plurality of information symbols representing a portion of an information identifier string; and a plurality of redundant symbols for providing error correction of the plurality of information symbols of the identifier partitions, the method comprising: positioning at least one of the identifier partitions in the biological asset. 56. The method of paragraph 55 wherein the method comprises storing at least one of the identifier partitions in a data storage facility. 57. A method of identifying a biological asset comprising a biological marker sequence, wherein the biological marker sequence comprises: a plurality of identifier partitions each positioned at a different site in the biological asset, wherein each identifier partition comprises a plurality of information symbols representing a portion of an information identifier string; and a plurality of redundant symbols for providing error correction of the plurality of information symbols of the identifier partitions, wherein the method comprises: reading the information symbols of the plurality of identifier partitions; and performing error correction on the plurality of identifier partitions using the redundant symbols to retrieve the information identifier string. 58. The method of any of paragraphs 55-57 wherein the method comprises: retrieving a further identifier partition from a data storage facility, the further partition comprising a plurality of further information symbols representing a further portion of the identifier string; and combining the further information symbols with the information symbols of the plurality of identifier partitions to retrieve the identifier string. 59. A method of producing one or more polymers corresponding to the sequences set out in any of the at least two discrete sub-informative polymer sequences (sub-IPS) and / or the at least one of the error-correcting discrete sub-IPSs, and / or the at least one discrete redundant polymer sequences generated according to the method of any of the preceding paragraphs so as to produce one or more or all of: a) at least one error-correcting discrete sub-IPS polymer corresponding to the at least one error-correcting discrete sub-IPS sequence; b) at least one discrete redundant polymer corresponding to the at least one discrete redundant polymer sequence; c) at least one sub-informative polymer corresponding to at least one the discrete sub-informative polymer sequence to produce a collection of polymeric information tag polymers optionally wherein the collection of polymeric information tag polymers comprises: at least one error-correcting discrete sub-IPS polymer and at least one discrete redundant polymer. Figure legends Figure 1 - (A) Flow chart showing the exemplary generation of a signature according to the methods provided herein, starting from optional digital information that is partitioned before being converted into part-alternative biological polymer sequences. (B) Flow chart showing the exemplary generation of a signature according to the methods provided herein, starting from optional digital information that is converted into informative biological polymer sequence before being partitioned into part-informative biological polymer sequences. (C) Flow chart showing the exemplary generation of a signature according to the methods provided herein, starting from an informative biological polymer sequence. In each of Figures 1 (A)-(C), the optional step of introducing of one or more characteristic symbols to impart desirable sequence characteristics and / or remove undesirable characteristics may be performed after intra-partition error correction, after inter-partition error correction, after both steps, or may not be performed. All sequences shown are exemplary. DNA and protein sequences displayed are unrelated. Figure 2 - Distribution of intergenic region lengths after motif prediction in desired length, calculated from exemplary species Clostridium butyricum. By way of exemplary explanation, there are potentially 487 neutral sites of length 50 to 99 nucleotides in the C. butyricum genome in which 5 partitions of 62 nucleotides encoding an error-corrected barcode of 168 bits, offering 223,741,613,032 options for storing the barcode in 5 different locations in the C. butyricum genome. Examples Figure 1(A) illustrates a method for generating a distributed polymeric information tag (dPIT) comprising a plurality of sub-informative polymer sequences (sub-IPS) according to an embodiment of the present disclosure. A first step 102 comprises receiving information. In this example the information comprises binary information. The information may be any length. A second step 104 comprises partitioning the information into at least two discrete information subsets. In this example, the second step comprises partitioning the information into three discrete information subsets. A third step 106 comprises encoding each discrete information subset in a plurality of information symbols, wherein the plurality of information symbols represents a plurality of information monomers, to form the at least two discrete sub-informative polymer sequences (sub-IPS). In this example, the information symbols represent one of Adenine (A), Thymine (T), Cytosine (C), Guanine (G). Although illustrated as A, T, C or G, the information symbols may be represented numerically, for example by 0, 1, 2, 3 such that they can be more easily processed with error correction coding. In this example, optional error correction coding is applied in fourth and sixth steps, 108, 112. In some examples, either or both steps may be omitted. The fourth step 108 comprises generating a plurality of redundant symbols by processing the information symbols present in each sub-IPS with an intra-partition error correction coding process. In this example, the fourth step comprises combining the plurality of sub-IPSs with their corresponding redundant symbols (a redundant sequence) to form a respective plurality of error correcting discrete sub-IPSs. In this example, the redundant symbols are appended to the end of each corresponding sub-IPS to form the respective error correcting discrete sub-IPS. An optional fifth step comprises 110 comprises introducing one or more characteristic symbols into at least one of (in this example each one of) the error-correcting discrete sub-IPS to: a) impart desirable sequence characteristics on the at least one or more of the at least two discrete part-informative biological polymer sequence; and / or b) remove undesirable sequence characteristics from the at least one or more of the at least two discrete part-informative biological polymer sequence. In some examples the fifth step may be performed after the sixth step, as illustrated. The sixth step 112 comprises generating at least one discrete redundant polymer sequence comprising a plurality of monomers that are redundant, by applying an interpartition error correction coding process to at least one of (in this example each of) the error-correcting discrete sub-IPSs. In this example, the inter-partition error coding process generates two redundant polymer sequences from the three error-correcting discrete sub-IPSs. A seventh optional step 114 comprises incorporating the one or more of the sub-IPSs into one or more target biological polymers. Figure 1(B) illustrates a method for generating a distributed polymeric information tag (dPIT) comprising a plurality of sub-informative polymer sequences (sub-IPS) according to another embodiment of the present disclosure. The method of Figure 1(B) is substantially the same to the method of Figure 1(A) other than the second and third steps have been swapped. That is the method comprises receiving information and encoding the information in a plurality of information symbols, wherein the plurality of information symbols represents a plurality of information monomers, to form an informative polymer sequence (IPS). The method comprises partitioning the IPS into at least two discrete sub-IPSs sequences. The fourth to seventh steps of Figure 1(B) are the same as those of Figure 1(A). Figure 1(C) illustrates a method for generating a distributed polymeric information tag (dPIT) comprising a plurality of sub-informative polymer sequences (sub-IPS) according to another embodiment of the present disclosure. The method of Figure IC is substantially the same to the method of Figure IB other than the first step of receiving information and the second step of encoding the information has been omitted. Instead Figure 1(C) comprises receiving the IPS comprising a plurality of information symbols, wherein the plurality of information symbols represents a plurality of information monomers. The second to sixth steps of Figure 1(C) are the same as the third to seventh steps of Figure 1(C). To address one or more of the challenges described herein, some embodiments of the solution involve a sophisticated error-correcting scheme that incorporates the following strategies: 1. Error Mitigation through Barcode Double Error Correction The barcode is partitioned into several sections, each section is then made errorresistant with an intra-partition (within partition) error correcting code. The ECC divided barcode can be distributed across different genomic locations within a single cell or different cells from an initially identical monoclonal culture. This partitioning has its own error correcting code that enhances redundancy and allows for error detection and correction across partitions. 2. Consulting "Tabu" and "Wish" Lists To navigate biological constraints, our coding scheme uses a 'tabu list' and a 'wish list'. The tabu list includes features that the error-corrected encoded barcode must avoid, such as restriction enzyme recognition sites, high GC content regions, and sequences prone to secondary structure formation. The wish list comprises features that are desirable, such as balanced GC content, secondary structure formations such as hairpins, kissing loops, etc., or other sequences / structures that do not disrupt gene function or that support genetic engineering. 3. Adversarial Resilience The partitioning of barcodes not only helps in error correction but also in mitigating adversarial attacks. If an adversary tampers with or removes a part of the barcode (i.e., a partition), the remaining sections can still provide enough information for error correction and validation. 4. Biological Compatibility Our solution ensures that the barcodes do not interfere with the bioasset's biology. Our barcodes / watermarks can be inserted at a number of neutral sites, at site of knockout genes or, if desired, next to knock-in genes or other regulatory sequences. The patterns on the tabu and wish lists help in optimizing the barcodes for minimal impact on the bioasset's functionality. For example, by avoiding insertion sites that could disrupt essential genes or regulatory elements. 5. Customizable Patterns The exact patterns in the tabu and wish lists can be optimized for different applications, species, and situations. This customization allows for the generation of barcodes that are tailored to specific bioassets and their intended uses. That is, our scheme is able to encode the same binary data snippet {e.g., barcode, watermark, etc.) differently for different {e.g.) cell species. 6. Partial Integration The integration of barcodes can be full or partial. Partial integration provides the opportunity to distribute the barcode information between the GitLife Biotech Ltd version control system and data bases and the asset owners (or their consumers). For example, for a 5-way partitioned barcode (see below), 3 partitions could be inserted into the bioasset, 1 partition retained fully digitally or molecularly encoded by GitLife Biotech Ltd and 1 partition provided either digitally or molecularly encoded to a GitLife customer who is the bioasset owner. A third party, to decode the barcode and ascertain, e.g., its provenance and right management access would need to decode the partitions in the bioassets and secure the other two partitions from GitLife and its customer respectively. This distribution can enhance security, support a simple authentication mechanism mediated by GitLife Biotech Ltd and enable new collaborative and business models for bioasset distribution, authentication, and IP rights managements. Solution Details Our workflow effectively combines in new ways that error correcting processes such as a BCH as an intra-partion error correction coding process and a Reed-Solomon error correcting code as an inter-partition error correcting coding process, with tabu list and wish lists to ensure robust data integrity and error correction within and across multiple nucleotide locations of a given number and size that is bioasset {e.g., species) specific. This encoding scheme leverages BCH codes for error correction within nucleotide locations and Reed-Solomon codes for error correction across multiple locations. The innovation ensures robust data transmission and recovery by efficiently encoding a 168-bit message into nucleotide sequences spread across five distinct locations that avoids sequences in the tabu list and generates sequences in the wish list. Reference to BCH and Reed-Solomon in the example below should be interpreted as referring to any suitable error correcting process, for example selected from any one or more of: Bose-Chaudhuri-Hocquenghem; Hamming Codes: Linear block codes over GF(2); LDPC Codes: Linear block codes over GF(2); and or Convolutional Codes: Codes generated by convolution of input bits; Reed-Solomon codes; non-binary cyclic code over finite fields (GF(q), typically 2m). Reference to BCH can for instance particularly also refer to Hamming Codes: Linear block codes over GF(2); LDPC Codes: Linear block codes over GF(2); and or Convolutional Codes: Codes generated by convolution of input bits. Reference to Reed-Solomon can also for instance particularly refer to non-binary cyclic code over finite fields (GF(q), typically 2m). Kev Components 0. Determination of position and number of barcodes partitions: - The system determines, for a given biological asset, a range of possible locations of where the barcode could be divided and placed depending on the size (information bits) that need to be stored. 1. Information Bits: - The system handles 168 information bits by distributing them across five nucleotide locations, utilizing error correction. 2. Error Correction: - Across Locations: Reed-Solomon codes enable correction of up to two missing locations or one erroneous location. - Within Locations: BCH codes correct up to five errors within each location. 3. Handling Tabu and Wish Sequences Lists: - A template scheme reserves every fifth base to break taboo sequences or to create target sequences, ensuring sequence and error correcting code integrity. Kev Parameters The workflow has various parameters that can be fine-tuned for different situations such as the size of the data snippet (barcode or watermark) to insert into the bioasset, the number of locations within the biological asset into which the barcode would be divided for insertion, the number of bases per location (e.g., how much barcode / watermark data plus redundancy information one may want to store in each location), the information symbols per location (how much of that data itself is actually barcode / watermark rather than redundancy). - Biological asset DNA (or RNA) sequence (e.g., strain genome, plasmid sequence or origami scaffold and staples): any length. - Tabu list and Wish list: two lists containing DNA or RNA sequences to avoid or generate respectively. - Number of Information Bits (e.g. barcode size): e.g. 168 bits (including padding of 4 bits if necessary) - Number of Locations (e.g. barcode partitions): e.g., 5 For example, the bar chart in Figure 2 shows that for this example species {E.coli} there at potentially 709 neutral sites of length between 50 and 99 nucleotides in its genome in which one could "hide" 5 chunks of 62 nucleotides encoding an errorcorrecting barcode of 168 bits. That is, our scheme allows us to choose one out of 709 taken by 5 combinatorial possibilities, i.e., 1,472,012,434,641 options, for storing our barcodes when divided into 5 locations. Alternatively, one may insert the 5 partitions into, e.g., 5 gene knock-outs or adjacent (with suitable insulator sequences) to 5 gene know-ins. - Number of Bases per Location: e.g., 62 - Alphabet size: e.g., 4 for nucleotides (DNA / RNA) or 20 for proteins although non-canonical nucleotides and amino acids can also be considered. - Information Symbols per Location: e.g. 50 (after reserving bases for taboo breaking) The example below describes the encoding process with the parameters values above as examples. Encoding Process The scheme leverages BCH for coding information bits within a location and Reed-Solomon codes for error correction between / across locations. It is composed of Input Message Start with a 164-bit message {e.g., a barcode, watermark, etc.}, pad it with 4 additional bits to make it 168 bits. Distribute Message Across Locations and Encode with BCH Split the 168-bit message into three chunks of 56 bits each. For each of the five locations, encode these chunks using the shortened BCH code [50, 28, >11]. Convert Bits to Symbols Convert each 56-bit chunk into 28 F4 symbols (since 1 symbol in F4 = 2 bits). F4 symbols represent the 4 nucleotide base and it could be for DNA or RNA. Encode these symbols to get 50 F4 symbols, incorporating 22 redundancy symbols. Incorporate Tabu Sequence-Breaking Template or Wish List Sequence-Creating Template Insert reserved bases every fifth position within the 50 F4 symbols to break taboo sequences or create wish list sequences, resulting in a final sequence length of 62 bases per location. Reed-Solomon Encoding Across Locations Encode the five 56-bit chunks (one per location) with Reed-Solomon code to generate redundancy, allowing for inter-location error correction. Detailed Codes Reed-Solomon Code Parameters: [5, 3, 3] over Fie Provides redundancy across locations, correcting up to two missing or one erroneous location. BCH Codes Original BCH Code Parameters: [63, 41, >11] over F4 - Information Symbols: 41 - Redundancy Symbols: 22 - Can correct up to 5 errors. Shortened BCH Code: [50, 28, >11] - Adjusted to fit the constraints of the nucleotide bases and taboo-breaking requirements. Pseudocode Explanation Function: Encode 1. Input: 168 bits (bo, bi, ..., bie?). 2. Divide Bits: - Split into three locations: b(t) <— (bsej, b56]+i, ..., bs6i+ss) for f = 0, 1, 2. 3. Encode Each Location: - For each location, encode using the function EncodeLocation. 4. Convert Symbols: - Convert adjacent pairs of elements from F4 to single elements in Fie. 5. Compute Reed-Solomon Redundancy: - Generate redundancy symbols across the locations. 6. Output: - The encoded bases in all five locations. Function: EncodeLocation 1. Input: 56 bits (xo, xi, ..., xss). 2. Convert to Symbols: - Convert pairs of bits from F2 to single symbols in F4. 3. Compute BCH Redundancy: - Generate redundancy symbols using shortened BCH code. 4. Break Taboo Sequences: - Insert reserved bases to break taboo sequences. 5. Output: - A 62-element vector over F4. Function: ComputeShortenedBCHRedundancv 1. Input: 28 symbols over F4. 2. Define Polynomial: - Compute the polynomial representation and redundancy symbols. 3. Output: - Redundancy symbols. Function: ComputeRSRedundancv 1. Input: 3 symbols over Fie. 2. Define Polynomial: - Compute the polynomial representation and redundancy symbols. 2.2. Break Taboo Sequences: - Insert reserved bases to break taboo sequences. 3. Output: - Redundancy symbols. Benefits of the Solution The key benefits of our solution include: 1. Enhanced Security The partitioning and distribution of barcodes increase the security of the system by making it more resilient to tampering and errors. 2. Improved Accuracy Error-correcting codes and the use of tabu and wish lists ensure high accuracy in maintaining the integrity of the barcodes. 3. Flexibility and Customization The ability to customize the barcode patterns for different applications and species makes this solution highly flexible and adaptable. 4. Compliance and IP Protection By linking physical samples to their digital counterparts and asserting IP rights, this solution supports compliance with regulatory requirements and protects intellectual property. 5. Collaboration and Sharing Partial integration facilitates better access and benefit sharing (ABS) management, ensuring fair distribution of benefits derived from bioassets. Summary Our innovative error-correcting scheme for digital snippets in biological assets provides a robust, secure, and flexible solution to the challenges posed by integrating DNA barcodes into bioassets. By addressing errors from DNA processes, biological constraints, and adversarial attacks, and by ensuring biological compatibility and customization, our solution enhances the utility and reliability of synthetic biology applications. This advancement supports the responsible and secure development of biotechnological innovations, paving the way for future breakthroughs in the field.

Claims

1. A method for generating a distributed polymeric information tag (dPIT) comprising a plurality of sub-informative polymer sequences (sub-IPS) said method comprising:a) partitioning information into at least two discrete information subsets and encoding each discrete information subset in a plurality of information symbols, wherein the plurality of information symbols represents a plurality of information monomers, to form the at least two discrete sub-informative polymer sequences (sub-IPS);b) providing an informative polymer sequence (IPS), wherein the informative polymer sequence comprises a plurality of information symbols representing a plurality of information monomers and partitioning the IPS into at least two discrete sub-IPSs sequences;and / orc) providing a plurality of discrete sub-informative polymer sequences (sub-IPS) wherein each discrete sub-IPS comprises a plurality of information symbols, wherein the plurality of information symbols represents a plurality of information monomers.

2. The method of claim 1 wherein:a) the information that is partitioned in (a) is binary information; and / orb) the informative polymer sequence of (b) was generated by encoding information in a plurality of information symbols to generate the informative polymer sequence, optionally wherein the information that is encoded is binary information.

3. The method of claim 1 or 2 wherein where the dPIT is for incorporation into a biological asset, the method comprises the step of obtaining a partition spectrum for said biological asset wherein said partition profile describes the frequency of available positions for insertion of discrete polymer fragments of different lengths into one or more corresponding polymers of said biological asset, optionally wherein the partition spectrum describes the frequency of available positions for insertion of discrete polymer fragments of between: 0-49, 50-99, 100-149, 150-199, 200-249, 250-299, 300-349, 350-399, 400-449, 450-499, 500-549, 550-599, 600-649, 650-699, 700-749, 750-799, 800-849, 850-899, 900-949, 950-999, 1000-1049, 1050-1099, 1100-1149,1150-1199, 1200-1249, 1250-1299, 1300-1349, 1350-1399, 1400-1449, 1450-1499, 1500-1549, 1550-1599, 1600-1649, 1650-1699, 1700-1749, 1750-1799, 1800-1849,1850-1899, 1900-1949, 1950-1999, 2000-2049, 2050-2099, 2100-2149, 2150-2199, 2200-2249, 2250-2299, 2300-2349, 2350-2399, 2400-2449, 2450-2499, 2500-2549,2550-2599,2600-2649, 2650-2699, 2700-2749, 2750-2799, 2800-2849, 2850-2899, 2900-2949, 2950-2999, 3000-3049, 3050-3099, 3100-3149, 3150-3199, 3200-3249, 3250-3299, 3300-3349, 3350-3399, 3400-3449, 3450-3499, 3500-3549, 3550-3599,3600-3649, 3650-3699, 3700-3749, 3750-3799, 3800-3849, 3850-3899, 3900-3949, 3950-3999, 4000-4049, 4050-4099, 4100-4149, 4150-4199, 4200-4249, 4250-4299,4300-4349, 4350-4399, 4400-4449, 4450-4499, 4500-4549, 4550-4599, 4600-4649,4650-4699, 4700-4749, 4750-4799, 4800-4849, 4850-4899, 4900-4949, 4950-4999, 5000-5049, 5050-5099, 5100-5149, 5150-5199, 5200-5249, 5250-5299, 5300-5349, 5350-5399, 5400-5449, 5450-5499, 5500-5549, 5550-5599, 5600-5649 monomers in length,optionally wherein the biological asset is selected from the group comprising or consisting of:a) one or more chromosomes of a cell, optionally:i) a prokaryotic cell, optionally a bacterial cell;ii) a eukaryotic cell, optionally a yeast cell, a fungal cell, a mammalian cell, a plant cell;b) A circular double-stranded DNA structure, optionally a plasmidc) A circular single-stranded DNA structured) A nanostructure, wherein the nanostructure comprises at least one DNA or RNA or protein or peptide molecule, optionally wherein the at least one DNA molecule is single-stranded or double-stranded and / or wherein the at least one RNA molecule is single-stranded or double-stranded;e) a phagemidf) a phageg) a virus, optionally a viral vector, optionally an AAV or lentiviral vector;h) a nanoparticle, for example a lipid nanoparticle, vesicle, micelle, etc, for example a nanoparticle that is for use as a vaccine.

4. The method of claim 3 wherein:a) where the method comprises partitioning information into at least two discrete information subsets, the length of and number of discrete information subsets is chosen based on the partition spectrum;b) where the method comprises providing an informative polymer sequence (IPS) and partitioning the IPS into at least two discrete sub-IPSs sequences, the length of and number of discrete sub-IPSs is chosen based on the partition spectrum; and / orc) where the method comprises providing a plurality of discrete sub-informative polymer sequences (sub-IPS), the length of and number of discrete sub-IPS is chosen based on the partition spectrumsuch that polymeric fragments corresponding to said discrete sub-IPS into said biological asset would be accommodated within the biological asset.

5. The method of any of claims 1-4 wherein said generating and / or partitioning and / or providing is in silico and said sequences are digital sequences.

6. The method of any of claims 1-5, wherein the method comprises generating a plurality of redundant symbols by processing the information symbols present in the at least two discrete sub-informative polymer sequences (sub-IPS) with an error correction coding process, optionally wherein the error correcting coding process is selected from any one or more or combination of: Bose-Chaudhuri-Hocquenghem; Reed-Solomon codes; Hamming Codes: Linear block codes over GF(2); LDPC Codes: Linear block codes over GF(2); and or Convolutional Codes: Codes generated by convolution of input bits; non-binary cyclic code over finite fields (GF(q), typically 2m).

7. The method of claim 6 further comprising combining the plurality of redundant symbols with the informative symbols, optionally wherein at least one or more or all of the at least two discrete sub-informative polymer sequences (sub-IPS) comprises redundant symbols to form at least one or more error-correcting discrete sub-IPS.

8. The method of any of claim 6 or 7 wherein:a) the error correction coding process comprises an intra-partition error correcting code process and is applied to the at least two discrete sub-informative polymer sequences (sub-IPS) to generate for each discrete sub-informative polymer sequences (sub-IPS) an additional sequence that is a redundant sequence comprising a plurality of redundant symbols that represent a plurality of redundant monomers for use in error correction of the information monomers; andb) combining plurality of redundant symbols with the informative symbols comprises combining each redundant sequence comprising the plurality of redundant symbols with the corresponding discrete sub-informative polymer sequences (sub-IPS) to form a corresponding error-correcting discrete sub-IPSs.

9. The method of any of claims 1-8 comprising the step of applying an interpartition error correction coding process to the at least two discrete sub-informative polymer sequences (sub-IPS) or to the at least one of the error-correcting discrete sub-IPSs to generate at least one discrete redundant polymer sequence comprising a plurality of monomers that are redundant.

10. The method of any of claims 1-9 further comprising introducing one or more characteristic symbols into one or more of the at least two discrete sub-informative polymer sequences (sub-IPS) and / or to at least one of the error-correcting discrete sub-IPSs, and / or to the at least one discrete redundant polymer sequences so as to:a) impart desirable sequence characteristics on the at least one or more of the at least two discrete part-informative biological polymer sequence; and / orb) remove undesirable sequence characteristics from the at least one or more of the at least two discrete part-informative biological polymer sequence;optionallywherein desirable sequence characteristics are characteristics are selected from: characteristics that camouflage the discrete part-informative sequence with respect to the target site in the one or more target biological polymers, balanced GC content, a secondary structure that is a hairpin, a secondary structure that is a kissing loop, or a sequence that supports genetic engineering; and / orwherein undesirable sequence characteristics are selected from: characteristics that are readily distinguishable from the surrounding sequence of the target site in the one or more target biological polymers, endonuclease recognition sites such as restriction endonuclease sites, greater than X% GC content, sequences prone to secondary structure formation, a sequence that may interfere with nearby gene expression.

11. The method of any of claims 1-10 whereina) the at least one error-correcting discrete sub-IPS and the at least one discrete redundant polymer sequence together collectively form the distributed polymeric information tag (dPIT);b) the at least two discrete sub-informative polymer sequences (sub-IPS) together collectively form the distributed polymeric information tag (dPIT);c) the error-correcting discrete sub-IPSs together collectively form the distributed polymeric information tag (dPIT); ord) the at least two discrete sub-informative polymer sequences (sub-IPS) and the at least one discrete redundant polymer sequence together collectively form the distributed polymeric information tag (dPIT).

12. The method of any of claims 8-11 wherein:a) the intra-partition error correction coding process comprises: a Bose-Chaudhuri-Hocquenghem error correction coding process; Hamming Codes: Linear block codes over GF(2); LDPC Codes: Linear block codes over GF(2); and or Convolutional Codes: Codes generated by convolution of input bits;and / orb) the inter-partition error correction coding process comprises a Reed-Solomon error correction coding process or non-binary cyclic code over finite fields (GF(q), typically 2m),13 The method of any of claims 1-12 wherein the polymer is selected from the group comprising or consisting of: Deoxyribonucleic acid (DNA), Ribonucleic acid (RNA), Protein or peptide or polypeptide, Peptide Nucleic Acids (PNAs), Polypeptides, Polysaccharides, Polyethylene Glycol (PEG), Polylactic Acid (PLA), Poly(ethylene terephthalate) (PET), Poly(methyl methacrylate) (PMMA), Polyaniline (PANI), Polycarbonate (PC), Polystyrene (PS), Poly(vinyl alcohol) (PVA), Polylactide-co-glycolide (PLGA), Polydimethylsiloxane (PDMS), Polyurethanes, Polycaprolactone (PCL), Polyhydroxyalkanoates (PHAs), optionallyWhere the polymer is DNA the monomer is selected from the group comprising or consisting of: Adenine (A), Thymine (T), Cytosine (C), Guanine (G), or a non-natural nucleotide, optionally wherein the nucleotide is a modified nucleotide;Where the polymer is RNA the monomer is selected from the group comprising or consisting of: Adenine (A), Uracil (U), Cytosine (C), Guanine (G), or a non-natural nucleotide, optionally wherein the nucleotide is a modified nucleotide;Wherein the polymer is a protein or peptide or polypeptide the monomer is selected from the group comprising or consisting of: an Amino acids optionally alanine (A), arginine (R), asparagine (N), aspartic acid (D), cysteine (C), glutamic acid (E), glutamine (Q), glycine (G), histidine (H), isoleucine (I), leucine (L), lysine (K), methionine (M), phenylalanine (F), proline (P), serine (S), threonine (T), tryptophan (W), tyrosine (Y), and valine (V), or a non-natural amino acid, optionally wherein the amino acid is a modified amino acid.Where the polymer is a PNA the monomer is selected from the group comprising or consisting of: Adenine (A), Thymine (T), Cytosine (C), Guanine (G), or a non-natural nucleotide, optionally wherein the nucleotide is a modified nucleotide;Where the polymer is a polysaccharide the monomer is selected from the group comprising or consisting of: Glucose, Fructose, Galactose, Mannose;Where the polymer is PEG the monomer is selected from the group comprising or consisting of: Ethylene glycol, Propylene glycol, Butylene glycol, Diethylene glycol;Where the polymer is PLA the monomer is selected from the group comprising or consisting of: L-lactic acid, D-lactic acid, Mesolactic acid, 2-Hydroxypropanoic acid;Where the polymer is PET the monomer is selected from the group comprising or consisting of: Ethylene glycol, Terephthalic acid, Isophthalic acid, Naphthalene dicarboxylate;Where the polymer is PMMA the monomer is selected from the group comprising or consisting of: Methyl methacrylate, Ethyl methacrylate, Butyl methacrylate, 2-Hydroxyethyl methacrylate;Where the polymer is PANI the monomer is selected from the group comprising or consisting of: Aniline, o-Toluidine, m-Toluidine, p-Toluidine;Where the polymer is PC the monomer is selected from the group comprising or consisting of: Bisphenol A, Bisphenol S, Cyclohexane dimethanol, Terephthalic acid;Where the polymer is PS the monomer is selected from the group comprising or consisting of: Styrene, a-Methylstyrene, Butadiene, Acrylonitrile;Where the polymer is PVA the monomer is selected from the group comprising or consisting of: Vinyl acetate, Vinyl alcohol, Ethylene, Acrylamide;Where the polymer is PLGA the monomer is selected from the group comprising or consisting of: L-lactic acid, D-lactic acid, Glycolic acid, 2-Hydroxypropanoic acid;Where the polymer is PDMS the monomer is selected from the group comprising or consisting of: Dimethylsiloxane, Diphenylsiloxane, Methylphenylsiloxane, Vinylsiloxane;Where the polymer is polyurethane the monomer is selected from the group comprising or consisting of: Toluene diisocyanate, Methylene diphenyl diisocyanate, Hexamethylene diisocyanate, Polycaprolactone polyol;Where the polymer is PCL the monomer is selected from the group comprising or consisting of: E-Caprolactone, 6-Valerolactone, y-Butyrolactone, L-Lactide;Where the polymer is PHA the monomer is selected from the group comprising or consisting of: 3-Hydroxybutyrate, 3-Hydroxyvalerate, 3-Hydroxyhexanoate, 4-Hydroxybutyrate.

14. The method of any of the preceding claims comprising producing one or more polymers corresponding to the sequences set out in the at least two discrete sub-informative polymer sequences (sub-IPS) and / or the at least one of the errorcorrecting discrete sub-IPSs, and / or the at least one discrete redundant polymer sequences so as to produce one or more or all of:a) at least one error-correcting discrete sub-IPS polymer corresponding to the at least one error-correcting discrete sub-IPS sequence;b) at least one discrete redundant polymer corresponding to the at least one discrete redundant polymer sequence;c) at least one sub-informative polymer corresponding to at least one the discrete sub-informative polymer sequenceto produce a collection of polymeric information tag polymersoptionally wherein the collection of polymeric information tag polymers comprises:at least one error-correcting discrete sub-IPS polymer and at least one discrete redundant polymer.

15. The method of claim 14 further comprising the step of applying, inserting or otherwise causing at least one of the polymers of the collection of the polymeric information tag polymers to be physically located in or on one or more assets.

16. The method of claim 15 wherein the asset is a biological asset and wherein the biological asset comprises one or more target polymers corresponding to the type of polymer of the collection of polymeric information tag polymers.

17. The method of claim 16 wherein the target polymers are selected from the group comprising or consisting of:a) one or more chromosomes of a cell, optionally:i) a prokaryotic cell, optionally a bacterial cell;ii) a eukaryotic cell, optionally a yeast cell, a fungal cell, a mammalian cell, a plant cell;b) A circular double-stranded DNA structure, optionally a plasmidc) A circular single-stranded DNA structured) A nanostructure, wherein the nanostructure comprises at least one DNA or RNA or protein or peptide molecule, optionally wherein the at least one DNA molecule is single-stranded or double-stranded and / or wherein the at least one RNA molecule is single-stranded or double-stranded;e) a phagemidf) a phageg) a virus, optionally a viral vector, optionally an AAV or lentiviral vector;h) one or more chromosomes or plasmids of separate cells in a cell culture;i) a polymer comprised by nanoparticle, for example a lipid nanoparticle, vesicle, micelle, etc, for example a nanoparticle that is for use as a vaccine.

18. The method of any of claims 15-17 wherein the method comprises the step of inserting or otherwise causing more than one polymer of the collection of polymeric information tag polymers to be physically located in one or more of the target biological polymers, optionally to be physically located in:a) The genome of the same cellb) The same circular double-stranded DNA structure, optionally a plasmidc) The same circular single-stranded DNA structured) The same nanostructure, wherein the nanostructure comprises at least one DNA or RNA or amino acid polymer molecule, optionally wherein the at least one DNA molecule is single-stranded or double-stranded and / or wherein the at least one RNA molecule is single-stranded or double-stranded;e) The same phagemidf) The same phageg) The same virus, optionally a viral vector, optionally an AAV or lentiviral vector h) members of a cell population, optionally one or more chromosomes or plasmids of separate cells in a cell population;i) the same nanoparticle, for example a lipid nanoparticle, vesicle, micelle, etc, for example a nanoparticle that is for use as a vaccine;optionally wherein all of the polymers of the collection of polymeric information tag polymers are physically located in one or more of the target biological polymers.

19. The method of any of claims 16-18 wherein at least one of the polymers of the collection of polymeric information tag polymers is not embedded in the same asset, optionally not embedded in the same one or more target biological polymers as the rest of the polymers of the collection of polymeric information tag polymers,optionally wherein:at least one of the polymers of the collection of polymeric information tag polymers is not embeddeda) in the genome of the same cell;b) in the same circular double-stranded DNA structure;c) in the same circular single-stranded DNA structure;d) in the same nanostructure;e) in the same phagemid;f) in the same phage;g) in the same virus, optionally not in the same viral vector, optionally not in the same AAV or lentiviral vector;h) in members of the same cell population;i) in the same nanoparticle, for example a lipid nanoparticle, vesicle, micelle, etc, for example a nanoparticle that is for use as a vaccine.as the rest of the polymers of the collection of polymeric information tag polymers.

20. A computer implemented method for generating a distributed polymeric information tag (dPIT) for incorporation into or application to one or more assets, the method comprising the method of any one or more of claims 1-13.

21. A digital distributed polymeric information tag (dPIT) produced according to any of the preceding methods.

22. A physical distributed polymeric information tag (dPIT) comprising the collection of polymeric information tag polymers of any of the preceding claims.

23. A method of labelling an asset with a distributed polymeric information tag (dPIT), said method comprising applying one or more of the polymers of the distributed polymeric information tag (dPIT) of claim 22 to said asset.

24. A biological asset comprising at least part of the distributed polymeric information tag (dPIT) or at least one or more polymers of the collection of polymeric information tag polymers as defined in any of the preceding claims.

25. The biological asset of claim 24 wherein the biological asset does not comprise all of the polymers of the collection of polymeric information tag polymers.

26. A method of producing one or more polymers corresponding to the sequences set out in any of the at least two discrete sub-informative polymer sequences (sub-IPS) and / or the at least one of the error-correcting discrete sub-IPSs, and / or the at least one discrete redundant polymer sequences generated according to the method of any of the preceding claims so as to produce one or more or all of:a) at least one error-correcting discrete sub-IPS polymer corresponding to the at least one error-correcting discrete sub-IPS sequence;b) at least one discrete redundant polymer corresponding to the at least one discrete redundant polymer sequence;c) at least one sub-informative polymer corresponding to at least one the discrete sub-informative polymer sequenceto produce a collection of polymeric information tag polymersoptionally wherein the collection of polymeric information tag polymers comprises:at least one error-correcting discrete sub-IPS polymer and at least one discrete redundant polymer.

Citation Information

Patent Citations

  • Systems for nucleic acid-based data storage

    US20240013063A1

  • Probes comprising a split barcode region and methods of use

    WO2023023484A1