Methods and compositions for providing identification and / or traceability of biological materials

Introducing unique DNA identifier sequences into biological entities' genomes facilitates rapid and reliable traceability, addressing the limitations of current food traceability systems by enabling swift identification of contamination sources and enhancing supply chain transparency.

JP7803542B2Active Publication Date: 2026-01-21INDEX BIOSYSTEMS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2022556695
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-11-26
Filing Date
2020-11-26
Publication Date
2026-01-21
Estimated Expiration
2040-11-26

AI Technical Summary

Technical Problem

Current food traceability systems are inadequate for quickly and affordably identifying the source of biological contaminants, particularly in cases of clonally propagated products and mixed items, leading to significant health risks and economic losses.

Method used

Introduce unique DNA identifier sequences into the genome of biological entities, allowing for rapid identification and traceability through sequencing and database matching, using oligonucleotide constructs and cassettes with primer annealing sequences for amplification and sequencing.

Benefits of technology

Enables rapid and reliable traceability of biological materials, reducing the time to identify contamination sources and enhancing supply chain transparency, thereby improving public health and economic efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007803542000025
    Figure 0007803542000025
  • Figure 0007803542000026
    Figure 0007803542000026
  • Figure 0007803542000027
    Figure 0007803542000027
Patent Text Reader

Abstract

[0003] Provided herein are methods and compositions for providing identification and / or traceability of biological material. In certain embodiments, the methods include determining the sequence of at least one unique identifier sequence in the genomic DNA of a biological entity, verifying the identity of the biological entity by establishing the presence of the unique identifier sequence in the genomic DNA and comparing the sequence of the unique identifier sequence to a database to confirm uniqueness, providing an indication of the validity of production of the biological material from the biological entity, and entering the unique identifier sequence into a database entry in the database and associating the unique identifier sequence with identification and / or tracking information, thereby providing traceability by reading the unique identifier sequence and retrieving the corresponding database entry to obtain the identification and / or tracking information. Also provided are oligonucleotides, cassettes, and compositions for providing identification and / or traceability of biological material.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates generally to the identification and / or tracking of biological substances. More particularly, the present invention relates to methods and agents for the identification and / or tracking of biological substances using nucleic acids. [Background technology]

[0002] Food system efficiency and output have reached unprecedented levels. While this development has brought significant public benefits in the form of cost savings and diversity, significant gaps remain that pose risks to public health, industry, and innovation. Traceability is one of the key technologies for efficiently managing and navigating these challenges.

[0003] The limitations of current food and beverage traceability systems are primarily exposed by contamination events. When such events occur, it can take months to trace affected products back to their source-of-origin. Clonally propagated products can pose additional challenges to source identification due to their lack of genetic variation. Additionally, because existing traceability best practices require adherence throughout the supply chain, converted and mixed products can pose problems for source identification. Failures in the ability to quickly and affordably trace such products pose significant risks to consumer safety, cause significant economic losses for stakeholders, and ultimately severe damage to the reputations of affected industries.

[0004] In 2015, the World Health Organization (WHO) completed a 10-year initiative to estimate the global burden of foodborne diseases. This initiative found that "...the global burden of foodborne diseases...was 33 million DALYs (95% UI, 25-46 million DALYs) in 2010, with 40% of the foodborne disease burden occurring in children under the age of 5" (p. 11) (WHO, Foodborne Disease Burden Epidemiology Reference Group. 2015. WHO Estimates of the Global Burden of Foodborne Diseases. Retrieved from https: / / academicanswers.waldenu.edu / faq / 73164). DALY stands for Disability-Adjusted Life Years. It can be thought of as years of healthy life lost. The estimates produced by this study were limited by data gaps. Improvements in surveillance, monitoring, and laboratory capacity were recognized as being needed for more accurate estimates. The need for surveillance was further highlighted by the Source Attribution Task Force (SATF).

[0005] The SATF is one of a number of expert task forces commissioned for this initiative. Their mandate was to estimate the effect of specific attribution points on disease transmission. Figure 1 (adapted from WHO, 2015, p. 101) illustrates the main attribution points (WHO, Foodborne Disease Burden Epidemiology Reference Group. 2015. WHO Estimates of the Global Burden of Foodborne Diseases. Retrieved from https: / / academicanswers.waldenu.edu / faq / 73164). The reference group for this study, FERG, determined that for the purposes of this study, the simplest attribution point was the end of the transmission chain: human contact. This simplicity characterizes the limitations of existing traceability practices. FERG also acknowledges that other points of attribution, such as primary production, may be more appropriate for risk management (p. 100) (WHO, Foodborne Disease Burden Epidemiology Reference Group. 2015. WHO Estimates of the Global Burden of Foodborne Diseases. Retrieved from https: / / academicanswers.waldenu.edu / faq / 73164). FERG identifies surveillance for repository-level attribution as desirable.

[0006] Modern techniques for food traceability in the food and beverage supply chain typically begin with the grower's harvest or within the production facility. Produce is often tracked at the case level, with a case containing many items. A physical barcode may be affixed to each item. A Global Trade Item Number (GTIN) and Global Location Number (GLN) are ideally associated with the case. A serial shipping container code (SSCC) may be generated for each pallet, i.e., collection of cases. Such traceability techniques are typically governed by standards, which for perishable foods are often GS1 standards. To progress through the supply chain with pallets, the above-mentioned identifiers found on the barcode are used in conjunction with key data elements (KDEs) recorded for critical tracking events (CTEs). The CTE may document the disposition of the produce from the grower to the packer / shipper. There is a commonly used adage, "one-step forward and one-step back," which suggests that each supply chain participant should be able to trace the produce. Unfortunately, this requirement has proven inadequate in many ways.

[0007] Once a food item reaches the point of sale, it may be converted or commingled with other items from different producers, such as fruit salad. Often, once the item is separated from its original case or item-level identifier, it is impossible to trace the item back to the producer. As seen in the recent romaine lettuce outbreak, lack of provenance information meant that it took investigators more than a month to identify the source of contamination, even though most of the produce is grown in the southwestern United States (FDA, 2019, p. 1) (FDA. (2019). Investigation Summary: Factors Potentially Contributing to the Contamination of Romaine Lettuce Implicated in the Fall 2018 Multi-State Outbreak of E. coli O157:H7. Retrieved from https: / / www.fda.gov / media / 120722 / download). As a result, the FDA urged "...the entire leafy greens supply chain to adopt traceability best practices and cutting-edge technologies to ensure fast, accurate, and easy access to key data elements from farm to fork if leafy greens are involved in a potential recall or infectious disease outbreak" (FDA, 2019, p. 8) (FDA. (2019). Investigation Summary: Factors Potentially Contributing to the Contamination of Romaine Lettuce Implicated in the Fall 2018 Multi-State Outbreak of E. coli O157:H7. Retrieved from https: / / www.fda.gov / media / 120722 / download ). The costs associated with this infectious disease outbreak have yet to be reimbursed. However, other contamination events are well known.

[0008] Since 2006, spinach recalls have been linked to five deaths and approximately 200 cases of fatal illness in 26 states, resulting in an economic loss of approximately $500 million (GS1, 2013, p. 3) (GS1 US. (2013). Integrated Traceability in Fresh Foods: Ripe Opportunity for Real Results. Retrieved from https: / / www.gs1us.org / DesktopModules / Bring2mind / DMX / Download.aspx?Command=Core_Download&EntryId=598). More generally, "...government agencies have also expressed concern about the health and economic impacts of recent food recalls, as foodborne illnesses affect 48 million people annually and cost the United States $152 billion in healthcare costs annually" (GS1, 2013, p. 2) (GS1 US. (2013). Integrated Traceability in Fresh Foods: Ripe Opportunity for Real Results. Retrieved from https: / / www.gs1us.org / DesktopModules / Bring2mind / DMX / Download.aspx?Command=Core_Download&EntryId=598). Full-chain traceability, which can be understood as tracking from seed to sale, was found to reduce the total amount of recalled produce by 12%, in the case of the Frontera Produce coriander recall.McKinsey found that a 25% improvement in recall accuracy could save the fresh food industry between $250 million and $275 million annually (GS1, 2013, p. 10) (GS1 US. (2013). Integrated Traceability in Fresh Foods: Ripe Opportunity for Real Results. Retrieved from https: / / www.gs1us.org / DesktopModules / Bring2mind / DMX / Download.aspx?Command=Core_Download&EntryId=598).

[0009] Whole-chain traceability has lacked an efficient form of item-level identification and lacked assurance of origin. Existing methods for item-level identification typically rely on physical branding (laser), radio frequency identification (RFID), and barcoding, i.e., external physical identifiers. There are still scaling challenges associated with such technologies. Each item requires a physical identifier, and there are costs associated with producing this. Additionally, there is an inherent risk of read errors and / or malicious tampering with their use, e.g., the risk of a sticker peeling off or being removed.

[0010] Food contamination, e.g., E. coli and / or Salmonella contamination affecting the food supply, is a threat to public health, and rapid action to identify and contain the source(s) of contamination is highly desirable. There is a long-standing unmet need in the art for reliable, cost-effective, and / or rapid strategies to enhance product traceability in the food supply. Traceability of biological entities and / or biological materials is not only desirable in the agricultural and food industries, but is also required in a wide variety of industries and sectors that handle biological entities and / or biological materials that contain or are derived from them.

[0011] Alternative, additional and / or improved methods and / or compositions for providing identification and / or traceability of biological entities and / or biological materials are desirable. [Prior art documents] [Non-patent literature]

[0012] [Non-Patent Document 1] WHO, Foodborne Disease Burden Epidemiology Reference Group. 2015. WHO Estimates of the Global Burden of Foodborne Diseases. Retrieved from https: / / academicanswers.waldenu.edu / faq / 73164 [Non-patent document 2] FDA. (2019). Investigation Summary: Factors Potentially Contributing to the Contamination of Romaine Lettuce Implicated in the Fall 2018 Multi-State Outbreak of E. coli O157:H7. Retrieved from https: / / www.fda.gov / media / 120722 / download [Non-patent document 3] GS1 US. (2013). Integrated Traceability in Fresh Foods: Ripe Opportunity for Real Results. Retrieved from https: / / www.gs1us.org / DesktopModules / Bring2mind / DMX / Download.aspx?Command=Core_Download&EntryId=598 Summary of the Invention [Means for solving the problem]

[0013] Provided herein are methods and compositions for providing identification and / or traceability of biological materials. In certain embodiments, the methods described herein may utilize unique identifier sequences (also referred to herein as DNA unique identifier sequences) exogenously introduced into the genome of a biological entity to provide identification and / or traceability of the biological entity and / or material, including biological entities and / or materials produced from and containing genomic DNA derived therefrom. In certain embodiments, the unique identifier sequences may be derived from a randomized pool of sequences. In certain embodiments, a database may maintain the association of the unique identifier sequences with corresponding identification and / or tracking information. Also provided herein are oligonucleotide constructs and cassettes comprising one or more unique identifier sequences for use in providing identification and / or traceability of biological materials. In certain embodiments, the oligonucleotide constructs and / or cassettes may include specific sequences for primer annealing sequence(s), which may be for amplification of the unique identifier sequence(s), sequencing of the unique identifier sequence(s), or both. In certain embodiments, the methods and compositions described herein may be used to provide food traceability, which may allow for rapid response and / or food recall in the event of contamination, for example.

[0014] In embodiments, a method for identifying a biological substance includes: receiving or providing a sample containing genomic DNA from a biological material; amplifying at least one DNA-specific identifier sequence in genomic DNA from the biological material and sequencing the DNA-specific identifier sequence; searching a database for the DNA unique identifier sequence and obtaining a database entry corresponding to the DNA unique identifier sequence, the database entry providing identification and / or tracking information for the biological material; A method is provided herein, comprising:

[0015] In another embodiment of the above method, the biological material may include plant-based material, fungus-based material, animal-based material, virus-based material, or bacterial-based material.

[0016] In certain embodiments, the biological material may include fungal-derived material. In certain embodiments, the biological material may include yeast. In certain embodiments, the yeast may be sporulated (i.e., the biological material may include yeast spores). In certain embodiments, the yeast may be added to, mixed with, or otherwise associated with a product whose identification and / or tracking is desirable, such as a food ingredient or food product.

[0017] In another embodiment, there is provided a method for providing traceability of a biological material, comprising: determining the sequence of at least one DNA-specific identifier sequence within the genomic DNA of the biological entity; verifying the identity of the biological entity by establishing the presence of the DNA-specific identifier sequence in the genomic DNA and comparing the sequence of the DNA-specific identifier sequence with a database to confirm that the DNA-specific identifier sequence is not already used in the database; providing an indication of the advisability of producing a biological material from the biological entity, the biological material comprising genomic DNA from the biological entity; entering the sequence of the at least one DNA unique identifier sequence into a database entry in a database, and associating the DNA unique identifier sequence with identification and / or tracking information for the biological material; whereby providing traceability of a biological material by reading a DNA unique identifier sequence in the biological material and retrieving a corresponding database entry to provide identification and / or tracking information for the biological material.

[0018] In another embodiment of the above method, the method may further comprise inserting at least one DNA-specific identifier sequence into the genomic DNA of the biological entity, or modifying by gene editing an identifier sequence already present in the genomic DNA of the biological entity to generate a DNA-specific identifier sequence in the genomic DNA of the biological entity, thereby achieving this identification.

[0019] In yet another embodiment of any of one or more of the above methods, the method may further comprise the step of providing at least one DNA-specific identifier sequence for insertion into the genomic DNA of the biological entity.

[0020] In yet another embodiment of any one or more of the above methods, the biological material may include plant-derived material, fungal-derived material, animal-derived material, viral-derived material, or bacterial-derived material.

[0021] In yet another embodiment of any one or more of the above methods, the biological entity may comprise a plant cell, a fungal cell, an animal cell, a virus, or a bacterial cell.

[0022] In another embodiment of any one or more of the above methods, the biological material, biological entity, or both may comprise fungal material or fungal cells. In certain embodiments, the biological material, biological entity, or both may comprise yeast. In certain embodiments, the yeast may be sporulated (i.e., may comprise yeast spores).

[0023] In yet another embodiment of any of one or more of the above methods, producing the biological substance from the biological entity may include growing the biological entity.

[0024] In another embodiment of any one or more of the above methods, the DNA-specific identifier sequence may be derived from a randomized pool of DNA-specific identifier sequences.

[0025] In yet another embodiment of any one or more of the above methods, the step of reading the DNA unique identifier sequence in the biological material and obtaining the corresponding database entry comprises: receiving or providing a sample containing genomic DNA from a biological material; amplifying at least one DNA-specific identifier sequence in genomic DNA from the biological material and sequencing the DNA-specific identifier sequence; comparing the DNA unique identifier sequence to a database and obtaining a database entry corresponding to the DNA unique identifier sequence, the database entry providing identification and / or tracking information for the biological material; may include:

[0026] In yet another embodiment of any one or more of the above methods, the DNA unique identifier sequence may comprise a unique nucleotide sequence inserted within an intergenic region of the genomic DNA.

[0027] In yet another embodiment of any of one or more of the above methods, the DNA unique identifier sequence may comprise a sequence up to about 1500 nt in length, up to about 1000 nt in length, from about 200 nt to about 600 nt in length, from about 200 nt to about 400 nt in length, or from about 400 nt to about 600 nt in length.

[0028] In another embodiment of any of one or more of the above methods, the DNA-specific identifier sequence may be flanked by one or more primer annealing sequences for PCR amplification of the DNA-specific identifier sequence, sequencing of the DNA-specific identifier sequence, or both.

[0029] In yet another embodiment of any one or more of the above methods, the biological material may include a food product.

[0030] In yet another embodiment of any of one or more of the above methods, the database entry identification and / or tracking information may include supply chain information for the biological material. In certain embodiments, the supply chain information may include supply chain information for food, agriculture, pharmaceutical, retail, textile, goods, chemicals, or other supply chain items with which the biological material may be associated.

[0031] In another embodiment of any one or more of the above methods, the identification and / or tracking information of the database entry may include source information of the biological material.

[0032] In yet another embodiment of any of one or more of the above methods, the identification and / or tracking information of the database entry may include grower, region, batch, lot, date, or other relevant supply chain information, or any combination thereof.

[0033] In yet another embodiment of any one or more of the above methods, a cassette may be incorporated into the genomic DNA that may include a DNA-specific identifier sequence flanked by one or more primer annealing sequences for PCR amplification of the DNA-specific identifier sequence, sequencing of the DNA-specific identifier sequence, or both.

[0034] In another embodiment of any of one or more of the above methods, the DNA unique identifier sequence can be a random sequence derived from a randomized pool of nucleic acid sequences up to about 1500 nt in length, up to about 1000 nt in length, from about 200 nt to about 600 nt in length, from about 200 nt to about 400 nt in length, or from about 400 nt to about 600 nt in length.

[0035] In another embodiment, provided herein are oligonucleotides comprising a DNA-specific identifier sequence flanked by one or more primer annealing sequences for PCR amplification of the DNA-specific identifier sequence, sequencing of the DNA-specific identifier sequence, or both.

[0036] In another embodiment of the above oligonucleotide, the DNA-specific identifier sequence may comprise a random sequence up to about 1500 nt in length, up to about 1000 nt in length, from about 200 nt to about 600 nt in length, from about 200 nt to about 400 nt in length, or from about 400 nt to about 600 nt in length.

[0037] In another embodiment, provided herein is a cassette comprising any one or more of the oligonucleotides described herein.

[0038] In yet another embodiment, provided herein is a cell or virus that has integrated into its genome any one or more of the oligonucleotides described herein or any one or more of the cassettes described herein.

[0039] In another embodiment, provided herein is a cell or virus that has incorporated a DNA unique identifier sequence into its genome.

[0040] In another embodiment of any of the above cells or viruses, the DNA-specific identifier sequence may be incorporated into an intergenic region of the genomic DNA of the cell or virus.

[0041] In yet another embodiment of any of the above cells or viruses, the cell may be a plant cell, a fungal cell, an animal cell, or a bacterial cell.

[0042] In another embodiment, the cell may be a fungal cell, for example, a yeast cell.

[0043] In another embodiment, DNA-specific identifier sequence, a randomized pool of DNA-specific identifier sequences; Any one or more of the oligonucleotides described herein; Any one or more of the cassettes described herein; one or more primer pairs for amplifying and / or sequencing the DNA unique identifier sequence; buffer solution, polymerase, or Instructions for carrying out any one or more of the methods described herein Kits are provided herein that include any one or more of the following:

[0044] In another embodiment, there is provided a method for identifying a biological substance, comprising: receiving, at a computing device, a DNA unique identifier sequence (DUID) extracted from a known biological material; searching a DUID database, which stores a plurality of DUIDs associated with respective biological material information, in the computing device to match the received DUID; if the search of the DUID database fails to match the received DUID, storing the received DUID in the DUID database in association with biological material information associated with known biological materials; receiving, at the computing device after storing the received DUID and information associated with the known biological material in a DUID database, a query DUID extracted from the unknown biological material; searching a DUID database at the computing device to match the received query DUID; if the search for the DUID results in a match with the received query DUID, returning the stored biological information associated with the DUID that matches the query DUID in response to the received query DUID; A method is provided herein, comprising:

[0045] In another embodiment of the above method, the step of searching the DUID database to match the received DUID comprises: searching a DUID database for an exact match with the received DUID; if no exact match is found, performing an alignment / identity search against DUIDs stored in the DUID database that closely match the received DUID; may include:

[0046] In yet another embodiment of any one or more of the above methods, the step of searching the DUID database to match the query DUID includes: searching the DUID database for an exact match with the query DUID; if no exact match is found, performing an alignment / identity search against DUIDs stored in the DUID database that closely match the query DUID; may include:

[0047] In yet another embodiment of any one or more of the above methods, the method further comprises: If the search results in a near match with the query DUID, storing the query DUID in association with the DUID that nearly matches the query DUID. It may further include:

[0048] In another embodiment, there is provided a computer system for identifying a biological substance, comprising: a processing unit capable of executing instructions; and a memory unit storing instructions that, when executed on a processing unit, configure a computer system to perform any one or more of the methods described herein; Provided herein is a computer system comprising:

[0049] In another embodiment, provided herein is a computer readable memory having stored thereon instructions that, when executed by a processing unit of a computer system, configure the system to perform any one or more of the methods described herein.

[0050] In another embodiment, there is provided a method for identifying a biological substance, comprising: receiving or providing a sample containing genomic DNA from a biological material; amplifying at least one DNA-specific identifier sequence in genomic DNA from the biological material and sequencing the DNA-specific identifier sequence; decoding or decoding the identification and / or tracking information of the biological material stored in the DNA unique identifier sequence; A method is provided herein, comprising:

[0051] In another embodiment, there is provided a method for providing traceability of a biological material, comprising: determining the sequence of at least one DNA-specific identifier sequence within the genomic DNA of the biological entity; verifying the identity of the biological entity by verifying the presence of the DNA-specific identifier sequence in the genomic DNA and decoding or decoding the identifying and / or tracking information stored in the DNA-specific identifier sequence to verify the DNA-specific identifier sequence; providing an indication of the adequacy of producing a biological material from the biological entity, the biological material comprising genomic DNA from the biological entity; whereby providing traceability of a biological material by reading a DNA unique identifier sequence in the biological material and decoding or deciphering the information stored in the DNA unique identifier sequence to provide identification and / or tracking information for the biological material.

[0052] In yet another embodiment, there is provided a method for identifying a biological substance, comprising: receiving, at a computing device, a DNA unique identifier sequence (DUID) extracted from an unknown biological material; decoding or decoding the identification and / or tracking information of the unknown biological material stored in the DNA unique identifier sequence; A method is provided herein, comprising:

[0053] In another embodiment, provided herein is a cassette comprising a DNA-specific identifier sequence, wherein the DNA-specific identifier sequence is flanked by at least one 5' primer annealing sequence and at least one 3' primer annealing sequence for amplification of the DNA-specific identifier sequence, sequencing of the DNA-specific identifier sequence, or both.

[0054] In another embodiment of the above cassette, the DNA-specific identifier sequence may be flanked by two 5' primer annealing sequences and two 3' primer annealing sequences to allow amplification of the DNA-specific identifier sequence by nested PCR.

[0055] In yet another embodiment of any of one or more of the above cassettes, the two 5' primer annealing sequences may partially overlap, the two 3' primer annealing sequences may partially overlap, or both.

[0056] In yet another embodiment of any of the one or more above cassettes, the cassette may further comprise a sequencing primer annealing sequence located 5' to the DNA-specific identifier sequence for sequencing the DNA-specific identifier sequence.

[0057] In yet another embodiment of any one or more of the above cassettes, a sequencing primer annealing sequence may be located between the two 5' primer annealing sequences.

[0058] In another embodiment of any one or more of the above cassettes, the sequencing primer annealing sequence may at least partially overlap with one or both of the two 5' primer annealing sequences.

[0059] In yet another embodiment of any of one or more of the above cassettes, the two 5' primer annealing sequences may partially overlap, and at least a portion of the sequencing primer annealing sequence may be located in the overlap.

[0060] In another embodiment of any of one or more of the above cassettes, the cassette sequence can be up to about 1500 nt in length, up to about 1000 nt in length, from about 200 nt to about 600 nt in length, from about 200 nt to about 400 nt in length, or from about 400 nt to about 600 nt in length.

[0061] In yet another embodiment of any one or more of the above cassettes, the primer annealing sequence may not naturally occur in the genome of the target biological entity.

[0062] In another embodiment, provided herein is a composition comprising a plurality of any one or more cassettes described herein, wherein each cassette comprises an identical primer annealing sequence, and each cassette comprises a randomized DNA-specific identifier sequence.

[0063] In yet another embodiment, provided herein is a composition comprising a plurality of any one or more cassettes described herein, wherein each cassette comprises an identical primer annealing sequence and an identical sequencing primer annealing sequence, and each cassette comprises a randomized DNA-specific identifier sequence.

[0064] In yet another embodiment, there is provided a method for providing traceability of biological material, comprising: Inserting at least one DNA-specific identifier sequence into the genomic DNA of a biological entity for use in the preparation of a biological material. A method is provided herein, comprising:

[0065] In another embodiment of the above method, the DNA unique identifier sequence may be inserted as any one or more of the cassettes described herein.

[0066] In another embodiment of any of one or more of the above methods, the method may further comprise determining the sequence of at least one DNA-specific identifier sequence within the genomic DNA of the biological entity.

[0067] In another embodiment of any of one or more of the above methods, the method may further comprise verifying the identity of the biological entity by establishing the presence of the DNA-specific identifier sequence in the genomic DNA and comparing the sequence of the DNA-specific identifier sequence with a database to confirm that the DNA-specific identifier sequence is not already used in the database.

[0068] In yet another embodiment of any one or more of the above methods, the method further comprises: Producing a biological material from a biological entity, the biological material comprising genomic DNA from the biological entity; and / or providing an indication of the advisability of producing a biological material from the biological entity, the biological material comprising genomic DNA from the biological entity; It may further include:

[0069] In yet another embodiment of any of one or more of the above methods, the method may further comprise entering the sequence of the at least one DNA-specific identifier sequence into a database entry and associating the DNA-specific identifier sequence with identification and / or tracking information for the biological entity and / or biological material.

[0070] In yet another embodiment of any one or more of the above methods, the method further comprises: providing traceability of the biological entity and / or biological substance by reading the DNA unique identifier sequence in the biological entity and / or biological substance and obtaining a corresponding database entry to provide identification and / or tracking information for the biological entity and / or biological substance. It may further include:

[0071] In another embodiment, provided herein is a plasmid or expression vector comprising any one or more of the oligonucleotides or one or more cassettes described herein.

[0072] In yet another embodiment, there is provided a method for providing traceability of a product of interest, comprising: Receiving or providing a sample from a product of interest, the sample comprising genomic DNA from biological material that is part of, mixed with, or otherwise associated with the product of interest; amplifying at least one DNA-specific identifier sequence in genomic DNA from the biological material and sequencing the DNA-specific identifier sequence; searching a database for the DNA unique identifier sequence and obtaining a database entry corresponding to the DNA unique identifier sequence, the database entry providing identification and / or tracking information for the product of interest; A method is provided herein, comprising:

[0073] In another embodiment of the above method, the method may include introducing or adding any one or more biological materials or one or more biological entities described herein to a product of interest, wherein the biological material or entity includes at least one DNA unique identifier sequence described herein as part of its genomic material.

[0074] In yet another embodiment of any of one or more of the above methods, the identification and / or tracking information of the database entry may include supply chain information for the product of interest.

[0075] In yet another embodiment of any of one or more of the above methods, the product of interest may include food, agricultural products, pharmaceuticals, retail products, textiles, goods, chemicals, or another supply chain item. [Brief explanation of the drawings]

[0076] These and other features will be better understood by consideration of the following description and accompanying drawings. [Figure 1] This is a diagram showing the transmission pathways identified by the World Health Organization (WHO) in its 2015 report (Source: WHO, 2015, p. 101). [Figure 2] FIG. 1 shows an example of a cassette described herein that includes a DUID sequence and its generation as described in Example 1. The sequence shown is SEQ ID NO:1. [Figure 3] FIG. 1 illustrates an example process overview of the DUID system described in Example 1. [Figure 4] FIG. 1 illustrates an example of the identification stage of the DUID system process described in Example 1. [Figure 5] FIG. 1 illustrates an example of the verification stage of the DUID system process described in Example 1. [Figure 6] FIG. 1 illustrates an example of the read stage of the DUID system process described in Example 1. [Figure 7] FIG. 1 illustrates another example of the DUID system and process described herein. [Figure 8] FIG. 1 illustrates another example of the DUID system and process described herein, which uses DUIDs and databases / registries to provide traceability of biological entities. [Figure 9]FIG. 10 illustrates yet another example of the DUID system and process described herein, which uses DUID sequences and databases / registries to obtain biological material identification and / or tracking information from a database. [Figure 10] FIG. 1 illustrates another example of a DUID system and process described herein that provides traceability of biological entities using a DUID that stores tracking and / or identification information. [Figure 11] FIG. 10 illustrates another example of the DUID system and process described herein for obtaining identification and / or tracking information for a biological substance using a DUID sequence that stores tracking and / or identification information. [Figure 12] FIG. 10 illustrates another example of the DUID system and process described herein for obtaining identification and / or tracking information for a biological substance using a DUID sequence that stores tracking and / or identification information. [Figure 13]

[0033] Figure 13 shows further examples of cassette designs described herein that include UID (unique identifier) ​​sequences: Figure 13(a) shows a dual primer design, 13(b) shows a single primer design, and 13(c) shows a stand-alone design. [Figure 14]

[0023] Figure 1 shows maps of two 370 pb DUID constructs described in Example 2. A) DUID construct design for PCR and qPCR amplification. The construct is 370 pb. This DUID construct contains two forward primers and two reverse primers. There are two identifiers (ID1 and ID2). ID1 is ideal for PCR amplification. ID2 is ideal for qPCR amplification. B) DUID construct design for loop-mediated isothermal amplification (LAMP) and PCR. This map contains primers for both PCR and LAMP. [Figure 15]Figure 1 shows detection of YCp-DUID in yeast genomic DNA by end-point PCR as described in Example 2. PCR amplification was performed using (A) the YCp-DUID vector, (B) gDNA extracted from BY4743, and (C) yeast strain BY4743 transformed with the YCp-DUID vector as template with DUID recall primers. Reactions were performed using serially diluted DNA template with input amounts of (1) 100 ng, (2) 10 ng, (3) 1 ng, (4) 100 pg, (5) 10 pg, (6) 1 pg, (7) 100 fg, and (8) 10 fg, and were separated on a 1% agarose gel using GeneRuler™ 100 bp Plus Ready-to-use Ladder as a standard. [Figure 16] Figure 1 shows the detection of DUID in yeast total DNA extracts as described in Example 2. Quantitative real-time PCR was performed on serial 10-fold dilutions of the YCp vector ranging from 50 ng to 500 ag and used to generate a standard curve (blue line) using MS Excel. Results from a similar qPCR experiment using DNA from BY4743 transformed with the YCp-DUID vector were plotted (orange bars) and compared to the standard curve values ​​to quantify the detection of DUID in yeast biomass. [Figure 17] FIG. 10 shows examples of homologies across identifier sequences that serve as a means to identify the version of the DUID, its origin, and subsequence protocols for interacting with the DUID, as further described in Example 2. DETAILED DESCRIPTION OF THE INVENTION

[0077] Methods and compositions for providing identification and / or traceability of biological materials are described herein. It is understood that the embodiments and examples are provided for illustrative purposes directed to those skilled in the art and are not intended to be limiting in any way.

[0078] Provided herein are methods and compositions for providing identification and / or traceability of biological materials. In certain embodiments, the methods described herein may utilize unique identifier sequences (also referred to herein as DNA unique identifier sequences) that can be exogenously introduced (i.e., inserted / integrated) into the genome of a biological entity to provide identification and / or traceability of the biological entity and / or material, including biological entities and / or materials produced from and containing genomic DNA derived therefrom. In certain embodiments, the strategies described herein may benefit from the durability and replicability of nucleic acids, e.g., DNA, that provide identification and / or traceability. In certain embodiments, the unique identifier sequences may be derived from a randomized pool of sequences. In certain embodiments, a database may maintain the association of unique identifier sequences with corresponding identification and / or traceability information.

[0079] Also provided herein are oligonucleotide constructs and cassettes comprising one or more unique identifier sequences for use in providing identification and / or traceability of biological materials. In certain embodiments, the oligonucleotide constructs and / or cassettes may comprise specific sequences of primer annealing sequence(s), which may be for amplification of the unique identifier sequence(s), sequencing of the unique identifier sequence(s), or both. In certain embodiments, the sequences of the primer annealing sequence(s) may be designed to reduce unintended and / or off-target amplification and / or sequencing events, which may result in, for example, increased fidelity and / or reduced errors in identification events, as described herein.

[0080] In certain embodiments, the methods and compositions described herein may be used to provide food traceability, e.g., allowing for rapid response and / or food recall in the event of contamination. Food contamination, e.g., E. coli and / or Salmonella contamination affecting the food supply, is a threat to public health, and rapid action to identify and contain the source(s) of contamination is highly desirable. There is a long-standing unmet need in the art for reliable, cost-effective, and / or rapid strategies to enhance traceability of products in the food supply. The strategies described herein may provide traceability in the food system from origin to digestion and beyond. Traceability of biological entities and / or biological materials is not only desirable in the agricultural and food industries, but is also required in a wide variety of industries and fields that handle biological entities and / or biological materials that contain or are derived from them. Thus, in addition to food safety, applications in food / seed security, IP tracking, certification (e.g., seed association, Kosher, Halal, etc.), GMO identification and / or characterization, and / or trade finance risk reduction are also contemplated herein.

[0081] In certain embodiments, food products or ingredients (e.g., fruits and vegetables, or other food products containing cells, etc.) may comprise the unique identifier sequence(s) described herein as part of the genome in at least some of their cells to provide identification and / or traceability. In other embodiments, the unique identifier sequence(s) described herein may be part of the genome of one or more biological entities or biological substances, including cells, that may be added to, mixed with, or otherwise associated with one or more products for which identification and / or tracking is desired. By way of example, in certain embodiments, food-safe yeast cells that comprise one or more unique identifier sequences described herein as part of one or more stably introduced artificial chromosome(s) may be added to or mixed with one or more food products or ingredients to provide identification and / or traceability thereof.

[0082] Methods for providing identification and / or traceability In certain embodiments, provided herein are methods for providing identification and / or traceability of biological materials or entities. Such methods may utilize unique identifier sequences to achieve such identification and / or traceability. Typically, a biological entity of interest, such as an agricultural crop (e.g., spinach), may be genetically modified to incorporate a unique identifier sequence into its genome. As a non-limiting and illustrative example, cells of a spinach plant may be genetically modified to incorporate, into the genome of a spinach cell, a cassette comprising the unique identifier sequence flanked by one or more primer annealing sequences for subsequent amplification and / or sequencing of the unique identifier sequence, at an intergenic site within the genome or at another innocuous site in the genome. The sequence of the unique identifier sequence may be known or may be derived from a randomized pool and subsequently determined after integration, and may be entered and recorded in a database or registry. The cells may then be used to grow / propagate one or more spinach crops, and identification and / or tracking information appropriate to the spinach crops (e.g., origin, batch / lot information, grower / producer, region, date, supplier, and / or any other supply chain information of interest) may be recorded in a database or registry in association with the corresponding unique identifier sequence. The database entries may be updated as supply chain events progress (i.e., harvest, shipment to supplier, sale, etc.). The spinach crops may be used to produce biological materials, such as bagged spinach or salads for sale in grocery stores. In the event of a contamination or foodborne illness event or suspicion, a sample of the suspect spinach or salad may be obtained, from which genomic DNA may be obtained and analyzed to determine whether the unique identifier sequence is present (i.e., whether the spinach is the spinach to be tracked by the system). If so, the unique identifier sequence may be sequenced to determine the nucleotide sequence, which may be used to query a database or registry to retrieve relevant database entries to provide identification and / or tracking information, thereby facilitating the recall of the contaminated spinach or salad.It will be appreciated that the above spinach example is provided for illustrative purposes, and that the methods described herein may be used to provide a wide variety of identification and / or traceability options for a wide variety of biological entities and / or biological materials in a wide variety of applications.

[0083] In embodiments, a method for identifying a biological substance includes: receiving or providing a sample containing genomic DNA from a biological material; amplifying at least one DNA-specific identifier sequence in genomic DNA from the biological material and sequencing the DNA-specific identifier sequence; searching a database for the DNA unique identifier sequence and obtaining a database entry corresponding to the DNA unique identifier sequence, the database entry providing identification and / or tracking information for the biological material; A method is provided herein, comprising:

[0084] A flow chart illustrating an embodiment of such a method is shown in FIG.

[0085] Of course, biological material can generally include any suitable biological material of interest. Biological material can include or consist of material that includes or consists of a biological entity, or any other suitable material of interest that is produced or derived from a biological entity or that contains genomic nucleic acid (i.e., genomic DNA) from a biological entity. In certain embodiments, biological material can include or consist of plant-derived material, fungal-derived material, animal-derived material, viral-derived material, or bacterial-derived material. By way of example, in certain embodiments, biological material can include or consist of a food or beverage that includes, consists of, or is produced from a plant or other biological entity, in which case the food or beverage contains genomic DNA from the biological entity. In certain embodiments, biological material can include or consist of, for example, lettuce, spinach, or other leafy green vegetables, or food products that include, consist of, or are produced from this biological material.

[0086] In certain embodiments of the methods described herein, one may receive or provide a sample containing genomic nucleic acid (i.e., genomic DNA, where the biological entity has a DNA-based genome) from a biological material of interest (e.g., biological material whose identification is desired). The sample may be received or provided in a purified or partially purified form so that the genomic DNA can be readily used, or may be provided as is (i.e., as a food product sample) or in another crude or precursor form, which may be subjected to one or more further processing or purification steps so that the genomic DNA contained therein can be readily used in a subsequent step. In certain embodiments, it is contemplated that any suitable standard technique for the purification and / or isolation of genomic nucleic acid may be used in sample preparation.

[0087] In certain embodiments, for example, one or more steps of nucleic acid (e.g., genomic) isolation, purification, and / or extraction may be performed as part of sample preparation for a subsequent step. DNA isolation or extraction may include, for example, one or more steps to obtain DNA from a sample. In certain embodiments, DNA isolation or extraction may include cell disruption (e.g., lysis) (e.g., by physical step(s), sonication, or chemical treatment), membrane removal using detergents, optional protein removal with proteases, and DNA precipitation using alcohol (e.g., ethanol (cold) or isopropanol). Thus, DNA precipitates may be obtained by centrifugation. In certain embodiments, DNAse enzymes may be inhibited by using chelating agents, as will be recognized by those of skill in the art. In certain embodiments, cellular and histone proteins may be removed using proteases, precipitated with sodium acetate or ammonium acetate, or by phenol-chloroform extraction prior to DNA precipitation. Those of skill in the art having regard to the teachings herein will appreciate that a wide variety of techniques are available for the desired sample preparation and / or isolation, purification and / or extraction of genomic nucleic acid.

[0088] In certain embodiments of the methods described herein, a unique identifier sequence (for convenience referred to herein as a DNA unique identifier sequence, DUID, although it will be understood that in certain embodiments where the biological entity has an RNA-based genome the unique identifier sequence may be RNA rather than DNA) inserted or incorporated within the genome of a biological entity / material may be amplified.

[0089] In certain embodiments, the integration into genome can comprise integration into natural chromosome.In certain embodiments, the integration into genome can comprise the stable introduction of artificial chromosome into genome, wherein the artificial chromosome has a centromere sequence and is heritable with natural genome material.The following example 2 describes the use of artificial chromosome in yeast, for example.

[0090] Such amplification may generally be performed using any suitable amplification technique known to those of skill in the art having regard to the teachings herein, such as by polymerase chain reaction (PCR). In certain embodiments, as described in further detail herein, the unique identifier sequence to be amplified may be accompanied by primer annealing sequences for amplification and / or sequencing in the genome. In certain embodiments, as described in further detail herein, the primer annealing sequences are selected and positioned to allow amplification by nested PCR, which may reduce the likelihood of unintended or off-target amplification.

[0091] In certain embodiments, PCR-based techniques may be used for amplification. PCR amplification may include forward and reverse primers, where the primers may be complementary (or substantially complementary) to the 5' and 3' regions toward the ends of the nucleic acid sequence to be amplified. Forward and reverse primers for specific primer annealing sequences may be generated by any suitable technique known to those of skill in the art. Examples of such techniques may be found, for example, in Dieffenbach CW, Dveksler GS. 1995. PCR primer: a laboratory manual, New York, NY: Cold Spring Harbor Laboratory Press; New England Biolabs Inc., 2007-08 Catalog & Technical Reference, incorporated herein by reference. In certain embodiments, in reading biological information, PCR primers may include multiple sets of forward and reverse primers that can operate independently of each other. In certain embodiments, the identities of some primers may be provided or distributed so that various parties can easily access various regions and / or nucleic acid sequence information as desired, while access to other primers may be controlled.

[0092] In certain embodiments of the methods described herein, the unique identifier sequence, e.g., a DNA unique identifier sequence (DUID), can comprise any suitable nucleic acid sequence exogenously introduced into the genome of a biological entity for identification purposes. Generally, the unique identifier sequence can be either DNA or RNA to match the genome type (DNA or RNA) of the biological entity. Of course, because the genomes of many biological entities, such as plants, are double-stranded, the unique identifier sequence is typically found in the double-stranded form of the genome. Thus, in certain embodiments, reference to a unique identifier sequence herein (e.g., when describing the sequencing of an identifier sequence, etc.) can be understood to be a reference to either or both strands of a double-stranded construct, as desired or appropriate.

[0093] In certain embodiments, the unique identifier sequence may be incorporated into a cassette or other such construct that includes one or more functional elements in addition to the unique identifier sequence. In certain embodiments, the cassette may include the unique identifier sequence adjacent to one or more primer annealing sequences for PCR amplification of the DNA unique identifier sequence, sequencing of the DNA unique identifier sequence, or both. It should be understood that a primer annealing sequence may refer to a predetermined sequence or region of a nucleic acid having a known nucleotide sequence, such that one or more primers may be designed or selected to anneal to such primer annealing sequence and initiate polymerization by a polymerase. Typically, the primer annealing sequence is selected to be unique in the genome of the biological entity of interest, thereby reducing or eliminating unintended or off-target amplification. In certain embodiments, the unique identifier sequence may be, for example, a known, predetermined sequence selected for specific amplification, or may be a random sequence derived from a randomized pool of nucleic acid sequences, which may be subsequently determined and recorded in a database, as described in detail herein. In certain embodiments, the unique identifier sequence, or a cassette comprising the unique identifier sequence, can have a length of up to about 1500 nt, a length of up to about 1000 nt, a length of about 200 nt to about 600 nt, a length of about 200 nt to about 400 nt, or a length of about 400 nt to about 600 nt, or any size, or a subrange between any two such sizes. Naturally, a longer unique identifier sequence may enable the generation of more unique sequences in a pool and reduce the risk of duplication. Furthermore, in embodiments where encoding or encryption of identifying information in the unique identifier sequence is desired, a longer length may allow, for example, the storage of relatively more information and / or the use of more complex encryption or coding schemes. Thus, maintaining a reasonable length as referred to herein may enable more reliable and / or rapid amplification and / or sequencing and / or relatively reduced costs.

[0094] In certain embodiments, the unique identifier sequence may comprise a sequence up to about 1500 nt in length, up to about 1000 nt in length, about 200 nt to about 600 nt in length, about 200 nt to about 400 nt in length, or about 400 nt to about 600 nt in length. In certain embodiments, the unique identifier sequence may be relatively short, for example, about 20 bp in length. Of course, in certain embodiments, the size of the unique identifier sequence may be selected to suit a particular implementation and its desired parameters. In certain embodiments, the unique identifier sequence may have a size of about 20 nt to about 1500 nt, or any size therebetween, or any subrange therein.

[0095] In certain embodiments, unique identifier sequences can be obtained, for example, from a random pool and may be screened for validity (e.g., screened for uniqueness and to avoid undesirable sequence motifs), or rationally designed (e.g., designed for uniqueness and to avoid undesirable sequence motifs).

[0096] In certain embodiments, the DNA-specific identifier sequence may be flanked by one or more primer annealing sequences for PCR amplification of the DNA-specific identifier sequence, sequencing of the DNA-specific identifier sequence, or both.

[0097] In certain embodiments, it is contemplated that the unique identifier sequence may be formed into a cassette or otherwise introduced or inserted into genomic nucleic acid such that it is adjacent to one or more primer annealing sequences for PCR amplification of the DNA unique identifier sequence, sequencing of the DNA unique identifier sequence, or both. Examples of suitable cassettes and configurations are described in further detail herein. In certain embodiments, the cassette may be incorporated into a plasmid, vector, or other such carrier suitable for use in inserting / incorporating / integrating the cassette into the genome of a biological entity.

[0098] It will be appreciated that any suitable genetic modification technique known to those of skill in the art armed with the teachings herein may be used for the introduction / insertion / incorporation / integration of a unique identifier sequence or a cassette / vector comprising a unique identifier sequence into the genome of a biological entity. It will also be appreciated that the genetic modification technique may be selected based on the unique identifier sequence or cassette / vector used and the particular biological entity to be modified. Techniques for genome modification of a wide variety of biological entities, including plants, animals, fungi, bacteria, and viruses, are well known and may be readily adapted to exogenously introduce the unique identifier sequences described herein.

[0099] By way of example, those skilled in the art, having regard to the teachings herein, will recognize vectors for incorporating DNA into organisms, which may be designed according to known principles of molecular biology. Such vectors may be designed, for example, to stably introduce a DNA sequence of interest into the genome of an organism. In certain embodiments, the vector may be, for example, of viral origin or derived therefrom. It is contemplated that when the organism is a plant, Agrobacterium tumefaciens-mediated integration of the DNA of interest may be used for introduction into the plant. Those skilled in the art, having regard to the teachings herein, will recognize several other transformation methods, such as ballistic or particle gun methods, among others, which may be adapted as desired or appropriate based on the particular application of interest. In certain embodiments, gene delivery systems may be used based on genetic engineering principles to introduce or insert a sequence of interest into the genome of a host organism. For example, in embodiments, a transposon system may be used to insert into the genome of a host, which may be, for example, a microorganism, an animal cell, or a plant cell (Insect Molecular Biology (2007), 16(1), 37-47, Plant Physiology Preview. 2007, DOI: 10.1104 / pp.107.111427, the American Society of Plant Biologists; research on production of lactoferrin from transformed silkworms and functionality thereof, the Ministry of Agriculture and Forestry, 2005).In certain embodiments, any suitable method in the fields of molecular biology and / or genetic engineering that allows for the insertion of one or more DNA fragments or components of interest into the genome of a host may be used (see, e.g., Transgenic Plants Methods and Protocols., Methods in Molecular Biology 2019, Editors: Kumar, Sandeep, Barone, Pierluigi, Smith, Michelle, ISBN 978-1-4939-8778-8, which is incorporated herein by reference in its entirety).

[0100] In certain embodiments, where it is desired to identify a biological substance or biological entity that contains a unique identifier sequence, the sequence of the unique identifier sequence can be determined by sequencing. Of course, the unique identifier sequence can generally be sequenced by any suitable sequencing technique known to those skilled in the art after considering the teachings herein. In certain embodiments, the sequencing can be aided by the inclusion or use of a sequencing primer annealing sequence that binds to the unique identifier sequence in the genomic nucleic acid. For example, examples of such sequencing primer annealing sequences that can be incorporated into a cassette containing a unique identifier sequence are described in detail herein.

[0101] In certain embodiments, sequencing may be performed using any suitable sequencing technology known to those of skill in the art having regard to the teachings herein, and the technology may be selected based on the particular application and / or configuration used. In certain embodiments, sequencing may be performed by any suitable sequencing method for determining the order of nucleotide bases in a DNA (or RNA) molecule. Examples of sequencing methods may include, for example, Maxam-Gilbert sequencing, chain termination, dye terminator sequencing, automated DNA sequencing, in vitro cloning and amplification, parallel sequencing by synthesis, sequencing by ligation, Sanger sequencing, e.g., microfluidic Sanger sequencing, and sequencing by hybridization.

[0102] In certain embodiments, once the sequence of a unique identifier sequence of a biological material is determined, this sequence may be used to generate a query to search a database (also referred to herein as a registry) that contains a collection of unique identifier sequences that are paired or otherwise associated with appropriate identifying and / or tracking information. If a matching database entry is found, retrieving the database entry may provide the identifying and / or tracking information for the biological material of interest. In such a manner, appropriate identifying and / or tracking information for the biological material may be identified, which may be used, for example, to signal an event such as a food recall or other action.

[0103] In another embodiment, there is provided a method for providing traceability of a biological material, comprising: determining the sequence of at least one DNA-specific identifier sequence within the genomic DNA of the biological entity; verifying the identity of the biological entity by establishing the presence of the DNA-specific identifier sequence in the genomic DNA and comparing the sequence of the DNA-specific identifier sequence with a database to confirm that the DNA-specific identifier sequence is not already used in the database; providing an indication of the advisability of producing a biological material from the biological entity, the biological material comprising genomic DNA from the biological entity; entering the sequence of the at least one DNA unique identifier sequence into a database entry in a database, and associating the DNA unique identifier sequence with identification and / or tracking information for the biological material; whereby providing traceability of a biological material by reading a DNA unique identifier sequence in the biological material and retrieving a corresponding database entry to provide identification and / or tracking information for the biological material.

[0104] A flow chart illustrating an embodiment of such a method is shown in FIG.

[0105] Of course, the biological entity may generally include any suitable biological entity of interest. The biological entity may comprise or consist of a cell (i.e., a plant cell, a fungal cell, an animal cell, or a bacterial cell), or a seed or tissue comprising one or more cells, or a virus, or an organism, such as a plant, an animal, or a fungus, or any part thereof. In certain embodiments, the biological entity may comprise a plant cell, a fungal cell, an animal cell, a virus, or a bacterial cell. When the biological entity is genetically modified to incorporate a unique identifier sequence, the biological entity may typically comprise a cell or a virus, which, after genetic modification, may be propagated to produce further biological entities, each comprising the inserted unique identifier sequence.

[0106] In certain embodiments, the verifying step may be carried out to verify the presence of the unique identifier sequence within the genomic DNA of the biological entity and / or to determine this sequence and / or to determine whether the unique identifier sequence is not already used in the database (i.e., it is a new sequence not already previously associated with a database entry). If the verification is successful (i.e., the unique identifier sequence is properly inserted and unique to the database), in certain embodiments, a database entry for the unique identifier sequence may be created in the database (which may be associated with appropriate identifying and / or tracking information and may be continually updated), and may provide an indication of the appropriateness of producing biological material from the biological entity to interested parties, such as growers, farmers, or other agricultural stakeholders who may then produce or cultivate the biological material.

[0107] In such methods, traceability of a biological material may be provided by reading (i.e., sequencing) the unique identifier sequence of the biological material, which may be used to obtain a corresponding database entry and obtain identification and / or tracking information.

[0108] In certain embodiments, the methods described herein may further comprise inserting at least one DNA-specific identifier sequence into the genomic DNA of the biological entity, or modifying by gene editing an identifier sequence already present in the genomic DNA of the biological entity to generate a DNA-specific identifier sequence in the genomic DNA of the biological entity, thereby achieving this identification.

[0109] In yet another embodiment, the methods described herein may further comprise providing at least one DNA-specific identifier sequence for insertion into the genomic DNA of the biological entity. In certain embodiments, the DNA-specific identifier sequence may be provided as a randomized pool of sequences as further described herein.

[0110] It is understood that in certain embodiments, the methods described herein may utilize a single unique identifier sequence or may use two or more identifier sequences incorporated into the genome to provide identification and / or traceability.

[0111] In certain embodiments, the unique identifier sequence may be derived from a randomized pool of unique identifier sequences. The identity of the inserted unique identifier sequence may not be determined until insertion (i.e., transformation or genetic modification) is achieved. In such methods, it is contemplated that an interested party may be provided with a randomized pool of unique identifier sequences and may perform genetic modification of the biological entity of interest such that one, two, or more unique identifier sequence(s) are inserted into the genome. After the genetic modification process, the inserted unique identifier sequence(s) may be sequenced to determine the nucleotide sequence of the inserted unique identifier sequence(s). Given that the typical length of a unique identifier sequence may typically be selected to be sufficiently long to generate a large number of different sequences within the randomized pool, the statistical likelihood of two different parties inserting the same unique identifier sequence is extremely low. Thus, such methods, in certain embodiments, contemplate providing samples from the same or similar pools of randomized sequences for insertion into biological entities of interest to many different parties seeking to benefit from the identification and / or traceability of the methods described herein. It is contemplated that such methods, in certain embodiments, may streamline processes and / or reduce costs.

[0112] In yet another embodiment of the method described herein, the step of reading the DNA unique identifier sequence in the biological material and obtaining the corresponding database entry comprises: receiving or providing a sample containing genomic DNA from a biological material; amplifying at least one DNA-specific identifier sequence in genomic DNA from the biological material and sequencing the DNA-specific identifier sequence; comparing the sequence of the DNA unique identifier sequence with a database to obtain a database entry corresponding to the DNA unique identifier sequence, the database entry providing identification and / or tracking information for the biological material; may include:

[0113] In certain embodiments, it is contemplated that the unique identifier sequence(s) may be inserted into a substantially innocuous site within the genome of a biological entity (i.e., have the potential to have substantially no effect on gene expression or phenotype). For example, in certain embodiments, it is contemplated that the unique identifier sequence(s) may be inserted into one or more intergenic region(s) of genomic DNA.

[0114] In certain embodiments, the identification and / or tracking information recorded in the database or registry may include supply chain information for the biological material. In certain embodiments, the identification and / or tracking information in the database may include source information for the biological material. In certain embodiments, the identification and / or tracking information in the database may include grower, region, batch, lot, date, or other relevant supply chain information, or any combination thereof. Those of skill in the art, armed with the teachings herein, will recognize and may select a variety of identification and / or tracking information to include in the database as desired or suitable for a particular application. In certain embodiments, existing supply chain tracking features, such as barcodes or lot or batch numbers, may be included in the database, for example. In certain embodiments, the database may contain / store information such as geographic region, date, purchaser, farmer, lot, sublot, harvest, batch, other DUID-enabled products, organisms, contractual obligations, certifications, nearby industry and commerce, sensor data, weather data, or any combination thereof.

[0115] In another embodiment, there is provided a method for identifying a biological substance, comprising: receiving, at a computing device, a DNA unique identifier sequence (DUID) extracted from a known biological material; searching a DUID database, which stores a plurality of DUIDs associated with respective biological material information, in the computing device to match the received DUID; if the search of the DUID database fails to match the received DUID, storing the received DUID in the DUID database in association with biological material information associated with known biological materials; receiving, at the computing device after storing the received DUID and information associated with the known biological material in a DUID database, a query DUID extracted from the unknown biological material; searching a DUID database at the computing device to match the received query DUID; if the search for the DUID results in a match with the received query DUID, returning the stored biological information associated with the DUID that matches the query DUID in response to the received query DUID; A method is provided herein, comprising:

[0116] A flowchart illustrating an embodiment of such a method is shown in Figure 7. In this figure, a DNA unique identifier sequence (DUID through DuID4 in the illustrated example) is extracted (i.e., read, determined, or sequenced) from a known biological material and provided to a computing device. The computing device then searches a DUID database (i.e., DuID data store) that stores multiple DUIDs in association with respective biological material information to match the received DUID4. If the DUID database search fails to match the received DUID, the DUID database stores the received DUID (DuID4) in association with biological material information (i.e., producer4 information) associated with the known biological material, thereby enabling registration of the DUID and biological material in the database. Notification of successful registration may then be provided to interested parties, who may then authorize the biological entity / material to proceed with propagation and produce the biological material, e.g., food product. After storing the received DUID and information associated with known biological material in the DUID database, a query DUID extracted (i.e., read, for example, by sequencing) from an unknown biological material (i.e., a biological material of interest, e.g., a food product suspected of being contaminated) may be received at the computing device, and a search of the DUID database may be performed to match the received query DUID. If the search of the DUID database results in a match with the received query DUID, the stored biological information associated with the matching DUID may be returned in response to the received query DUID, thereby providing tracking and / or identification information for the biological material, which may be used, for example, to respond to a food recall, etc.

[0117] In another embodiment, the step of searching the DUID database to match the received DUID comprises: searching a DUID database for an exact match with the received DUID; if no exact match is found, performing an alignment / identity search against DUIDs stored in the DUID database that closely match the received DUID; may include:

[0118] In yet another embodiment, the step of searching the DUID database to match the query DUID comprises: searching the DUID database for an exact match with the query DUID; if no exact match is found, performing an alignment / identity search against DUIDs stored in the DUID database that closely match the query DUID; may include:

[0119] Of course, because nucleic acid sequences are used, there may be the possibility of sequence mutations of the unique identifier sequence during propagation, and / or amplification and / or sequencing errors may occur. Thus, in certain embodiments, such alignment / identity searches may be performed to identify whether there may be nearly or very similarly matching entries. A variety of sequence comparison algorithms exist for performing such alignment / identity / similarity assessments (see, for example, the BLAST tool available from NCBI), and those skilled in the art, having regard to the teachings herein, may select or adapt an appropriate algorithm as desired to suit a particular application.

[0120] In yet another embodiment, the methods described herein comprise: If the search results in a near match with the query DUID, storing the query DUID in association with the DUID that nearly matches the query DUID. It may further include:

[0121] In this way, the database can be updated, for example, as sequence variations are identified.

[0122] In another embodiment, there is provided a computer system for identifying a biological substance, comprising: a processing unit capable of executing instructions; and a memory unit storing instructions that, when executed on a processing unit, configure a computer system to perform any one or more of the methods described herein; Provided herein is a computer system comprising:

[0123] In another embodiment, provided herein is a computer readable memory having stored thereon instructions that, when executed by a processing unit of a computer system, configure the system to perform any one or more of the methods described herein.

[0124] In another embodiment, there is provided a method for identifying a biological substance, comprising: receiving or providing a sample containing genomic DNA from a biological material; amplifying at least one DNA-specific identifier sequence in genomic DNA from the biological material and sequencing the DNA-specific identifier sequence; decoding or decoding the identification and / or tracking information of the biological material stored in the DNA unique identifier sequence; A method is provided herein, comprising:

[0125] Such method embodiments may be similar to those described herein that utilize databases or registries, except that the identification and / or tracking information may not be stored in a database, but instead may be encoded (encrypted or unencrypted) in the unique identifier sequence itself. Techniques for storing information in nucleic acid sequences are known in the art and may typically include the use of A, T, G, C nucleotides, as well as 0 and 1 bits in digital data storage. Examples of techniques for storing / encoding / encrypting information may be found, for example, in Clelland, C., Risca, V. & Bancroft, C. Hiding messages in DNA microdots. Nature 399, 533-534 (1999) doi:10.1038 / 21092 (incorporated herein by reference).

[0126] A flow chart illustrating an embodiment of such a method is shown in FIG.

[0127] It is contemplated that in certain embodiments, the unique identifier sequence may be used to encode a key, which is to be stored in a database in association with tracking and / or identification information. Thus, it is understood that references herein to storing a DUID in a database and searching a DUID in a database may be considered to encompass both direct (i.e., storing and searching the primary nucleic acid sequence of the unique identifier sequence itself) and indirect (i.e., deriving a key from the primary nucleic acid sequence of the unique identifier sequence and storing it in a database and using the key to search the database) options. Those skilled in the art, armed with the teachings herein, will recognize a variety of combinations that may be used, all of which are intended to be encompassed herein.

[0128] In another embodiment, there is provided a method for providing traceability of a biological material, comprising: determining the sequence of at least one DNA-specific identifier sequence within the genomic DNA of the biological entity; verifying the identity of the biological entity by verifying the presence of the DNA-specific identifier sequence in the genomic DNA and decoding or decoding the identifying and / or tracking information stored in the DNA-specific identifier sequence to verify the DNA-specific identifier sequence; providing an indication of the adequacy of producing a biological material from the biological entity, the biological material comprising genomic DNA from the biological entity; whereby providing traceability of a biological material by reading a DNA unique identifier sequence in the biological material and decoding or deciphering the information stored in the DNA unique identifier sequence to provide identification and / or tracking information for the biological material.

[0129] A flow chart illustrating an embodiment of such a method is shown in FIG.

[0130] Such method embodiments may be similar to those described herein that utilize databases or registries, except that the identification and / or tracking information may not be stored in a database, but instead may be encoded (encrypted or unencrypted) in the unique identifier sequence itself. Techniques for storing information in nucleic acid sequences are known in the art and may typically include the use of A, T, G, C nucleotides, as well as 0 and 1 bits in digital data storage. Examples of techniques for storing / encoding / encrypting information may be found, for example, in Clelland, C., Risca, V. & Bancroft, C. Hiding messages in DNA microdots. Nature 399, 533-534 (1999) doi:10.1038 / 21092 (incorporated herein by reference).

[0131] In yet another embodiment, there is provided a method for identifying a biological substance, comprising: receiving, at a computing device, a DNA unique identifier sequence (DUID) extracted from an unknown biological material; decoding or decoding the identification and / or tracking information of the unknown biological material stored in the DNA unique identifier sequence; A method is provided herein, comprising:

[0132] A flow chart illustrating an embodiment of such a method is shown in FIG.

[0133] In yet another embodiment, there is provided a method for providing traceability of biological material, comprising: Inserting at least one DNA-specific identifier sequence into the genomic DNA of a biological entity for use in the preparation of a biological material. A method is provided herein, comprising:

[0134] In another embodiment of the above method, the DNA unique identifier sequence may be inserted as any one or more cassettes described herein.

[0135] In another embodiment of any of one or more of the above methods, the method may further comprise determining the sequence of at least one DNA-specific identifier sequence within the genomic DNA of the biological entity.

[0136] In another embodiment of any of one or more of the above methods, the method may further comprise verifying the identity of the biological entity by establishing the presence of the DNA-specific identifier sequence in the genomic DNA and comparing the sequence of the DNA-specific identifier sequence with a database to confirm that the DNA-specific identifier sequence is not already used in the database.

[0137] In yet another embodiment of any one or more of the above methods, the method further comprises: Producing a biological material from a biological entity, the biological material comprising genomic DNA from the biological entity; and / or providing an indication of the advisability of producing a biological material from the biological entity, the biological material comprising genomic DNA from the biological entity; It may further include:

[0138] In yet another embodiment of any of one or more of the above methods, the method may further comprise entering the sequence of the at least one DNA-specific identifier sequence into a database entry and associating the DNA-specific identifier sequence with identification and / or tracking information for the biological entity and / or biological material.

[0139] In yet another embodiment of any one or more of the above methods, the method further comprises: providing traceability of the biological entity and / or biological substance by reading the DNA unique identifier sequence in the biological entity and / or biological substance and obtaining a corresponding database entry to provide identification and / or tracking information for the biological entity and / or biological substance. It may further include:

[0140] Oligonucleotide constructs, cassettes, plasmids, vectors, cells, and kits In another embodiment, provided herein is a cassette comprising a unique identifier sequence, wherein the unique identifier sequence is flanked by at least one 5' primer annealing sequence and at least one 3' primer annealing sequence for amplification of the DNA unique identifier sequence, sequencing of the DNA unique identifier sequence, or both.

[0141] Of course, in certain embodiments, such cassettes may be for use in any one or more of the methods described herein.

[0142] In certain embodiments of the cassette, the DNA-specific identifier sequence may be flanked by two 5' primer annealing sequences and two 3' primer annealing sequences, allowing amplification of the DNA-specific identifier sequence by nested PCR. In certain embodiments, a nested design may be used, for example, to improve recall fidelity. In still further embodiments of the cassette, the two 5' primer annealing sequences may partially overlap, or the two 3' primer annealing sequences may partially overlap, or both. In still further embodiments of the cassette, the cassette may further comprise a sequencing primer annealing sequence located 5' to the DNA-specific identifier sequence for sequencing the DNA-specific identifier sequence. In yet further embodiments of the cassette, the sequencing primer annealing sequence may be located between the two 5' primer annealing sequences. In further embodiments of the cassette, the sequencing primer annealing sequence may at least partially overlap one or both of the two 5' primer annealing sequences. In still further embodiments of the cassette, the two 5' primer annealing sequences may partially overlap, and at least a portion of the sequencing primer annealing sequence may be located in the overlap. In further embodiments of the cassette, the cassette sequence may be up to about 1500 nt in length, up to about 1000 nt in length, from about 200 nt to about 600 nt in length, from about 200 nt to about 400 nt in length, or from about 400 nt to about 600 nt in length.

[0143] An example of an embodiment of a cassette described herein and a process for its generation is shown in FIG. 2, where the cassette may be generated using a pool of oligonucleotides of randomized sequences. Randomized pools of oligonucleotides may be obtained commercially or synthesized on demand. They may be constructed, for example, by enzymatic polymerization or ligation, or may be chemically synthesized. The randomized oligonucleotide fragments may be purified, for example, by column separation to isolate fragments of approximately the same or similar size (e.g., approximately 300 nt to 400 nt in size in the example shown) and inserted into the cassette. A pool of cassettes containing a variety of different unique identifier sequences (i.e., approximately 10 in some examples) may be generated using a pool of oligonucleotides of randomized sequences. 7 ) may be generated. The cassette may include a primer annealing sequence (i.e., a primer binding site) and at least one sequencing primer annealing sequence (i.e., a sequencing primer binding site) in an appropriate arrangement to allow amplification and / or sequencing of the DUID, such as the configuration shown in FIG. 2. The primers and sequencing sites may be verified against the host genome to verify the absence of natural amplification. Cassettes with different primers for different organisms or different genomes may be utilized as desired. The cassette may include a restriction enzyme array site and may be formed, for example, in the form of an insertion cassette carrier plasmid or vector. In certain embodiments, the cassette may be about 500 bp in length, for example, formed within a plasmid or carrier vector about 1200 bp in size.

[0144] It will be appreciated that a primer annealing sequence of a cassette may refer to a predetermined sequence or region of a nucleic acid having a known nucleotide sequence, such that, for example, one or more primers may be designed or selected to anneal to such primer annealing sequence and initiate polymerization by a polymerase. The primer annealing sequence may be used for amplifying the unique identifier sequence, sequencing the unique identifier sequence, or both.

[0145] Figure 13 shows further examples of cassette designs described herein that include a UID (unique identifier) ​​sequence. Figure 13(a) shows a dual primer design, 13(b) shows a single primer design, and 13(c) shows a standalone design. The dual primer insertion cassette design of Figure 13(a) includes, in the illustrated embodiment, a restriction enzyme array, a 5' "primer A" region, and a 5' "primer B" region (where a 5' sequencing primer may anneal in the region spanning between the "primer A" and "primer B" regions), followed by a blunt-end ligation site. The UID region (e.g., variable bp random DNA or another identifier sequence) may then be formed, followed by a CAS9 PAM site, as shown. The blunt-end ligation site follows, followed by a 3' "primer B" region and a 3' "primer A" region, followed by a restriction enzyme array. The single primer insertion cassette design of Figure 13(b) includes, in the embodiment shown, a restriction enzyme array, a 5' "primer A" region (which may anneal with a 5' sequencing primer in this case), followed by a blunt-end ligation site. A UID region (e.g., variable bp random DNA or another identifier sequence) may then be formed, and a CAS9 PAM site may then be formed, as shown. A blunt-end ligation site follows, followed by a 3' "primer B" region, which then forms the restriction enzyme array. Figure 13(c) shows an embodiment of a stand-alone insertion cassette design, which includes a restriction enzyme array, a UID region (e.g., variable bp random DNA or another identifier sequence), an optional forming CAS9 PAM site, and a restriction enzyme array, as shown.

[0146] A wide variety of cassette designs are contemplated, as shown in Figure 13. Cassettes can differ, for example, with respect to existing elements, size, and amplification efficiency. Depending on whether primer pairs are present (see Figures 13(A)-(C)), the total cassette size can vary. For example, if individual primer pairs are removed, the total cassette size can be reduced (e.g., to approximately 40 bp in certain embodiments). Of course, in certain embodiments, the amplification efficiency of the UID may be reduced as a result of removing primer pairs. For example, in a dual primer design, any permutation of primers may be used for amplification, thereby providing four possible variations rather than those found in a single primer pair design. It should also be understood that in certain embodiments, a reduced cassette size may, for example, reduce the likelihood of unintended effects. In certain embodiments, any CAS9 PAM site may be used to enable efficient CRISPR-based editing of the UID sequence in progeny of the transformed organism. In certain embodiments where all primers are removed from the cassette design, it is contemplated that CAS9 PAM sites may be provided, in which case, in certain embodiments, the CAS9 PAM sites may allow, for example, a stand-alone cassette to be composed entirely of host genomic DNA, such as when using DNA degradation / ligation techniques. In certain embodiments, the UID sequence may be variable in length. In certain embodiments, it is contemplated that, for example, even short UID sequences may be safely used, especially when a validation step is performed that includes checking for any collisions in UIDs present in the registry and the newly inserted UID.

[0147] In yet another embodiment of the cassettes described herein, the primer annealing sequences may not naturally occur in the genome of the target biological entity, in such a way that unintended and / or off-target amplification and / or sequencing may be reduced or avoided.

[0148] In another embodiment, provided herein is a composition comprising a plurality of any one or more cassettes described herein, wherein each cassette comprises an identical primer annealing sequence and each cassette comprises a randomized DNA-specific identifier sequence. Such a composition may represent an example of a randomized pool of sequences described herein.

[0149] In yet another embodiment, provided herein is a composition comprising a plurality of any one or more cassettes described herein, wherein each cassette comprises the same primer annealing sequence and the same sequencing primer annealing sequence, and each cassette comprises a randomized DNA-specific identifier sequence. Such a composition may represent an example of a randomized pool of sequences described herein.

[0150] In another embodiment, provided herein is a plasmid, expression vector, or other single- or double-stranded oligonucleotide construct comprising any one or more of the oligonucleotides described herein or any one or more of the cassettes described herein.

[0151] In another embodiment, provided herein is a cassette comprising any one or more of the oligonucleotides described herein.

[0152] In yet another embodiment, provided herein is a cell or virus that has incorporated into its genome one or more of any of the oligonucleotides described herein or any of the cassettes described herein. In another embodiment, provided herein is a cell or virus that has incorporated into its genome a unique identifier sequence. In another embodiment of any of the cells or viruses described herein, the unique identifier sequence may be incorporated into an intergenic region of the genomic nucleic acid of the cell or virus. In yet another embodiment of any of the cells or viruses, the cell may be a plant cell, a fungal cell, an animal cell, or a bacterial cell.

[0153] In another embodiment, DNA-specific identifier sequence, a randomized pool of DNA-specific identifier sequences; Any one or more of the oligonucleotides described herein; Any one or more of the cassettes described herein; one or more primer pairs for amplifying and / or sequencing the DNA unique identifier sequence; buffer solution, polymerase, or Instructions for carrying out any one or more of the methods described herein Provided herein are kits that include any one or more of the following, or any combination thereof:

[0154] In yet another embodiment, there is provided a method for providing traceability of a product of interest, comprising: Receiving or providing a sample from a product of interest, the sample comprising genomic DNA from biological material that is part of, mixed with, or otherwise associated with the product of interest; amplifying at least one DNA-specific identifier sequence in genomic DNA from the biological material and sequencing the DNA-specific identifier sequence; searching a database for the DNA unique identifier sequence and obtaining a database entry corresponding to the DNA unique identifier sequence, the database entry providing identification and / or tracking information for the product of interest; A method is provided herein, comprising:

[0155] In another embodiment of the above method, the method may include introducing or adding any one or more biological materials or one or more biological entities described herein to a product of interest, wherein the biological material or entity includes at least one DNA unique identifier sequence described herein as part of its genomic material.

[0156] In yet another embodiment of any of one or more of the above methods, the identification and / or tracking information of the database entry may include supply chain information for the product of interest.

[0157] In yet another embodiment of any of one or more of the above methods, the product of interest may include food, agricultural products, pharmaceuticals, retail products, textiles, goods, chemicals, or another supply chain item. [Example]

[0158] An exemplary DUID system for providing food traceability This example describes an exemplary food traceability system embodiment, referred to herein as a DNA Unique Identifier (DUID) system. This example takes advantage of the durability and replicability of DNA sequences to securely encode unique identifiers within the nuclear genome of an organism. By encoding identifying information within an organism's DNA using the methods described above, detailed information in traceability can be provided throughout the supply chain. In particular, the DUID system: 1. The ability to safely achieve DNA-level population identification without affecting the genetic traits of the target organism; 2. Ability to build logical relationships between DUIDs and reference information; 3. The ability to reduce the time it takes to trace produce back to its origin from several months to about one day; 4. The ability to rapidly identify both the origin of a product and its ultimate path through the supply chain; 5. The ability to provide useful information to medical professionals and industry watchdogs; 6. The ability to foster consumer and industry confidence in the stability, transparency, and efficiency of the food supply chain; and / or 7. Mechanisms for enforcing the obligations of members in the association and the ability to support intellectual property rights in food products. may have:

[0159] In certain embodiments, for example, we consider that a DUID system can be used to significantly enhance the oversight capabilities of food system stakeholders. In addition to providing traceability, the DUID system described herein can challenge traditional thinking about points of belonging, from the bottom up rather than the top down. Such an approach described herein may be particularly desirable given that improved supply chain enforcement is becoming commonplace. The DUID system described herein can provide substantial assurance of provenance traceability, generally from anywhere along the supply chain, within about one day if needed. The system can benefit from the replicable and stable cellular properties of organisms, resulting in marginal costs approaching zero when generating progeny. The economic costs and risk of tampering and / or fraud in traditional tracking systems are prohibitive, and the legal implications of malicious activity can be significant. The DUID system described herein can be edited in interesting ways, for example, so that progeny of a population retain part of their original identifier. DUIDs can also be utilized by medical professionals who may wish to test human waste, for example, to identify recently consumed foods.

[0160] Consider that the population-level identification described above may include further reference to legal arrangements. Consider, for example, that a product IP owner may intentionally associate a growing material with, for example, a specific grower and / or region. Population-level genetic identification combined with traditional whole-chain traceability techniques may enable a remarkable level of control over product movement. Consider, for example, a spinach plant variety that has been genetically modified to be resistant to various pests. Using a DUID system may significantly reduce costs in detection. In addition to cost reduction, a DUID system may, for example, serve as a registry, providing a centralized point of contact for IP tracking.

[0161] Many organisms are regulated. For example, plant species may be precursors to narcotics. Such organisms may, in certain embodiments, benefit from close association with, for example, authorized legal entities. Therefore, it is contemplated that such cases may benefit from the strategies described herein.

[0162] Consider, as an example, the regulation of cannabis in Canada. The production and distribution of cannabis plants, as well as their propagation materials, are regulated. In certain embodiments, it is contemplated that, for example, licensed cannabis producers may include DUIDs within their products, which may be used to aid in regulation. In certain embodiments, such DUIDs may be useful for regulation by identifying and / or tracking cannabis, even in complex cases where cannabis is mixed with other things (i.e., for example, in edible products).

[0163] As another example, consider a spinach growers' association. In certain embodiments, membership in the association may be required to grow and sell spinach. In certain embodiments, consider that such propagation material may be derived from DUID-prepared plants. Random audits, for example, may then be conducted at the retail level to ensure that all spinach sold is certified.

[0164] DUID System: The DUID system may include, for example, product identification, DUID verification, DUID reading, and subsequent product population tracking. It may also serve as a central registry of all DUID data.

[0165] In the following example, the DUID platform may include a collection of actors, business services, tasks, events, and systems. Actors may perform or trigger business services and tasks. Systems and business services may be understood in terms of the events they generate. Events may be directly related to the tracking status of food products.

[0166] Actor: As an example, a Consumer Safety Officer (actor) at the FDA (actor) may request that the DUID Platform (actor) attempt to read the DUID (business service) from a supplied organic substance of interest. Actors are the engine of the DUID Platform. Actors can be systems, institutions, and / or individuals. They can trigger events and make requests to business services. Actors can also perform tasks. The following list provides some examples of actors; however, this is a non-exhaustive list intended for illustrative purposes. - DUID Platform DUID Registry DUID API analytical chemist microbiologist DNA sequencer - Producer botanist CFO Traceability Software - grower food safety officer Enterprise Resource Planning System - Packer / Shipper Truck driver executive CEO - Retailers food safety officer CFO Advisor - Government regulatory authorities consumer safety officer administrator - Insurance companies Underwriter assessor

[0167] Business services: As an example, a read (event) may be logged in a registry (system) upon validation / acknowledgement by a consumer safety officer (actor) and successful completion of the read (business service). Business services may encompass critical processes and tasks that may ultimately result in events. Such services may be designed to be stateless in that they do not require any specific prior conditions to exist in order to trigger them. They may indicate that a specific event has occurred for successful completion. In either case, business services may utilize a system, but most typically involve some human involvement. As an example, in certain embodiments, this would be requested or triggered by an actor. Business services may also be specified, as well as the events they result in, for example, Validate (business service) → Validate (event).

[0168] System: As an example, when a read (event) is recorded in a registry (system), a stream processor (system) can read the newly created read (event) from the registry and broadcast it to validated / authorized listeners (systems). One of the listeners may update a notification dashboard used by the product trademark owner (actor). On the other hand, a system may only interact with other systems or otherwise with human-operated clients. In other words, a system may typically be a digital system. An example of a system in a DUID platform may be an API. An API may expose an interface to authorized actors operating outside the platform boundary. Another example of a system may be a DUID registry (i.e., database), which may serve as a persistent data store for all DUID data. The registry may not be directly exposed to external actors.

[0169] Event: As an example, a read (business service) may be requested by a consumer safety officer (actor) at the FDA (actor). After approval / validation, a successful read (event) may occur by the business service. Events may refer to the results of business services and systems. Events are typically recorded in association with a DUID; that is, an organism may be identified, verified, or read by a business service and tracked by internal or external systems. The following table outlines each event in this example and its relationship to various business services, actors, systems, and tasks.

[0170] [Table 1]

[0171] As described, the DUID platform, in this example, may include various actors, business services, events, systems, and / or tasks. All such components may follow a particular process flow. This section describes an example flow in detail. The diagrams used to illustrate such processes use BPMN 2.0 notation (BPMN2.0 - https: / / www.omg.org / spec / BPMN / 2.0 / PDF; incorporated herein by reference in its entirety). These diagrams can be used in figures described in more detail below.

[0172] Process Overview: FIG. 3 provides an overall view of the exemplary processes of the DUID ecosystem of this embodiment.

[0173] Process begins: Prior to initiating the example process, it may be expected that appropriate terms of service arrangements are in place. This may include know-your-customer (KYC) verification, e.g., proof of ownership, legal entity identification, and payment. In addition to KYC requirements, customers may be able to specify user access roles and other system / account settings via an administrative dashboard.

[0174] Generation of primers and sequencing sites This can be an ongoing / ongoing task that can occur independently of the process. The development of DUID primers can depend, for example, on the customer's host organism requirements, or R&D efforts, or both. Existing available primers can be used for identification business services.

[0175] Identification: The identification business service can be seen in more detail in Figure 4. The physical production of this business service can be a cassette based on a DNA sequence that can be used by a producer in the transformation of an organism. There can be two scenarios that can unfold in this activity.

[0176] First, if there is an existing cassette, consider using standard CRISPR and / or related technology to modify parts of the existing identifier. For example, if the existing identifier is mapped to a geographic region, a few bases at the end of the sequence can be edited. This edit can then be mapped to further specific information, for example, the expected transformation state after processing. Once this is complete, an identification event can be triggered.

[0177] If no pre-existing cassettes exist, they may be generated. See FIG. 2 for details on such a process. As shown in FIG. 2, cassettes may be generated using a pool of oligonucleotides of randomized sequences. Randomized pools of oligonucleotides may be obtained commercially or synthesized on demand. They may be constructed, for example, by enzymatic polymerization or ligation. The randomized oligonucleotide fragments may be purified, for example, by column separation to isolate fragments of approximately the same or similar size (e.g., approximately 300 nt to 400 nt in size in the example shown), which may be inserted into the cassette. A pool of cassettes containing a variety of different unique identifier sequences (i.e., approximately 10 in some examples) may be generated. 7) The cassette may include a primer annealing sequence (i.e., primer site) and at least one sequencing primer annealing sequence (i.e., sequencing site) in a suitable arrangement to allow amplification and / or sequencing of the DUID, such as the configuration shown in FIG. 2. The primers and sequencing sites may be verified against the host genome to verify the absence of natural amplification. Cassettes with different primers for different organisms or different genomes may be utilized as desired. The cassette may include a restriction enzyme array site and may be provided, for example, in the form of an insertion cassette carrier plasmid. In certain embodiments, the cassette may be about 500 bp in length and may be formed into a plasmid or carrier vector about 1200 bp in size.

[0178] Once the cassette is complete, an identification event can be triggered and the cassette can be sent to a customer. The customer is typically a producer, e.g., a grower in the agricultural industry. The producer can use suitable transformation and reproduction techniques to reproduce the organism of interest, which now contains the inserted cassette in its genome. The producer can then generate a validation package containing at least a sample of genomic DNA from the transformed biological entity, which can then be returned.

[0179] verification: After receiving the verification package, verifying the requester, and checking authorizations, the verification process can begin. An example verification process is outlined in Figure 5. The DUID can be verified for: - Stable integration into the host nuclear genome The DUID can be easily amplified from total DNA extracts. The DUID sequence may be recoverable within predictable specifications from the DUID cassette. - Specific value validity If the value already exists in the registry, the transform event may be discarded. - Number of copies of the built-in Transformation events where there are more than one copy of a DUID may be discarded (although it is contemplated that more than one DUID may be used in some embodiments). - Built-in location DUIDs can be targeted to non-coding / intergenic regions to reduce the likelihood of insertions affecting native coding regions. Additionally, the location of the DUID can be mapped to a particular chromosome and chromosome arm. - Non-expression evaluation If there is any RNA expression of the DUID, the transformation event can be discarded.

[0180] The DUID can be independently amplified by both sets of primers (e.g., using two or more sets as in the example of Figure 2), and the random ID can be sequenced. In certain embodiments, this process can be repeated three times to mitigate sequencing errors. The validation business service can utilize a step-wise flow of success or failure for each cassette validation step. In certain embodiments, this can reduce the cost of validation. If a failure occurs, the result can be recorded. If each sequence validation is successful, the result can be recorded and a recall test can be initiated.

[0181] In certain embodiments, such recall simulations may involve introducing the organic material of interest into various environmental conditions. These environments may result in variations in the organic material, which may then be read and passed to a business service. In this example, any of four parallel tests may occur: - A completely fresh environment - A completely dry environment - Simulated GI acid environment This may simulate the digestion of the substance. Potential recall from feces can be simulated. - UV ionizing radiation environment This can simulate exposure of organic materials to sunlight or other food processing sterilization techniques, such as gamma irradiation or e-beam sterilization.

[0182] Upon completion of such testing, the resulting organic matter may be independently passed to and triggered by a reading business service. After the reading business service, all results may be recorded. Not all organic matter resulting from such environmental condition testing must be successfully read for verification to be successfully completed; such determination may be made, for example, on a case-by-case basis.

[0183] Postmortem: Depending on the validation outcome, there can be a number of potential outflows. If the validation is unacceptable, the DUID service may terminate and the appropriate parties may be notified. If one of the sequence validation checks fails, a post-mortem review may be initiated. The post-mortem review may attempt to identify the cause of the failure. There can be two outcomes (cassette error or transformation error), and depending on the cause, the flow may trigger a retry of the identification business service or request a retry of the transformation from the producer.

[0184] If the result of the verification business service is a verification event, the DUID registry (i.e., database) may be updated with the appropriate information. This event may also trigger a propagation authorization message or notification, which may be received by the grower. This may then prompt the grower to produce and propagate the material, and the grower may then continue with business as usual.

[0185] Pre-read supply chain activities As described herein, the rest of the supply chain may continue business as usual. However, supply chain participants may have the option to incorporate the DUID into these existing processes. If they choose not to, the presence of the DUID may, at a minimum, provide provenance traceability. In certain embodiments, it is contemplated that the DUID may be incorporated into an existing barcode. Note that in certain embodiments, the unique identifier (UID) portion of the DUID may essentially be a string characterized by its nucleotides (A, T, G, C). In certain embodiments, if an unambiguous read is not required, they may independently track this DUID-prepared organism using their own data acquisition technology (e.g., barcoding). This may result in unconfirmed tracking events.

[0186] When a read event is required, the relevant stakeholder may send a request to the read business service. In this example, there may be two types of requests: one may be mandatory and the other may be optional. The contents of the read package may depend on the type. For example, if the read request is mandatory, there may be specific requirements that must be met to fulfill the stakeholder's requirements, such as an organic material sample on a specific date.

[0187] reading: The Read business service is shown in detail in Figure 6. As with other business services, authorization can be checked immediately. Most likely, the read package may contain various organic materials. Depending on the material, purification and / or amplification may occur. If the primer is detected, sequencing (and in some cases the UID decoding step) may begin. If the primer is not detected, the result and failure are logged.

[0188] Once the UID has been sequenced and / or decoded, an attempt may be made to find all suitable data in the DUID registry. It is anticipated that in certain cases the DUID may not be found in the registry, and in such cases, a post-mortem review may be performed. This review may attempt to find the cause of the error. On the other hand, if the DUID is found, the result may be logged and a read event may be created.

[0189] Also, in certain embodiments, authorized integration partners, such as the FDA, may require reading business services. Some jurisdictions may have regulations that may require, for example, the sharing of traceability data.

[0190] After reading After the reading business service is completed, a reading data package may be generated and returned to the requesting stakeholder. The reading package may include all tracking events to date, verification results, and primer data. It may also include contractual obligations that required the use of the DUID in the first place. It may also include KYC information for each party involved.

[0191] Support system: There may be two supporting systems described above in this example DUID global view diagram. Neither of these plays an essential role in the overall process, but instead may act as an interface and processor for the DUID registry.

[0192] API: An API may act as an interface to the DUID registry. This may allow authorized integration parties to access authorized data. In some cases, they may be able to modify this data. See User Access Roles above.

[0193] Stream Processor: A stream processor reads from the registry in real time and can trigger functionality as a result, for example, automatically notifying the DUID owner if an unauthorized actor requests a read business service.

[0194] Thus, this example details embodiments of DUID systems, methods, and compositions that may be used in accordance with the teachings provided herein. It should be understood that this example is provided for illustrative purposes directed to those skilled in the art and is not intended to be limiting. [Example]

[0195] Stable integration of DUID into yeast species Stable integration of DUID into yeast species This example describes the stable integration of a DUID into a yeast species.

[0196] Methods and Materials for Stable Integration of DUID into Yeast Species Overview: This example describes techniques for designing, integrating, and validating DNA sequence-based unique identifiers (DUIDs) into the model organism yeast. Such techniques include the use of both laboratory and industrial yeast strains. The methods of the present invention validate the utility and effectiveness of integrating DUIDs into the genome for traceability activities. Such molecular biology experimental methods include: 1. In silico design of DUID, DUID vector, and DUID primer 2. Methods for stable genomic integration using yeast centromeric plasmids (YCp) 3. Methods for stable genome integration by insertion into native yeast chromosomes 4. Methods for verifying DUID inclusion 5. Methods for DUID Signal Detection and Signal Detection Limits

[0197] We consider this method to be applicable to a wide range of research and industrial yeast strains, including prototrophic strains. The YCp approach allows for cellular and nuclear manipulation of the genome to integrate the DUID construct as an independent chromosome through spindle attachment of a centromere sequence engineered into the vector backbone. For insertion into native yeast chromosomes, four genomic sites were selected to minimize interference with the normal coding capacity and gene expression in the genome. These sites included subtelomeric regions, commonly considered heterochromatic when genes are typically silenced, and euchromatic regions with low coding capacity, which act as positive controls. For insertion into native yeast chromosomes, we focused on 1) cotransformation of a plasmid carrying antibiotic resistance for selection of transformants with a linear fragment containing the DUID flanked by homologous regions flanking the selected target site, and 2) a CRISPR-based method to target the integration site using a specific guide RNA (gRNA) and a specific homology-directed repair template (HRT) that acts as a template for the PAM site targeted by Cas9 degradation.

[0198] Construct and vector design and development DUID construct design Figure 14 shows maps of two 370 pb DUID constructs. A) Design of the DUID construct for PCR and qPCR amplification. The construct is 370 pb. This DUID construct contains two forward primers and two reverse primers. There are two identifiers (ID1 and ID2). ID1 is ideal for PCR amplification. ID2 is ideal for qPCR amplification. B) Design of the DUID construct for loop-mediated isothermal amplification (LAMP) and PCR. This map includes primers for both PCR and LAMP. Aside from traditional amplification design decisions, note the pink features, which are optional CAS PAM sites that allow for editing and detection of the DUID construct sequence using a CRISPR-based system. Such PAM sites may allow for editing of the integrated DUID construct.

[0199] Figure 17 shows the ID to registry mapping example described herein. Note that the figure shows a simplified example, and consider that full DUID sequences are typically not as short as those shown in the table.

[0200] In this example, there is no more than one alignment of an ID sequence within the database. In this example, an ID sequence is always unique to a single DUID construct, but a single DUID construct may have multiple ID sequences. However, an ID sequence may have one or more segments within it that are homologous to other DUID sequences. Sequences that can be used across a DUID construct may be present within the DUID construct, but it is important to consider that the ID itself, and by extension, the DUID, is unique. This design decision to have homologous segments within an ID sequence that span any number of DUIDs may allow for versioning of DUIDs in a number of ways.

[0201] Example of a homologous ID section - one homologous section spanning three DUIDs: There are several reasons why it may be desirable to have homologous sequences across multiple identifiers. In some cases, an identifier may have homologous sequences for the purpose of creating versions associated with this identifier. The ability to manage identifier versions may allow a user to reference an associated protocol that conveys how to interact with the DUID. For example, a particular version of a DUID identifier may contain a public key within it, in the context of cryptography, so that subsequent interactions with the DUID can be meaningfully conveyed. In other cases, the homologous sequence may reference, for example, the system or entity that originally created this identifier. The following table shows three DUIDs with such homologous segments; in this example DUID, 1:10 is homologous and 11:50 is unique.

[0202] [Table 2]

[0203] YCp and co-transformation The plasmid used in the cotransformation procedure was the yeast centromeric vector YCp41K (Taxis & Knop, 2006). Four target sites for integration were identified: a subtelomeric region of Chr 6 and a euchromatic region of chromosome 2 (Appendix C). Linear fragments targeting these sites contained the DUID flanked by 75-nt regions of homology to the regions flanking each integration site (Figure 14). The exact linear fragment sequences for each integration site are listed in Appendix D. These fragments were both synthesized as linear fragments by Twist Bioscience (https: / / www.twistbioscience.com / ) and inserted into the pRS41K (https: / / bip.weizmann.ac.il / plasmid / pics / 106.jpg) and pRS42K (https: / / bip.weizmann.ac.il / plasmid / pics / 109.jpg) vectors.

[0204] Generation of linear DUID fragments for co-transformation Linear DNA fragments for homologous recombination (HR) were generated by PCR using the linear fragments generated by Twist Bioscience as templates. For the specific fragments generated, see Appendix A under "Co-transformation" below. pRS41K-Chr6 and pRS41K-Euch, provided by Twist Bioscience, were used. The primers used to generate the HR fragment were Chr6_DUID F and DUID-synthR for the Chr6 target region, and EuchDUID F and DUID-synthR for the Euch target region, respectively (Appendix A). The PCR reaction composition (Table 2) and reaction conditions (Table 3) are detailed below.

[0205] [Table 3]

[0206] [Table 4]

[0207] Primers were validated and annealing temperatures were optimized using a 20 μL reaction volume. For generation of HR linear integration fragments, two 50 μL reactions were performed and the products were purified using the Qiagen PCR Purification Kit (https: / / www.qiagen.com / ie / shop / pcr / qiaquick-pcr-purification-kit / ). The purified DNA fragments were eluted in 50 μL of elution buffer (10 mM Tris-Cl, pH 8.5). The products were verified by electrophoresis of 5 μL on a 1% agarose gel.

[0208] CRISPR vector and HRT generation CRISPR experiments were performed using the plasmid pCC-036, which contains CAS9 expressed by TDH3p, SNR52p to drive gRNA expression, and hygR for selection against hygromycin, as described in Krogerus et al., 2019. Three gRNAs were designed for each of the target integration sites using Benchling software (https: / / www.benchling.com / ). Primers containing the gRNA sequences (Appendix B) were used in PCR reactions using pCC-036 as a template. The reaction composition and conditions are outlined below (Tables 3 and 4). These PCR reactions were transformed into E. coli. Plasmids were isolated from transformants and screened by sequencing to confirm correct clones (Figure 14 and Appendix B). One gRNA clone was constructed for Chr6 (Chr6_2) and two for Euch (Euch_1; Euch_2). The primers were designed to partially overlap the mutation in the center of both primers (approximately 8-10 bp at both ends were non-overlapping), and PCR was performed according to the protocol of Zheng, et al., 2004.

[0209] [Table 5]

[0210] [Table 6]

[0211] Primers were validated and annealing temperatures were optimized using a 20 μL reaction volume. For integration, five 50 μL reactions were performed, followed by digestion of the vector with HindIII and BamHI (NEB). DNA was purified using phenol / chloroform / isoamyl alcohol followed by ethanol precipitation in the presence of 0.1 M ammonium acetate and glycogen. DNA was resuspended in 30 μL of nuclease-free water. Amplification was verified by electrophoresis of 5 μL on a 1% agarose gel.

[0212] [Table 7]

[0213] [Table 8]

[0214] After PCR, 10 μL was electrophoresed on a 1% agarose gel (the yield of SDM DNA by Phusion is low). 10 μL of PCR amplification product in a 30 μL reaction volume was digested with DpnI (NEB) overnight at 37°C to linearize the methylated template DNA. Another 10 μL was separated on a gel, and then 5 μL was transformed into E. coli. Minipreps were performed on 12 colonies and sequenced.

[0215] Yeast transformation: Transformation of YCp-DUID vector Both the CRISPR plasmid and repair template were transformed into the target strain using a standard lithium acetate-based yeast transformation protocol as described by Mertenes et al. 2017 and completed as follows: 1. Yeast was grown overnight in 100 mL of 2% YPD growth medium at 30°C to an OD of approximately 0.7-0.8. 2. The yeast cell culture was then centrifuged (3000 rpm for 3 minutes), washed once with sterile water, and the cells were resuspended in 200 μL of 0.1 M lithium acetate solution. 3. After 10 minutes of incubation at room temperature, 50 μL of cell culture was mixed with 500 ng of plasmid, 300 μL of PLI (142 M polyethylene glycol, 0.12 M lithium acetate, 0.01 M Tris (pH 7.5), and 0.001 M EDTA), and 5 μL of salmon sperm DNA (1 mg mL). A negative control transformation without DNA (sterile water) was performed in parallel. 4. The yeast suspension was incubated at 42°C for 30 minutes. 5. Cells were centrifuged (3000 rpm for 3 minutes) and resuspended in fresh YPD 2%, after which the cells were allowed to recover during one overnight incubation at 30°C. 6. 200 μL of the yeast suspension was plated onto YPD + G418 300 μg / mL and then incubated for 2 days at 30° C. 200 μL was also plated onto YPD without antibiotics to check cell viability after such treatment.

[0216] Co-transformation The method involved DNA transformation using electroporation after generation of competent cells with lithium acetate, as described in Bernardi et al., 2019 .

[0217] [Table 9]

[0218] Co-transformation step 1. Cells were grown in 100 mL of YPD with shaking to the desired growth phase (based on growth curve or OD). 2. Mid-logarithmic growth phase (OD 600 The cells were harvested at 0.7-0.8°C. The culture was centrifuged and the supernatant was discarded. 3. The pellet was washed once with sterile water. The culture was spun down, the supernatant was discarded, and the pellet was resuspended in 25 mL of 0.1 M lithium acetate / 10 mM DTT / 10 mM TE solution (TrisHCl:EDTA=10:1). The culture was incubated at room temperature for 1 hour. The culture was spun down, and the supernatant was discarded. 4. Note: When working with clumping strains, be sure to invert the tube several times every 10 minutes to prevent cells from settling to the bottom of the tube. 5. The pellet was washed with 25 mL of ice-cold distilled sterile water, the culture was spun down at 4°C and the supernatant was removed. The step was repeated (for a total of two washes). 6. The precipitate was washed with 10 mL of ice-cold sorbitol, centrifuged at 4°C, and the supernatant was removed. The precipitate was resuspended in 100 μL of ice-cold sorbitol. 7. 100 μL of cell suspension was used for transformation. 8. 15 μL of transforming DNA (1 μg of pRS41K [YCp plasmid] + 1 μg of linear DUID fragment; molar ratio 1:10; and molar ratio 1:20) was mixed with the cell suspension and incubated on ice for 5 minutes. 9. The cell suspension was electroporated in a 0.1 cm cuvette at 1.8 kV. 10. Add 1 mL of chilled sorbitol to the electroporation cuvette and mix with the cell suspension. The suspension was transferred to a tube containing 300 μL of YPD. 11. Note: If you are using an antibiotic marker, you will want to incubate the suspension at 30°C for 3 hours to allow expression of the antibiotic to occur. *Do not add antibiotic to this culture. This addition will kill all cells, as they are not yet expressing the plasmid that confers antibiotic resistance.* 12. 100 μL of the transformed culture was plated onto selective (YPD+300 mg / L G418) plates and incubated at 30° C. for 5 days to allow colonies to appear. 13. The transformed culture was also plated onto YPD plates without any marker / antibiotics to ensure that the cells were viable.

[0219] Transformation with CRISPR vectors and HRT Both the CRISPR plasmid and repair template were transformed into the target strain using a standard lithium acetate-based yeast transformation protocol, as described in Mertens et al., 2019. The protocol described below is based on a standard transformation procedure, in which cells are made competent by treatment with LiOAc solution, then incubated with DNA molecules (plasmid and repair template) and carrier DNA (salmon sperm DNA), and then heat-shocked to allow DNA absorption. After recovery, the cells are plated on hygromycin to select against all non-transformed cells. Plating on YPD without hygromycin demonstrates cell growth following the transformation procedure; for example, the procedure itself did not kill the cells. Transformation of the CRISPR plasmid without the HRT should result in cell death, as the DSB is not repaired. This confirms successful function of the CRISPR plasmid, indicating that Cas9 is expressed and the gRNA is targeted to the genome by Cas9. Transformation with the CRISPR plasmid and the HRT should repair the DSB and support cell growth.

[0220] Plasmids pCC-036_Chr6_2 / Chr6_HRT and pCC-036_Euch_1 / Euch-HRT were the respective combinations of DNA molecules transformed into yeast strains S288c, Vermont and French Saison using the following protocol: 1. Yeast was grown overnight in 5 mL of YPD at 30°C and 200 rpm, after which 1 mL of the preculture was transferred to 50 mL of YPD and incubated for an additional 4 hours (30°C, 200 rpm). 2. The yeast cell culture was then centrifuged (3 minutes at 3000 rpm) and the cells were resuspended in 200 μL of 0.1 M lithium acetate solution. 3. After 10 min of incubation at room temperature, mix 50 µL of cell culture with 500 ng of plasmid, in this case the corresponding sgRNA, 5–25 µg of HRT DNA (prepared protocol), 300 µL of PLI (142 M polyethylene glycol, 0.12 M lithium acetate, 0.01 M Tris (pH 7.5), and 0.001 M EDTA) and salmon sperm DNA (1 mg mL -1 ) Cloned both with and without 5 μL. Incubated at 4.42°C for 30 minutes. 5. The cells were centrifuged (3000 rpm for 3 minutes) and resuspended in fresh YPD, after which the cells were allowed to recover overnight at 30°C. 6. A volume of 200 μL of the yeast suspension was plated onto YPD containing 300 mg / L hygromycin and then incubated at 30°C for 3 to 5 days.

[0221] Screening of transformants Genomic DNA extraction protocol 1. Plate duplicate co-transformation plates on YPD+G418 (300 mg / L). 2. Split the primary plate into 4-8 colonies per split. Scrape the colonies into sterile tubes with 3 mL of YPD and grow overnight at 30°C with shaking. Precipitate 2 mL of culture into a 3.2 mL screw-cap tube. 4. Wash once with 1 mL of MQ water. 5. Resuspend in 200 μL of Disintegration Buffer (2% TX-100, 1% SDS, 100 mM NaCl, 100 mM Tris pH 7.5). 6. Add 200 μL of glass beads and 200 μL of phenol / chloroform / isoamyl alcohol. 7. Vortex on high for 3 minutes. 8. Centrifuge at maximum speed for 5 minutes. 9. Transfer the top aqueous layer to a clean microcentrifuge tube. 10. Add 1 mL of 100% EtOH and mix by inversion. 11. Centrifuge at maximum speed for 3 minutes. 12. Decant the ethanol, dry the precipitate and resuspend in 400 μL of 1×TE. Add 30 μL of 13.1 mg / mL RNase A. 14. Incubate at 37°C for 5 minutes. Add 10 μL of 15.4 M ammonium acetate and 1 mL of 100% EtOH. Mix by inversion. 16. Centrifuge at maximum speed for 3 minutes. Wash the pellet with 1 mL of 70% EtOH and dry. 17. Resuspend in 100 μL of water.

[0222] Identification of integrants and PCR screening of transformants: The gDNA isolated as described above served as a template for PCR reactions using primers that bound genomic DNA at specific regions upstream and downstream of the HRT homology region flanking the target integration site (see Appendix A for primer details; reaction compositions and conditions are outlined in Tables 9 and 10 below). For the euchromatic target integration site on Chr2, primers Euch_SeqF / R were used, and for the subtelomeric heterochromatic target integration site on Chr6, primers Chr6_SeqF / R were used. These primers generated a DNA fragment of approximately 600 bp from gDNA that did not contain any insertion at the integration site. Upon integration, this fragment size increases to approximately 970 bp. Controls included reactions without any gDNA template, and gDNA was isolated from untransformed strains (e.g., S288c / BY4743). PCR reactions were separated by gel electrophoresis using GeneRuler 100 bp Plus molecular weight markers to confirm the size of the generated DNA fragments.

[0223] Once an integrant was identified, correct integration was verified by PCR primers, one binding to the genome outside the integration fragment and the other binding within the transformation fragment, which yield a DNA fragment if integration had occurred at the correct target site and no DNA fragment if integration had not occurred.

[0224] The DNA fragments generated by both the integration confirmation and verification assays are sequenced to confirm integration.

[0225] [Table 10]

[0226] [Table 11]

[0227] After PCR, 10 μL of the reaction was separated on a 1% agarose gel (1× TAE, containing SYBR Safe nucleic acid stain).

[0228] Confirmation of insert copy number and position WGS was performed on the parent and integrants to confirm the insert copy number and identify any off-target integration events. Short-read (Illumina) and long-read (PacBio) sequencing data were combined to construct the genomes of both the parent and transformed strains. The combination of these two approaches generated the overall genomic structure of the integrant, thereby identifying whether multiple insertions occurred or whether any off-target integration events were present. The entire genomes of the integrant(s) and parent strain(s) were sequenced by Genome Quebec (Montreal, Canada) as previously described (Preiss et al., 2018). Briefly, DNA was isolated and used as a template for library construction for Illumina and PacBio applications. Sequencing reads were quality analyzed using FastQC (version 0.11.5) (Andrews, 2010) and trimmed and filtered using Trimmomatic (version 0.36) (Bolger, Lohse, & Usadel, 2014). Reads were aligned to the S. cerevisiae S288c (R64-2-l) reference genome using SpeedSeq (0.1.0) (Chiang et al., 2015). Alignment quality was assessed using QualiMap (2.2.1) (Garcia-Alcalde et al., 2012). Variant analysis was performed on aligned reads using FreeBayes (1.1.0-46-g8d2b3a0l) (Garrison & Marth, 2012). Variants from all strains are called simultaneously (multiple samples). Prior to variant analysis, alignments are filtered to a minimum MAPQ50 using SAMtools (1.2) (Li et al., 2009). Variant annotation and effect prediction are performed using SnpEff (1.2) (Cingolani et al., 2012).Chromosome and gene copy number variation is estimated based on coverage by Control-FREEC (11.0) (Boeva ​​et al., 2012). Statistically significant copy number variation is identified using the Wilcoxon rank sum test (p<0.05). Median coverage and windows of heterozygous SNP numbers greater than 10,000 bp are calculated by BEDTools (2.26.0) (Quinlan & Hall, 2010) and visualized in R.

[0229] Determining the expression of DUID in integrants using droplet digital PCR We used droplet digital PCR (ddPCR), which allows for the quantification of absolute numbers of molecules within a sample. This allows for the quantification of copy number or low-expression genes in particular. The procedure involves the isolation of gDNA-free RNA from yeast, followed by cDNA synthesis and the final generation of S288c with and without the pRS41K-Euch plasmid. The integrant strains were grown in triplicate in YPD. RNA was extracted using the commonly used hot acid phenol method (Collart and Olivierro 2001) and quantified using a NanoDrop 2000C spectrophotometer (NanoDrop Technologies Inc.). RNA samples were treated with a RapidOut DNA Removal Kit (Thermo Fisher), tested for DNA contamination, and quality was assessed using an Agilent 2100 Bioanalyzer. RNA (1000 ng per sample) was used to generate cDNA using the High Capacity cDNA Reverse Transcription Kit (Applied BioSystems).

[0230] These samples and diluted pRS41K-Euch were subjected to ddPCR analysis at the University of Guelph Genomics Facility. These samples and a "no-template control" were used as templates in ddPCR reactions using ddPCR EvaGreen Supermix (which allows emulsification) and qPCR primers for DUID, GAT3 (low-expression control), and ACT1 (high-expression control) for all reactions. Nanoliter-sized droplets were generated on an AutoDG™ Instrument (Bio-Rad), and PCR amplification was then performed using a C1000 Touch thermal cycler (Bio-Rad). After PCR cycling, the ddPCR plate was read using a Bio-Rad QX200 Droplet Reader, and data were analyzed using QuantaSoft Analysis Pro software version 1.0.596 (Bio-Rad Laboratories).

[0231] Protocol for LOD / LOQ analysis: gDNA prepared using the gDNA isolation protocol was used to screen for inserts. The vector was prepared using a QiaQuick miniprep kit from DH5αK12 cultures grown in the presence of ampicillin.

[0232] Standard PCR Protocol Primers: S288C DUID F and R Dilution series: 100ng, 10ng, 1ng, 100pg, 10pg, 1pg, 100fg, 10fg, 1fg, 100ag

[0233] [Table 12]

[0234] [Table 13]

[0235] After PCR, 10 μL of the reaction was separated on a 1% agarose gel (1× TAE, containing SYBR Safe nucleic acid stain).

[0236] Quantitative PCR (qPCR) protocol: qPCR reactions were performed on a StepOnePlus Real-Time PCR system using SensiFAST Hi-ROX SYBR Master Mix by the University of Guelph AAC Genomics Facility. qPCR cycling conditions are listed in Table 12. Analysis was completed using Applied Biosystems StepOnePlus software. gDNA was prepared using the gDNA isolation protocol described above. Control DUID vectors were prepared using a QiaQuick miniprep kit from DH5αK12 cultures grown in the presence of ampicillin.

[0237] Amplification was performed on both the plasmid and YCp yeast gDNA samples across a dilution series using the following primers: Primers: S288C DUIDqPCR F and R Dilution series: 100ng, 10ng, 1ng, 100pg, 10pg, 1pg, 100fg, 10fg, 1fg, 100ag

[0238] Results and Discussion Verification of transformation: DUID was stably transformed into the genome of a yeast strain (BY4743) using the YCp vector. The transformed yeast was cultured and genomic DNA was extracted as described above. Stable integration. Cingolani P, Platts A, Wang le L, Coon M, Nguyen T, Wang L, Land SJ, Lu X, Ruden DM. SnpEff, a program for annotating and predicting the effect of single nucleotide polymorphisms: SNPs in the genome of Drosophila melanogaster fly strain (w1118; iso-2; iso-3, Austin. 2012 Apr-Jun;6(2):80-92. doi: 10.4161 / fly.19695. PMID: 22728672; PMCID: PMC3679285.on) were verified by end-point PCR (Figure 15, B1-B3) and qPCR (Figure 16). Endpoint analysis of PCR amplification using DUID recall primers flanking the DUID construct resulted in robust amplification for both the YCp-DUID vector (Figure 15A) and yeast genomic DNA extracts from cells transformed with the YCp-DUID vector (Figure 15C). When 100 pg to 100 ng of YCp-DUID vector were used as template (Figure 15A, lanes 1-4), a 370 bp band indicating DUID amplification was clearly visible, whereas with any input DNA amount, with untransformed BY4743 genomic DNA, there was no detectable amplification (Figure 15B, lanes 1-8). A similar assay using total DNA isolated from cells transformed with the YCp-DUID vector yielded robust amplification of input DNA ranging from 1 to 100 ng, with very faint signals detected from 100 pg of input DNA (Figure 15C, lanes 1-4), demonstrating that DUID, present at 1-2 copies per cell and whose copy number reflects that of a chromosomal feature, can be readily detected in yeast gDNA isolates using standard end-point PCR procedures.

[0239] Figure 15 shows detection of YCp-DUID in yeast genomic DNA by end-point PCR. PCR amplification was performed using (A) the YCp-DUID vector, (B) gDNA extracted from BY4743, and (C) yeast strain BY4743 transformed with the YCp-DUID vector as templates with DUID recall primers. Reactions were performed using serially diluted DNA template with input amounts of (1) 100 ng, (2) 10 ng, (3) 1 ng, (4) 100 pg, (5) 10 pg, (6) 1 pg, (7) 100 fg, and (8) 10 fg, and were separated on a 1% agarose gel using GeneRuler™ 100 bp Plus Ready-to-use Ladder as a standard.

[0240] LOD / LOQ analysis Quantitative real-time PCR was performed using serial 10-fold dilutions of purified YCp-DUID vector (Figure 16), demonstrating that such assays detect DUID amplification at all measured concentrations, thereby enabling reliable identification of DUID down to concentrations as low as 500 ag. A standard curve was generated using MS Excel by plotting the mean Cq values ​​versus known DNA input concentrations. Based on this standard curve, R 2The Cq was calculated as 0.9993, yielding a primer efficiency of 105.5% (calculated using the Agilent QPCR Standard Curve to Slope Efficiency calculator, https: / / www.chem.agilent.com / store / biocalculators / calcSlopeEfficiency.jsp?_requestid=1116919), indicating that the reaction efficiency was within acceptable standards for high-quality qPCR analysis (https: / / www.gene-quantification.de / roche-rel-quant.pdf). A similar qPCR assay using DNA isolated from BY4743 transformed with YCp-DUID demonstrated that the DUID was detectable in 50 ng of total yeast DNA with an average Cq value of 29.02. This Cq value was plotted against the standard curve (Figure 3; orange bar). These results validate that the DUID recall method can amplify the DUID from yeast cell culture matrix.

[0241] These results demonstrated the following: 1. DUID was successfully designed and can be stably transformed into yeast. 2. For traceability purposes, DUIDs can be recalled from biological matrices by both standard end-point PCR and qPCR techniques.

[0242] Figure 16 shows the detection of DUID in yeast total DNA extracts. Quantitative real-time PCR was performed on serial 10-fold dilutions of the YCp vector ranging from 50 ng to 500 ag and used to generate a standard curve (blue line) using MS Excel. Results from a similar qPCR experiment using DNA from BY4743 transformed with the YCp-DUID vector were plotted (orange bars) and compared to the standard curve values ​​to quantify the detection of DUID in yeast biomass.

[0243] [Table 14] JPEG0007803542000015.jpg103170

[0244] [Table 15] JPEG0007803542000017.jpg188170

[0245] Appendix C: Homologous Recombination Constructs The entire sequence containing the homologous arms for use in homologous recombination Yeast chromosomal integration site ChrII:809650..809799

[0246] Left homologous arm: ACACAAACTGGCGTAGAAGGGGAAACGGAAATAGGGTCTGACGAGGAAGATAGCATAGAGGACGAGGGAAGCAGC (SEQ ID NO: 47)

[0247] Right homologous arm: AGTGGAGGAAATAGTACGACAGAAAGACTAGTACCACACCAGCTGAGGGAACAAGCAGCCAGACATATAGGAAAA (SEQ ID NO: 48)

[0248] ChrVI:261123..261272

[0249] Left homologous arm: TGGAGTTGCAAAAAACAAGGGAAAGGAAAATCAATCAAATTAGAATTAAGGTTTTTTTTGGACAGTGCAGCGTCA (SEQ ID NO: 49)

[0250] Right homologous arm: ATGCGCACGTAATGGCTTCGAAGAAAAAAAGAAGGCAAATACAATGAAGCTGAGATCTTGTTTTATCATGAGGGG (SEQ ID NO: 50)

[0251] ChrXIV:764201..764350

[0252] Left homologous arm: CAAATAAATTAGGCTCATAACCGTAATTTTATTCGAGACATTTTTGGTTACTTCAAAATATTGTTATTATATAAA (SEQ ID NO: 51)

[0253] Right homologous arm: GATCATATAAAGTTCTTGGACAAGATTGGATACATTTAGTTTTATTTTTGAAAATCACAAAGATGAAACAAAATA (SEQ ID NO: 52)

[0254] [Table 16]

[0255] S288c

[0256] [Table 17]

[0257] [Table 18]

[0258] s288c-chromosome 2 ACACAAACTGGCGTAGAAGGGAAACGGAAATAGGGTCTGACGAGGAAGATAGCATAGAGGACGAGGGAAGCAGCGCGTGATGGTTAGGCGTACACGAATCCTGGTTTCAACGCGCTGCAAACCTACCCCTGCTCCAAACTGCTGTTCAACGCCACTCTAACTGGCAGGCAAATTATTAGTTTCCAAGTTCCCAGGTGCTGAAGAGCAGTCATCAACGCCCTCAGTACATCCAGTCCATCCCGGCGTTTGTCCGGAGGATCGTGTCGTACAACAACAACCATCTGACTATCAACCCTCCaggCGTATAGAGCGGGTCATCGATGCGCTCAGGGAACAACAACGATAGCCTGCGGGCTGGTCACCATCGGGAAGTTTTGCTGGGAGAATCTGCTGCTTGTAGGaggTCTCTACAGCCAAACGAACCAAGTCATTTCCAGGAGTGGGAGGAAAATAGTACGACAGAAAGACTAGTACCACCACGAGCTGGGAAACAACAAGCCAGCCAGAGCCAGACATATAGGAAAA (sequence number 57)

[0259] s288c- chroma6 TGGAGTTGCAAAAAACAAAGGGAAAGGAAAATCAAATTAGAATTAAGGTTTTTTTTGGACAGTGCAGCGTCAGCTGATGGTTTAGGCGTACACGAGATCCTGGTTCAACGCGCTGCAAACCTACCCCTGCTCCAAACTGCTGTTCAACGCCACTCTAACTGGCAGGCAAATTATTAGTTTCCAAGTTCCCCAGGTGCTGGAAGAGCAGTCATCAACGCCCTCAGTAGATCATCATCGCCCTCAGGGAAACAACAACGATAGGCCTGCGCGGTTTGTCCGGAGGAATCGTGTCGTACAACAACAACCATCTGACTACTCAACCCTCCaggCGTATAGAGCGGGTCATCGATGCGCTCAGGGAACAACAACGATAGGCCTGCGGGCTGGTCACCATCGGGAAGTTTTGCTGGGAGAATCTGCTGTGTAGGaggTCTCTACAGCCAAACGACCAGACCAAGTGCATTTCCAGGGATGCGCACGTAATGGCTTCGAAGAAAAAAAGAAGGCAAATACAATGAAGCTGAGATCTGTTTTATCATGAGGGG (sequence number 58)

[0260] s288c- chroma14 CAAATAAATTAGGCTCATAACCGTAATTTTATTCGAGACATTTTTGGTTACTTCAAAATATTGTTATTATATAAAGCTGATGGTTTAGGCGTACACGAGATCCTGGTTCAACGCGCTGCAAACCTACCCTGCTCCAAACTGCTGTTCAACGCCACTCTAACTGGCAGGCAAATTATTAGTTTCTAAGTTCCCCAGGTGCTGAAGAGCAGTCATTCAACGCCCTCAGATCATCCCGGCAAGTTGGCTGGCGCGTTTGTCCGGAGGATCGTGTCGTACAACAACCATCTGACTATCAACCCTCCaggCGTATAGAGCGGGTCATCGATGCGCTCAGGGAACAACAACGATAGGCCTGCGGCTGGTCACCATCGGGAAGTTTTGCTGGAGATCTGCTGCTGTAGGaggTCTCTACAGCCAAACGACCAGACCAAGTGCATTTCCAGGGGATCATATAAAGTTCTTGGACAAGATTGGATACATTTAGTTTTATTTTTGAAAATCACAAAGATGAAACAAAATA (SEQ ID NO: 59)

[0261] Vermont

[0262] [Table 19]

[0263] [Table 20]

[0264] Vermont - Chromosome 2 ACACAAACTGGCGTAGAAGGGAAACGGAAATAGGGTCTGACGAGGAAGATAGCATAGAGGACGAGGGAAGCAGCACTTCCCCATTAGTCGGCAGCACGTTCGCCAGTATTACCGGACAGAAAATCTCGGAACAGTTTATCCGCAATTCTGAGGAAATCGTCGTCCCGCAAGCTCCGTGCACAGCTAGTAGTAGTCTCCGGTGCGGGGGGGGCGGAGTGGTCTCCCACGATACGCAGTTGTCTAGATACGTACCCACCCGCTGTTGTCTCTGGCTATCTGAACGTCACTCAGAaggGGCCCTATCAGTACAGCAGTCATAGCCGCACACAAGTCCCCCCAAACCTCCTGACCACGTCGCCACCGGCGCACACACTATTTCTCGTaggTTCATTCTCTCGCCAGCACTTGTCGGAACAAAGCGGTCTTAGTGGAGAAAATAGTACGACAGAAAGACTAGTACCACCACGCTGGGAAACAAGCAGCCAGACATATAGGAAAA (sequence number 64)

[0265] バーモント- chroma6 TGGAGTTGCAAAAAACAAGGGAAAAGGAAAATCAAATTAGAATTAAGGTTTTTTTTGGACAGTGCAGCGTCAACTTCCCCATTAGTCGGCAGCAGCAGTTCGCCAGTATTACCGGACAGAAAAAATCTCGGAACAGTTTATCCGCAATTCTGAGGAAATCGTCCGCAAGCTCCGTGCACAGCTAGTAGTAGTCTCCGGTGCGGTGGGGGGGCGGAGTGGTCTCCCACGATACGACGTTGTCTAGATACGTACCCACCCCCGCTGTTGTCTCTGGCTATCTGAACGTCCTCGCAAGGGGCCCTATCAGTACAGCAGTCATAGCCGCACACAAGTCCCCCCAAACCTCCTGACCACGTCGCCACCGGCGCAGACTATTTCTCGTaggTTCATTCTCTCGCCAGCACTTGTCGGAACAAAGCGGTCTTATGCGCACGTAATGGCTTCGAAGAAAAAGAAAGAAAGCAAATACAATGAAGCTGAGATCTGTTTTATCATGAGGGG (sequence number 65)

[0266] バーモント- chroma14 CAAATAAATTAGGCTCATAACCGTAATTTTATTCGAGACATTTTTGGTTACTTCAAAATATTGTTATTATATAAAACTCTCCCATTAGTCGGCAGCACGTTCGCCAGTAATTACCGGAGACAGAAAAATCTCGGAACAGTTTATCCGCAATTCTGAGGAAATCGTCGTCCGCAAGCTCCGTGCACAGCTAGTAGTAGTCTCCGGTGCGGGGGGGGGCGGAGTGGTCTCCCACGATACGACGTTGTCTAGATACGTACCCACCTCGCTGTGTGCTCTCTGGCTATCTGAACGTCCACTCCAGAaggGGCCCTATCAGTACAGCAGTCATAGCCGCACACAAGTCCAACGTCCCCCAAACCTCCTGACCACGCAGTCGCCACCGGCGCAGACACTATTTCTCGTaggTTCATTCTCTCGCCAGCACTTGTCGGAACAAAGCGGTCTTGATCATATAAAGTTCTTGGACAAGATTGGATACATTTAGTTTTATTTTTGAAAATCACAAAGATGAAACAAAATA (SEQ ID NO: 66)

[0267] French season

[0268]

Table 21

[0269]

Table 22

[0270] French - Chromosome 2 ACACAAACTGGCGTAGAAGGGGAAACGGAAATAGGGTCTGACGAGGAAGATAGCATAGAGGACGAGGGAAGCAGCGCGTACAATGCCCTGAAGAATTACTTCCGTACTGGAAGCGGATAGCACCAGACTGTAAGCTAACGAACGCCTGTTTGAGGCTCAGTCTGCTAAATTGGAACCGCGTCGCTCCTAGGCATATTTTGGTGAAAGCACTCTGCCCAAAAGCCTGTAGAATTCCGGACCGACGCTCTCTTCACTCGAAGATTCCGGGTAAGAAGTTTCAGCCAGGGCTGTCTCCATTAGAAaggAGCGGGTCATCGAAAGGTTACGTTGGTTGTATCTGATTAGACGGTAGACATCCAGCTCATCTCTGATTACTAAAGTTCTCCGCCGCTCCATCGGGCGaggTACAGCCAAACGACCAAGTGCCAAGTGCATTTCCAGGGAGAGTGGAGGAAATAGTACGACAGAAAGACTAGTACCACACCAGCTGAGGGAACAAGCAGCCAGACATATAGGAAAA (SEQ ID NO: 71)

[0271] French - Chromosome 6 TGGAGTTGCAAAAAACAAGGGAAAGGAAAATCAATCAAATTAGAATTAAGGTTTTTTTTGGACAGTGCAGCGTCAGCGTACAATGCCCTGAAGAATTACTTCCGTACTGGAAGCGGATAGCACCAGACTGTAAGCTAACGAACGCCTGTTTGAGGCTCAGTCTGCTAAATTGGAACCGCGTCGCTCCTAGGCATATTTTGGTGAAAGCACTCTGCCCAAAAGCCTGTAGAATTCCGGACCGACGCTCTCTTCACTCGAAGATTCCGGGTAAGAAGTTTCAGCCAGGGCTGTCTCCATTAGAAaggAGCGGGTCATCGAAAGGTTACGTTGGTTGTATCTGATTAGACGGTAGACATCCAGCTCATCTCTGATTACTAAAGTTCTCCGCCGCTCCATCGGGCGaggTACAGCCAAACGACCAAGTGCCAAGTGCATTTCCAGGGAGATGCGCACGTAATGGCTTCGAAGAAAAAAAGAAGGCAAATACAATGAAGCTGAGATCTTGTTTTATCATGAGGGG (SEQ ID NO: 72)

[0272] French - Chromosome 14 CAAATAAATTAGGCTCATAACCGTAATTTTATTCGAGACATTTTTGGTTACTTCAAAATATTGTTATTATATAAAGCGTACAATGCCCTGAAGAATTACTTCCGTACTGGAAGCGGATAGCACCAGACTG TAAGCTAACGAACGCCTGTTTGAGGCTCAGTCTGCTAAATTGGAACCGCGTCGCTCCTAGGCATATTTTGGTGAAAGCACTCTGCCCAAAAGCCTGTAGAATTCCGGACCGACGCTCTCTTCACTCGAAG ATTCCGGGTAAGAAGTTTCAGCCAGGGCTGTCTCCATTAGAAaggACGGGGTCATCGAAAGGTTACGTTGGTTGTATCTGATTAGACGGTAGACATCCAGCTCATCTCTGATTACTAAAGTTCTCCGCCG CTCCATCGGGCGaggTACAGCCAAACGACCAAGTGCCAAGTGCATTTCCAGGGAGGATCATATAAAGTTCTTGGACAAGATTGGATACATTTAGTTTTATTTTTGAAAATCACAAAGATGAAACAAAATA (Sequence number 73)

[0273] One or more exemplary embodiments have been described by way of example, and it will be understood by those skilled in the art that numerous variations and modifications can be made thereto without departing from the scope of the invention, as defined in the appended claims.

[0274] (References) 1 FDA. (2019). Investigation Summary: Factors Potentially Contributing to the Contamination of Romaine Lettuce Implicated in the Fall 2018 Multi-State Outbreak of E. coli O157:H7. Retrieved from https: / / www.fda.gov / media / 120722 / download 2 Food and Drug Regulations C.R.C. c. 870 (2019) 3 GS1 US. (2013). Integrated Traceability in Fresh Foods: Ripe Opportunity for Real Results. Retrieved from https: / / www.gs1us.org / DesktopModules / Bring2mind / DMX / Download.aspx?Command=Core_Download&EntryId=598 4 Introduction of Organisms and Products Altered or Produced Through Genetic Engineering Which Are Plant Pests or Which There Is Reason to Believe Are Plant Pests C.F.R. §340.1 (2019) 5 WHO, Foodborne Disease Burden Epidemiology Reference Group. 2015. WHO Estimates of the Global Burden of Foodborne Diseases. Retrieved from https: / / academicanswers.waldenu.edu / faq / 73164 6 System of centromeric, episomal, and integrative vectors based on drug resistance markers for Saccharomyces cerevisiae. Christof Taxis and Michael Knop EMBL, Heidelberg, Germany. BioTechniques 40:73-78 (January 2006) doi 10.2144 / 000112040 7 Krogerus, K., Magalhaes, F., Kuivanen, J. et al. A deletion in the STA1 promoter determines maltotriose and starch utilization in STA1+ Saccharomyces cerevisiae strains. Appl Microbiol Biotechnol 103, 7597-7615 (2019). https: / / doi.org / 10.1007 / s00253-019-10021-y 8 Zheng L, Baumann U, Reymond JL. An efficient one-step site-directed and site-saturation mutagenesis protocol. Nucleic Acids Res. 2004;32(14):e115. Published 2004 Aug 10. doi:10.1093 / nar / gnh110 9 Mertens S, Steensels J, G. B, V Kevin J. Rapid Screening Method for Phenolic Off-Flavor (POF) Production in Yeast. J Am Soc Brew Chem. 2017; 75(4):318-23 10 Mertens S, Gallone B, Steensels J, et al. Reducing phenolic off-flavors through CRISPR-based gene editing of the FDC1 gene in Saccharomyces cerevisiae x Saccharomyces eubayanus hybrid lager beer yeasts [published correction appears in PLoS One. 2019 Oct 24;14(10):e0224525]. PLoS One. 2019;14(1):e0209124. Published 2019 Jan 9. doi:10.1371 / journal.pone.0209124 11 Mertens S, Gallone B, Steensels J, et al. Correction: Reducing phenolic off-flavors through CRISPR-based gene editing of the FDC1 gene in Saccharomyces cerevisiae x Saccharomyces eubayanus hybrid lager beer yeasts. PLoS One. 2019;14(10):e0224525. Published 2019 Oct 24. doi:10.1371 / journal.pone.0224525 12 Bolger AM, Lohse M, Usadel B. Trimmomatic: a flexible trimmer for Illumina sequence data. Bioinformatics. 2014 Aug 1;30(15):2114-20. doi: 10.1093 / bioinformatics / btu170. Epub 2014 Apr 1. PMID: 24695404; PMCID: PMC4103590 13 Chiang C, Layer RM, Faust GG, Lindberg MR, Rose DB, Garrison EP, Marth GT, Quinlan AR, Hall IM. SpeedSeq: ultra-fast personal genome analysis and interpretation. Nat Methods. 2015 Oct;12(10):966-8. doi: 10.1038 / nmeth.3505. Epub 2015 Aug 10. PMID: 26258291; PMCID: PMC4589466 14 Garcia-Alcalde F, Okonechnikov K, Carbonell J, Cruz LM, Gotz S, Tarazona S, Dopazo J, Meyer TF, Conesa A. Qualimap: evaluating next-generation sequencing alignment data. Bioinformatics. 2012 Oct 15;28(20):2678-9. doi: 10.1093 / bioinformatics / bts503. Epub 2012 Aug 22. PMID: 22914218 15 Erik Garrison and Gabor Marth 2012. Haplotype-based variant detection from short-read sequencing 16 Li H, Handsaker B, Wysoker A, Fennell T, Ruan J, Homer N, Marth G, Abecasis G, Durbin R; 1000 Genome Project Data Processing Subgroup. The Sequence Alignment / Map format and SAMtools. Bioinformatics. 2009 Aug 15;25(16):2078-9. doi: 10.1093 / bioinformatics / btp352. Epub 2009 Jun 8. PMID: 19505943; PMCID: PMC2723002 17 Cingolani P, Platts A, Wang le L, Coon M, Nguyen T, Wang L, Land SJ, Lu X, Ruden DM. A program for annotating and predicting the effects of single nucleotide polymorphisms, SnpEff: SNPs in the genome of Drosophila melanogaster strain w1118; iso-2; iso-3. Fly (Austin). 2012 Apr-Jun;6(2):80-92. doi: 10.4161 / fly.19695. PMID: 22728672; PMCID: PMC3679285 18 Boeva V, Popova T, Bleakley K, et al. Control-FREEC: a tool for assessing copy number and allelic content using next-generation sequencing data. Bioinformatics. 2012;28(3):423-425. doi:10.1093 / bioinformatics / btr670 19 Quinlan AR, Hall IM. BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics. 2010 Mar 15;26(6):841-2. doi: 10.1093 / bioinformatics / btq033. Epub 2010 Jan 28. PMID: 20110278; PMCID: PMC2832824

[0275] All references cited herein and elsewhere herein are hereby incorporated by reference in their entirety.

Claims

1. 1. A method for providing traceability of biological material, comprising: inserting at least one DNA-specific identifier sequence into genomic DNA of a biological entity, wherein the biological entity comprises a plant cell, a fungal cell, an animal cell, a virus, or a bacterial cell; determining the sequence of at least one DNA-specific identifier sequence within the genomic DNA of said biological entity; verifying the identity of the biological entity by verifying the presence of the DNA-specific identifier sequence in the genomic DNA and comparing the sequence of the DNA-specific identifier sequence with a database to confirm that the DNA-specific identifier sequence is not already used in the database; providing an indication of the advisability of producing a biological material from the biological entity, wherein the biological material comprises genomic DNA from the biological entity, and wherein the biological material comprises a food product, a plant-derived material, a fungal-derived material, an animal-derived material, a viral-derived material, or a bacterial-derived material; entering the sequence of at least one of said DNA unique identifier sequences into a database entry in said database, and associating said DNA unique identifier sequence with identification and / or tracking information for said biological material; whereby providing traceability of the biological material by reading the DNA unique identifier sequence in the biological material and retrieving the corresponding database entry or decoding or decrypting information stored in the DNA unique identifier sequence to provide the identification and / or tracking information of the biological material.

2. 10. The method of claim 1, further comprising providing at least one DNA-specific identifier sequence for insertion into the genomic DNA of the biological entity.

3. The method of claim 1 , wherein producing a biological substance from a biological entity comprises growing the biological entity.

4. reading a DNA unique identifier sequence in the biological material and obtaining a corresponding database entry, receiving or providing a sample comprising genomic DNA from said biological material; amplifying at least one DNA-specific identifier sequence in genomic DNA from said biological material and sequencing said DNA-specific identifier sequence; comparing the DNA unique identifier sequence with a database to obtain the database entry corresponding to the DNA unique identifier sequence, the database entry providing identification and / or tracking information for the biological material; The method of claim 1 , comprising:

5. 2. The method of claim 1, wherein the DNA unique identifier sequence comprises a unique nucleotide sequence inserted within an intergenic region of the genomic DNA.

6. 2. The method of claim 1, wherein the DNA-specific identifier sequence comprises a sequence up to 1500 nt in length, up to 1000 nt in length, 200 nt to 600 nt in length, 200 nt to 400 nt in length, or 400 nt to 600 nt in length, and is flanked by one or more primer annealing sequences for PCR amplification of the DNA-specific identifier sequence, sequencing of the DNA-specific identifier sequence, or both.

7. 10. The method of claim 1, wherein the identification and / or tracking information of the database entry includes supply chain information of the biological material, source information of the biological material, grower, region, batch, lot, date, or associated supply chain information, or any combination thereof.

8. 2. The method of claim 1, wherein the DNA unique identifier sequence is a random sequence derived from a randomized pool of nucleic acid sequences up to 1500 nt in length, up to 1000 nt in length, 200 nt to 600 nt in length, 200 nt to 400 nt in length, or 400 nt to 600 nt in length.

9. 10. The method of claim 1, receiving, at a computing device, a DNA unique identifier sequence (DUID) extracted from a known biological material; searching a DUID database, which stores a plurality of DUIDs associated with respective biological material information, in the computing device to match the received DUID; if the search of the DUID database fails to match the received DUID, storing the received DUID in the DUID database in association with biological material information associated with the known biological material; receiving, at the computing device after storing the received DUID and information associated with the known biological material in the DUID database, a query DUID extracted from an unknown biological material; searching the DUID database at the computing device to match the received query DUID; If the search for the DUID results in a match with the received query DUID, returning, in response to the received query DUID, the stored biological information associated with the DUID that matches the query DUID; The method further comprises:

10. searching a DUID database to match the received DUID or the query DUID, wherein the step of searching the DUID database to match the received DUID or the query DUID includes searching the DUID database to match the received DUID or the query DUID; 10. The method of claim 9, further comprising: if an exact match is not found, performing an alignment / identity search on DUIDs stored in the DUID database that closely match the received DUID or the query DUID; or, if the search results in a close match with the query DUID, storing the query DUID in association with the DUID that closely matches the query DUID.

11. Producing a biological material from a biological entity, said biological material comprising genomic DNA from said biological entity; and / or providing an indication of the advisability of producing said biological material from said biological entity, said biological material comprising genomic DNA from said biological entity; The method of claim 1 further comprising:

12. 10. The method of claim 1, receiving or providing a sample from a product of interest, including food, agricultural products, pharmaceuticals, retail products, textiles, goods, chemicals, or another supply chain item, wherein the sample contains genomic DNA from biological material that is part of, mixed with, or otherwise associated with the product of interest; amplifying at least one DNA-specific identifier sequence in genomic DNA from said biological material and sequencing said DNA-specific identifier sequence; searching for the DNA unique identifier sequence in a database and obtaining a database entry corresponding to the DNA unique identifier sequence, the database entry providing identification and / or tracking information for the product of interest, the identification and / or tracking information of the database entry including supply chain information for the product of interest; The method comprising:

13. 10. The method of claim 1, wherein the one or more DNA-specific identifiers are present at one or more harmless sites within the genome of the biological entity.

14. 14. The method of claim 13, wherein the one or more DNA unique identifiers in the one or more innocuous sites do not affect gene expression or phenotype of the biological entity.

15. an oligonucleotide comprising a DNA-specific identifier sequence, wherein the DNA-specific identifier sequence comprises a sequence up to 1500 nt in length, up to 1000 nt in length, 200 nt to 600 nt in length, 200 nt to 400 nt in length, or 400 nt to 600 nt in length, and is flanked by one or more primer annealing sequences for PCR amplification of the DNA-specific identifier sequence, sequencing of the DNA-specific identifier sequence, or both; one or more primer pairs for amplifying and / or sequencing the DNA unique identifier sequence; and buffer solution, polymerase, or Instructions for carrying out the method according to any one of claims 1 to 14 one or more of: A kit for use in the method according to any one of claims 1 to 14, comprising:

16. 16. The kit of claim 15, wherein the DNA-specific identifier sequences comprise a randomized pool of DNA-specific identifier sequences.

17. 16. The kit of claim 15, wherein the DNA-specific identifier sequence is flanked by at least one 5' primer annealing sequence and at least one 3' primer annealing sequence for amplifying the DNA-specific identifier sequence, sequencing the DNA-specific identifier sequence, or both, and wherein the primer annealing sequences do not naturally occur in the genome of the target biological entity.

18. 16. The kit of claim 15, wherein the DNA-specific identifier sequence is flanked by two 5' primer annealing sequences and two 3' primer annealing sequences, the two 5' primer annealing sequences partially overlapping, the two 3' primer annealing sequences partially overlapping, or both, such that the DNA-specific identifier sequence can be amplified by nested PCR.

19. 16. The kit of claim 15, wherein the DNA-specific identifier sequence further comprises a sequencing primer annealing sequence located 5' to the DNA-specific identifier sequence for sequencing the DNA-specific identifier sequence, and / or the sequencing primer annealing sequence is located between two 5' primer annealing sequences.

20. 16. The kit of claim 15, wherein the sequencing primer annealing sequence at least partially overlaps with one or both of the two 5' primer annealing sequences, and / or the two 5' primer annealing sequences partially overlap, with at least a portion of the sequencing primer annealing sequence being located in the overlap.

Citation Information

Patent Citations

  • Cell labelling, tracking and retrieval

    WO2018154027A1