Polymerase x-based genetic physical unclonable functions

By employing TdT to promote random insertions during NHEJ repair and optimizing indel patterns, genetic PUFs achieve efficient, cost-effective, and reliable cell line authentication without barcoding, addressing limitations in existing PUF technologies.

WO2026050313A1PCT designated stage Publication Date: 2026-03-05BOARD OF RGT THE UNIV OF TEXAS SYST
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Existing genetic physical unclonable functions (PUFs) rely on deletions during non-homologous end joining (NHEJ) repair, which limits sequence entropy and necessitates barcoding for unique identifier generation, complicating production and increasing costs.

Method used

Utilizing polymerase X family proteins, particularly Terminal deoxynucleotidyl Transferase (TdT), to favor random insertions during NHEJ repair, eliminating the need for barcoding and enhancing sequence entropy, combined with a post-sequencing feature selection methodology to optimize indel patterns for authentication.

Benefits of technology

Genetic PUFs with increased entropy and reduced production time and cost, achieving robust, unique, and unclonable identifiers for cell line authentication and provenance verification, with improved classification accuracy through logistic regression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025043622_05032026_PF_FP_ABST
    Figure US2025043622_05032026_PF_FP_ABST
Patent Text Reader

Abstract

A Physical Unclonable Function (PUF) is a security primitive that exploits inherent variations in manufacturing protocols to generate unique, random-like identifiers. These identifiers are used for authentication and encryption purposes in hardware security applications in the semiconductor industry. Inspired by the success of silicon PDFs, herein we leverage Terminal deoxynucleotidyl Transferase (TdT), a template-independent polymerase belonging to the X-family of DNA polymerases, to augment the intrinsic entropy generated during DNA lesion repair and rapidly produce genetic PUFs that satisfy the following properties: robustness (i.e., they repeatedly produce the same output), uniqueness (i.e., they do not coincide with any other identically produced PUF), and unclonability (i.e., they are virtually impossible to replicate). Furthermore, we develop a post-sequencing feature selection methodology based on logistic regression to facilitate PUF classification. Our experimental and computational pipeline drastically reduces production time and cost compared to conventional genetic barcoding without compromising the stringent PUF criteria of uniqueness and unclonability. Our results provide novel insights into the function of TdT and represent a major step towards utilization of PUFs as a biosecurity primitive for cell line authentication and provenance attestation.
Need to check novelty before this filing date? Find Prior Art

Description

Polymerase X-based Genetic Physical Unclonable FunctionsRELATED APPLICATION

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 687,257 filed on August 26, 2024, which is incorporated herein by reference in its entirety.STATEMENT REGARDING FEDERALLY FUNDED RESEARCH

[0002] This invention was made with government support under Grant number 2300340, 2029121, and 2114192 awarded by the National Science Foundation and Grant number R41HG012884 awarded by the National Institutes of Health. The government has certain rights in this invention.REFERENCE TO SEQUENCE LISTING

[0003] The present application contains a Sequence Listing which has been submitted in electronic format via EFS-Web and is hereby incorporated by reference in its entirety. The Sequence Listing, created on August 26, 2025, is named UTD-P0007 and is 95 kilobytes in size. The Sequence Listing complies with the requirements of WIPO Standard ST.26.FIELD OF THE INVENTION

[0004] Embodiments relate to the field of genetic engineering and biotechnology, specifically to methods and systems for creating and authenticating genetic physical unclonable functions (PDFs) in cell lines.BACKGROUND

[0005] A physical unclonable function (PUF) (Gao et al., Nature Electronics, 3:81-91, 2020; Gao et al., IEEE Access, 4:61-80, 2016; Herder et al., Proceedings of the IEEE, 102:1126-41, 2014) is a security measure that enables the unique identification and authentication of a device. Silicon PUFs exploit the inherent randomness of semiconductor manufacturing processes and are innately unique and irreproducible. These variations are random but can be exploited to generate a unique “response” to a specific input, known as a “challenge”. The response is based on thephysical characteristics of the chip that are unique to each unit. A database of challenge-response pairs (CRPs) serves as a basis for verification. When the authenticity of the product is to be evaluated, a challenge is presented to compare the product’s response against the reference stored in the database. The standard metrics for evaluating a PUF include robustness (i.e., the ability to produce consistent responses to the same challenges) and uniqueness (i.e., the ability to produce mappings that no other identically manufactured PUF can replicate).

[0006] We previously introduced the first generation of genetic PUF (Li et al., Sci Adv, 8:4106, 2022) by employing amplicon sequencing of a predefined engineered genomic locus as a distinctive signature; a cell line owner can juxtapose these (nucleotide frequency) signatures against a database for verification. Drawing parallels to the utilization of silicon PUFs, this methodology affirms the bona fide procurement of the cell line, concurrently assuring the enduser of both the provenance and the quality of the biological product.

[0007] To develop PUFs, genome engineering using Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) was leveraged (Jinek et al., Elife, 2:2, 2013; Shalem et al., Science, 343:84-7, 2014; Mali et al., Science, 339:823-6, 2013). CRISPR is an immune response mechanism (Makarova et al., Nat Rev Microbiol, 13:722-36, 2015) against bacteriophage infections in bacteria and archaea that has revolutionized the field of genome editing (Li et al., CRISPR J, 1:286-93, 2018; Tabebordbar et al., Science, 351:407-11, 2016; Ran et al., Nature, 520: 186-91, 2015; Qi et al., Cell, 152: 1173-83, 2013; Gilbert et al., Cell, 154:442, 2013; Yang et al., Methods in Molecular Biology), 1114:245-67, 2014). We demonstrated that a two-step process that combines molecular barcoding with non-homologous end joining (NHEJ) repair can exploit the inherent stochasticity of the latter, yielding measurable genetic changes that satisfy all PUF conditions. For the first stage of the protocol, we integrated a 5-nucleotide barcode library into the AAPSI locus of human HEK293 cells, using CRISPR-mediated homologous recombination (HR). We then transiently transfected the barcoded cell line with a sgRNA targeting a locus adjacent to the barcodes to induce NHEJ repair. Finally, we amplified this region via PCR, sequenced the amplicons by NGS, associated the observed indels with their corresponding barcodes from the same reads, and produced the resulting two-dimensional matrix by the frequencies of barcoded (Li et al., Sci Adv, 8:4106, 2022). The two-dimensional matrix of the most frequently detected barcode and indel sequences satisfies the PUF metrics of robustness, uniqueness, and unclonability (Li et al., Sci Adv, 8:4106, 2022).

[0008] There remains a need for additional compositions and methods for making and assessing genetic physical unclonable functions (PUFs).SUMMARY

[0009] The following presents a summary to provide a basic understanding of some aspects of the invention. This summary is not an extensive overview and is not intended to identify critical elements or delineate the scope of the invention. Its purpose is to present key concepts as a prelude to the detailed description that follows.

[0010] The reliability of genetic PUFs hinges on the generation of unpredictable insertions and deletions (indels) following NHEJ repair of the CRISPR-mediated cleavage. However, as the NHEJR indels are dominated by deletions and thus do not generate abundant genetic variability (i.e., entropy) among cell populations, the first generation of genetic PUFs relied on the random association of the indels with genetic barcodes in individual cells (FIG. 1). In this disclosure, the focus is on modulating NHEJ events to prioritize random insertions and subsequently increase sequence entropy at the target locus, potentially removing the need for barcoding (FIG. 1). The error-prone nature of NHEJ repair can involve the recruitment of a group of proteins known as the polymerase-X (Pol X) family of proteins, which are known to catalyze DNA end- and gapfilling processes in template-dependent and template-independent manner. It is contemplated that by manipulating the expression of a Pol X family member, specifically the overexpression of Terminal deoxynucleotidyl Transferase (TdT), one can direct the repair outcome towards frequent random insertions. TdT is known to catalyze the addition of nucleotides to the 3' ends of DNA strands, thus favoring random insertions during NHEJ (Motea et al., Biochrm Biophys Acta, 1804: 1151-66, 2010; Ramsden et al., Environ Mol Mutagen, 53:741-51, 2012; Lutze et al., Methods Mol Biol, 2394:299-317, 2022; Loveless et al., Nat Chem Biol, 17,739-47, 202115- 18). Random TdT-based insertions incorporate one of four possible nucleotides and thus inherently are of higher entropy than deletions, a potentially valuable strategy towards eliminating the need for barcoding in PUF production.

[0011] The genetic signatures of PUFs are obtained through amplicon next-generation sequencing (NGS). Analysis of the first generation PUFs revealed that fewer than 10% of the most frequently observed barcode-indel reads were sufficient to provide an indisputable identification signature (robustness) while being large enough to prevent unauthorizedreproduction (unclonability). Another observation was that specific indels appear across all PUFs (presumably due to the location of editing) and excluding these indels from the PUF signatures improved the ability to distinguish intra-PUFs (same cell line / PUF after it undergoes freeze-thaw cycles) from inter-PUFs (different cell lines / PUFs). Accordingly, instead of using the raw frequency of genetic signatures as PUFs, we developed an algorithm to select specific genetic signatures for the verification of PUFs. More specifically, the algorithm objective function minimizes the differences between duplicates of identical genetic PUFs (i.e. “intra-PUFs”), while maximizing the difference between different genetic PUFs (i.e. “inter-PUFs”).

[0012] Any embodiment described herein may apply to other aspects of the invention unless explicitly stated otherwise. Methods, compositions, systems, or kits disclosed herein may be used interchangeably to achieve the objectives of the invention.

[0013] As used herein, the term “about” indicates a value includes the inherent variation of error for the device, method, or measurement used to determine the value.

[0014] The term “or” in the claims means “and / or” unless explicitly stated to refer to mutually exclusive alternatives.

[0015] Terms such as “comprising,” “including,” “having,” or “containing” are open-ended and do not exclude additional, unrecited elements or steps. For example, a method or composition that “comprises” listed elements may include other elements not expressly recited but inherent to the method or composition.

[0016] The phrases “consists of’ or “consisting of’ exclude any element, step, or component not specified in the claim, except for impurities ordinarily associated with the recited components. When “consists of’ or “consisting of’ appears in a claim clause, it limits only the elements in that clause; other elements are not excluded from the claim as a whole.

[0017] The phrases “consists essentially of’ or “consisting essentially of’ include the recited elements, steps, or components and any additional elements, steps, or components that do not materially affect the basic and novel characteristics of the invention.

[0018] Negative limitations may be used in the claims to exclude specific elements, features, or embodiments to define the invention’s scope more precisely. These exclusions distinguish the invention from prior art or alternative configurations and are supported by the specification, which provides context for the recited exclusions.

[0019] Other objects, features, and advantages of the invention will become apparent from the following detailed description and drawings. The detailed description and examples, while indicating specific embodiments, are provided for illustration only, and various modifications within the spirit and scope of the invention will be apparent to those skilled in the art.Definitions:

[0020] As used herein, the following terms shall have the meanings set forth below, unless otherwise indicated by the context. These definitions are provided to clarify the scope of the invention and to aid in understanding the detailed description and claims. Terms not explicitly defined herein shall be given their ordinary meaning as understood by one of ordinary skill in the art.

[0021] Authentication: Authentication refers to the process of confirming that a cell line in question was derived from a cell line that was previously modified with a genetic physical unclonable function (PUF) by a prior holder. This confirmation is typically achieved by comparing the indel frequency distribution of the cell line against an established authentication classifier.

[0022] Authentication Classifier: An authentication classifier is a dataset comprising a selected list of indels and their corresponding frequency distributions, derived from replicate sequencing of a cell line containing a genetic PUF, along with an associated threshold dissimilarity value. The classifier is used to verify the authenticity and provenance of a cell line by measuring statistical dissimilarity against an unknown sample's indel profile.

[0023] Bray-Curtis Dissimilarity (BCD): Bray-Curtis Dissimilarity is a statistical metric used to quantify the difference in composition between two frequency distributions, such as indel frequencies in different cell line samples. Lower BCD values indicate greater similarity (e.g., intra-PUF comparisons), while higher values indicate dissimilarity (e.g., inter-PUF comparisons).

[0024] Cell Line: A cell line refers to a population of cells derived from a single cell or a group of cells that can be propagated indefinitely in culture under controlled conditions. Examples include human cell lines such as HEK293 or HCT116, and generally mammalian cells such as CHO which may be modified to incorporate genetic PUFs.

[0025] Challenge-Response Pair (CRP): A Challenge-Response Pair is a security mechanism in which a specific input (challenge), such as a genomic locus targeted for sequencing, elicits aunique output (response), such as an indel frequency distribution. In the context of genetic PUFs, CRPs are used to verify authenticity by comparing the response against a stored database.

[0026] CRISPR-Cas Endonuclease: CRISPR-Cas endonuclease refers to a programmable nuclease system, such as SpCas9, used for targeted genome editing. It comprises a Cas protein and a guide RNA (gRNA) that directs the nuclease to cleave DNA at a specific locus, inducing double-strand breaks that trigger repair mechanisms like NHEJ.

[0027] Double-Strand DNA Lesion (dsDNA Lesion): A double-strand DNA lesion is a break in both strands of the DNA double helix, induced by endonucleases such as CRISPR-Cas, TALENs, or zinc finger nucleases. These lesions are repaired through mechanisms like NHEJ, leading to the formation of indels.

[0028] Endonuclease: An endonuclease is an enzyme that cleaves DNA at specific internal sites. In this invention, endonucleases include CRISPR-Cas systems, transcription activator-like effector nucleases (TALENs), and zinc finger nucleases, used to introduce targeted dsDNA lesions.

[0029] Physical Unclonable Function (PUF): A physical unclonable function is a security measure that exploits inherent manufacturing randomness to create unique, irreproducible identifiers for authentication. In silicon-based PUFs, this randomness arises from semiconductor variations; in genetic PUFs, it arises from stochastic DNA repair.

[0030] Genetic Physical Unclonable Function (Genetic PUF): A genetic physical unclonable function is a biological security primitive created in a cell line by exploiting stochastic DNA repair processes to generate a unique set of random indels at a targeted genomic locus. It satisfies PUF criteria of robustness, uniqueness, and unclonability, serving as an identifier for cell line authentication and provenance verification.

[0031] Indel (Insertion or Deletion): An indel refers to a genetic alteration involving the insertion or deletion of nucleotides at a DNA repair site, typically resulting from NHEJ repair of dsDNA breaks. Indels in this invention are predominantly insertions when modulated by Pol X proteins like TdT, contributing to sequence entropy.

[0032] Inter-PUF: Inter-PUF refers to comparisons or dissimilarities between different genetic PUFs (i.e., between distinct cell lines or the same cell line carrying different genetic PUFs), which should exhibit high statistical dissimilarity to ensure uniqueness.

[0033] Intra-PUF: Intra-PUF refers to comparisons or dissimilarities within the same genetic PUF (i.e., between replicate samples or passages of the same cell line), which should exhibit low statistical dissimilarity to ensure robustness.

[0034] Next-Generation Sequencing (NGS): Next-generation sequencing is a high- throughput DNA sequencing technology, such as Illumina MiSeq, used to characterize indel populations at targeted loci by generating amplicon sequences and frequency data.

[0035] Non-Homologous End Joining (NHEJ): Non-homologous end joining is an error- prone DNA repair pathway that ligates broken DNA ends without requiring a homologous template, often resulting in indels. In this invention, NHEJ is modulated to favor insertions through Pol X protein activity.

[0036] Nucleic Acid-Repair Protein: A nucleic acid-repair protein is an enzyme involved in DNA repair processes. In this context, it specifically refers to members of the Pol X family that facilitate noisy repair during NHEJ.

[0037] Polymerase-X (Pol X) Family: The polymerase-X family is a group of specialized DNA polymerases involved in error-prone DNA synthesis during repair processes like NHEJ. Members include Terminal deoxynucleotidyl Transferase (TdT), which catalyzes templateindependent nucleotide additions.

[0038] Provenance: Provenance refers to the documented origin and history of a cell line. In this specification, determining provenance involves verifying whether a cell line originated from a source that incorporated a genetic PUF, requiring authentication as an essential step.

[0039] Robustness: Robustness is a PUF metric referring to the ability of a genetic PUF to produce consistent indel frequency responses to the same challenge (e.g., sequencing replicates) across conditions like freeze-thaw cycles or passages.

[0040] Safe-Harbor Locus: A safe-harbor locus is a genomic site suitable for targeted modifications without disrupting essential cellular functions, such as AAVS1, hRosa26, or CCR5 in human cells.

[0041] Sequence Entropy: Sequence entropy refers to the measure of genetic variability or randomness introduced at a repair site, increased by random insertions (each potentially adding one of four nucleotides) compared to deletions.

[0042] Set- Sei ection: Set-selection is a post-sequencing heuristic methodology that iteratively selects a subset of indels to minimize intra-PUF dissimilarity while maximizing inter- PUF dissimilarity, optimizing authentication classifiers.

[0043] Terminal Deoxynucleotidyl Transferase (TdT): Terminal deoxynucleotidyl Transferase is a Pol X family protein that adds random nucleotides to the 3' ends of DNA strands in a template-independent manner, favoring insertions during NHEJ repair and enhancing indel entropy.

[0044] Unclonability: Unclonability is a PUF metric referring to the practical impossibility of replicating a genetic PUF due to the stochastic nature of indel generation, requiring an impractically large number of attempts (e.g., engineering 380 or more cell lines) to reproduce.

[0045] Uniqueness: Uniqueness is a PUF metric referring to the ability of a genetic PUF to produce an indel profile distinct from other PUFs generated by the same method, ensuring no two PUFs can be confused.

[0046] SEQ ID NO: 1 - Reference sequence for the mKate2 gene integrated into a cell line, used as a target for indel generation in genetic physical unclonable function (PUF) experiments.

[0047] SEQ ID NO:2 - Indel sequence generated in the mKate2 gene by Cas9-mediated non- homologous end joining (NHEJ) repair in a cell line for genetic PUF creation.

[0048] SEQ ID NO:3 - Indel sequence generated in the mKate2 gene by Cas9-mediated NHEJ repair in a cell line for genetic PUF creation.

[0049] SEQ ID NO:4 - Indel sequence generated in the mKate2 gene by Cas9-mediated NHEJ repair in a cell line for genetic PUF creation.

[0050] SEQ ID NO:5 - Indel sequence generated in the mKate2 gene by Cas9-mediated NHEJ repair in a cell line for genetic PUF creation.

[0051] SEQ ID NO:6 - Description: Indel sequence generated in the mKate2 gene by Cas9- mediated NHEJ repair in a cell line for genetic PUF creation.

[0052] SEQ ID NO:7 - Indel sequence generated in the mKate2 gene by Cas9-mediated NHEJ repair in a cell line for genetic PUF creation.

[0053] SEQ ID NO:8 - Indel sequence generated in the mKate2 gene by Cas9-mediated NHEJ repair in a cell line for genetic PUF creation.

[0054] SEQ ID NO:9 - Indel sequence generated in the mKate2 gene by Cas9 and Terminal deoxynucleotidyl Transferase (TdT)-mediated NHEJ repair in a cell line for genetic PUF creation.

[0055] SEQ ID NO: 10 - Indel sequence generated in the mKate2 gene by Cas9 and TdT- mediated NHEJ repair in a cell line for genetic PUF creation.

[0056] SEQ ID NO: 11 - Indel sequence generated in the mKate2 gene by Cas9 and TdT- mediated NHEJ repair in a cell line for genetic PUF creation.

[0057] SEQ ID NO: 12 - Indel sequence generated in the mKate2 gene by Cas9 and TdT- mediated NHEJ repair in a cell line for genetic PUF creation.

[0058] SEQ ID NO: 13 - Indel sequence generated in the mKate2 gene by Cas9 and TdT- mediated NHEJ repair in a cell line for genetic PUF creation.

[0059] SEQ ID NO: 14 - Indel sequence generated in the mKate2 gene by Cas9 and TdT- mediated NHEJ repair in a cell line for genetic PUF creation.

[0060] SEQ ID NO: 15 - Indel sequence generated in the mKate2 gene by Cas9 and TdT- mediated NHEJ repair in a cell line for genetic PUF creation.

[0061] SEQ ID NO: 16 - Plasmid sequence (pCMV-SpCas9-U6-sgRNAmKate2) containing the pCMV promoter, SpCas9 open reading frame, U6 promoter, spacer targeting mKate2, and scaffold for SpCas9, used for CRISPR-mediated indel generation in genetic PUF experiments.

[0062] SEQ ID NO: 17 - Plasmid sequence (pCMV-TdT) containing the pCMV promoter and Terminal deoxynucleotidyl Transferase (TdT) open reading frame, used to enhance insertion- biased NHEJ repair in genetic PUF experiments.

[0063] SEQ ID NO: 18 - Plasmid sequence (pCMV-SpCas9-T2A-mKate2-U6- sgRNA_AAVSl) containing the pCMV promoter, SpCas9 open reading frame, T2A peptide, mKate2 open reading frame, U6 promoter, spacer targeting AAVS1, and scaffold for SpCas9, used for CRISPR-mediated indel generation at the AAVS1 locus in genetic PUF experiments.

[0064] SEQ ID NO: 19 - Plasmid sequence (pCMV-SpCas9-T2A-mKate2-U6- sgRNA_CCR5) containing the pCMV promoter, SpCas9 open reading frame, T2A peptide, mKate2 open reading frame, U6 promoter, spacer targeting CCR5, and scaffold for SpCas9, used for CRISPR-mediated indel generation at the CCR5 locus in genetic PUF experiments.

[0065] SEQ ID NO:20 - Plasmid sequence (pCMV-SpCas9-T2A-mKate2-U6- sgRNA_Rosa26) containing the pCMV promoter, SpCas9 open reading frame, T2A peptide,mKate2 open reading frame, U6 promoter, spacer targeting Rosa26, and scaffold for SpCas9, used for CRISPR-mediated indel generation at the Rosa26 locus in genetic PUF experiments.

[0066] SEQ ID NO:21 - Plasmid sequence (pCMV-SpCas9-T2A-mKate2-U6- sgRNA_Rosa26 (CHO)) containing the pCMV promoter, SpCas9 open reading frame, T2A peptide, mKate2 open reading frame, U6 promoter, spacer targeting Rosa26 in Chinese hamster ovary (CHO) cells, and scaffold for SpCas9, used for CRISPR-mediated indel generation in genetic PUF experiments.

[0067] SEQ ID NO:22 - Plasmid sequence (pCMV-TdT-T2A-EGFP) containing the pCMV promoter, TdT open reading frame, T2A peptide, and EGFP open reading frame, used to enhance insertion-biased NHEJ repair and mark transfected cells in genetic PUF experiments.

[0068] SEQ ID NO:23 - Forward primer (Pl) sequence for next-generation sequencing (NGS) of the integrated mKate2 region in genetic PUF experiments, as referenced in Figure 2.

[0069] SEQ ID NO:24 - Reverse primer (P2) sequence for NGS of the integrated mKate2 region in genetic PUF experiments.

[0070] SEQ ID NO:25 - Forward primer (P3) sequence, including Illumina partial adapter, for NGS of the AAVS1 amplicon in genetic PUF experiments.

[0071] SEQ ID NO:26 - Reverse primer (P4) sequence, including Illumina partial adapter, for NGS of the AAVS1 amplicon in genetic PUF experiments.

[0072] SEQ ID NO:27 - Forward primer (P5) sequence, including Illumina partial adapter, for NGS of the hRosa26 amplicon in genetic PUF experiments.

[0073] SEQ ID NO:28 - Reverse primer (P6) sequence, including Illumina partial adapter, for NGS of the hRosa26 amplicon in genetic PUF experiments.

[0074] SEQ ID NO:29 - Forward primer (P7) sequence, including Illumina partial adapter, for NGS of the CCR5 amplicon in genetic PUF experiments.

[0075] SEQ ID NO:30 - Reverse primer (P8) sequence, including Illumina partial adapter, for NGS of the CCR5 amplicon in genetic PUF experiments.

[0076] SEQ ID NO:31 - Forward primer (P9) sequence, including Illumina partial adapter, for NGS of the Rosa26 amplicon in Chinese hamster ovary (CHO) cells in genetic PUF experiments.

[0077] SEQ ID NO:32 - Reverse primer (P10) sequence, including Illumina partial adapter, for NGS of the Rosa26 amplicon in Chinese hamster ovary (CHO) cells in genetic PUF experiments.

[0078] SEQ ID NO:33 - Reference amplicon sequence for the integrated mKate2 region, including the target sequence for CRISPR-mediated cleavage, used in genetic PUF experiments.

[0079] SEQ ID NO:34 - Reference amplicon sequence for the AAVS1 locus, including the target sequence for CRISPR-mediated cleavage, used in genetic PUF experiments.

[0080] SEQ ID NO:35 - Reference amplicon sequence for the hRosa26 locus, including the target sequence for CRISPR-mediated cleavage, used in genetic PUF experiments.

[0081] SEQ ID NO:36 - Reference amplicon sequence for the CCR5 locus, including the target sequence for CRISPR-mediated cleavage, used in genetic PUF experiments.

[0082] SEQ ID NO:37 - Reference amplicon sequence for the Rosa26 locus in Chinese hamster ovary (CHO) cells, including the target sequence for CRISPR-mediated cleavage, used in genetic PUF experiments.DESCRIPTION OF THE DRAWINGS

[0083] The following drawings form part of the specification and are included to illustrate certain aspects of the invention. The invention may be better understood by reference to these drawings in combination with the detailed description.

[0084] FIG. 1. PUF engineering methodology. Overview of barcode-indel PUFs, Cas9- induced NHEI indel generation and frequency-based selection for signature verification. Compared against TdT PUF consisting of a single Cas9-induced NHEJ indel generation and Pol X enzyme (TdT) for entropy increase, alongside optimized feature selection techniques for signature verification.

[0085] FIG. 2A-2H. Characterization of TdT-mediated indels. (a) Illustration of the hypothesized mechanisms involved in TdT-induced insertion-biased NHEI repair, (b) Plasmid constructs used to target the integrated mKate2. (c) Tables of the seven most frequent indels of Cas9 alone and Cas9 with TdT, generated by CRISPResso2. SEQ ID NO: 1 = reference, SEQ ID NO:2-8 = Cas9, SEQ ID NO:9-15 = Cas9 +TdT. (d) Bar graph (top) comparing indel size distributions between Cas9 alone and Cas9 alongside TdT. Bar graph (bottom) showing the expected contribution to diversity, (e) Comparison between numbers of possible indels andobserved indels in Cas9 alone and Cas9 with TdT samples, (f) Plasmid constructs used to target the AAVS1 site, (g, h) Frequency plots of the top 20 indels at the AA S7 site with Cas9 (g) and with Cas9-TdT (h). Characterization of TdT-Mediated Insertions.

[0086] FIG. 3A-3D. The relationship between TdT-mediated indels and transfection time points and location, (a) Illustration of the experimental timeline. The experiment was conducted at three time points: TdT-Tl, TdT-T2, and TdT-T3. For each time point, cells were seeded in a six-well plate on the first day, transfections were performed on the second day, and cells were enriched via FACS on the fifth day (3 days post-transfection), after which gDNAs were harvested. Cells were transfected at 3 time points (TdT-Tw), with 3 PUFs at each time point, (b) Frequency plot of the top 20 indels at different transfection time points, (c) Plasmid constructs used to target the Rosa26 and the CCR5 loci, (d) Frequency plots of the top 20 indels at the hRosa26 and CCR5 sites.

[0087] FIG. 4A-4D. Exclusion of common indels for enhanced PUF classification, (a) Overview of the computational processes, including exclusion of singular indel reads, removal of common indels, and use of the top 25% most frequent indels to perform pairwise BCD calculation, (b) Schematic of the generation of common indels among the samples, (c) Frequency plots of the top 20 indels from samples seq-TdTi-6 after common indel exclusion, (d) BCD between seq-TdTi and the rest of the samples at the AAVS1 locus after common indel removal. The fold changes shown are ratios between intra-PUF BCD (gray) and the average of inter-PUFs BCD (black).

[0088] FIG. 5A-5F. Multiclass logistic regression for classification of intra- and inter- PUFs. (a) The indel sequences and frequencies from the original samples (TdTi, TdT2...TdT / 2) were used as training data to build a logistic regression model. The technical replicates of these samples (e.g. TdTir, TdT2r...TdTi2r) were used as test data, and the trained model predicted the classes to which they belonged, (b) Classification results of each test sample against every other sample as predicted by the logistic regression model trained in (a), (c) Plasmid constructs used to target various loci in human cell lines (HEK293, HCT116, A549, HeLa) and Chinese Hamster Ovary cell line. The resulting PUFs created were used to train the multiclass logistic regression algorithm as described in (a), (d) Classification results of each TdT-PUF sample against every other sample as predicted by the logistic regression model trained in (c). (e) To generate simulated PUFs, the sequences and corresponding frequencies of indels in every PUF wereconcatenated into a single comprehensive distribution of all observed PUFs. Simulated PUFs were then generated by randomly subsampling from this distribution. Technical replicates of these simulated PUFs were then created using random subsampling to reproduce the similarity levels measured experimentally, (f) Classification accuracy of simulated PUFs generated by concatenating all indel s.

[0089] FIG. 6. Poly-X based PUF attributes. Comparison of Barcode- Indel PUFs and TdT PUFs. The TdT PUFs significantly reduced the production time and cost associated with the complexity of implementing PUFs, while maintaining their robustness, uniqueness, and unclonability.DESCRIPTION

[0090] The following discussion is directed to various embodiments of the invention. The term “invention” is not intended to refer to any particular embodiment or otherwise limit the scope of the disclosure. Although one or more of these embodiments may be preferred, the embodiments disclosed should not be interpreted, or otherwise used, as limiting the scope of the disclosure, including the claims. In addition, one skilled in the art will understand that the following description has broad application, and the discussion of any embodiment is meant only to be an example of an embodiment(s) and not intended to imply that the scope of the disclosure, including the claims, is limited to that embodiment.

[0091] The present invention provides methods and systems for generating and authenticating genetic physical unclonable functions (PUFs) in cell lines using polymerase X (Pol X) family proteins, particularly Terminal deoxynucleotidyl Transferase (TdT), to produce unique genetic signatures through non-homologous end joining (NHEJ) repair. These genetic PUFs serve as robust, unique, and unclonable identifiers for cell line authentication and provenance verification, eliminating the need for prior barcoding techniques used in earlier genetic PUF methodologies. The invention further includes a post-sequencing feature selection methodology to enhance the identification of indel patterns that meet PUF criteria, thereby improving efficiency and reducing production costs.I. Overview of Genetic PUFs

[0092] A physical unclonable function (PUF) is a security primitive that leverages inherent randomness in a process to generate unique identifiers for authentication purposes. In the context of this invention, genetic PUFs are created in cell lines by exploiting the stochastic nature of DNA repair mechanisms, specifically NHEJ, to produce random insertions and deletions (indels) at targeted genomic loci. These indels form a unique genetic signature for each cell line, satisfying three key PUF criteria: robustness (consistent output under identical conditions), uniqueness (distinct from other PUFs produced by the same method), and unclonability (resistance to unauthorized replication).

[0093] Unlike prior genetic PUF approaches that relied on random associations between engineered barcodes and indels, the present invention utilizes the error-prone repair activity of TdT, a Pol X family protein, to favor random insertions over deletions during NHEJ repair. This shift increases sequence entropy at the target locus, enabling the creation of genetic PUFs without the need for barcoding. The invention further employs a novel post-sequencing feature selection methodology based on logistic regression to identify optimal indel combinations that facilitate PUF classification.II. Methods for Creating Genetic PUFs

[0094] In one embodiment, the method for creating a genetic PUF in a cell line comprises two primary steps: (1) introducing double-strand DNA (dsDNA) lesions at a targeted genomic locus using an endonuclease, and (2) modulating DNA repair through NHEJ using a nucleic acid-repair protein, specifically a Pol X family member such as TdT. The targeted locus is preferably a safe-harbor site in the genome, such as AAVS1, hRosa26, or CCR5, to minimize unintended effects on cell function.

[0095] The endonuclease used to induce dsDNA lesions may include, but is not limited to, a CRISPR-Cas endonuclease (e.g., SpCas9), a transcription activator-like effector nuclease (TAEEN), or a zinc finger nuclease. In a preferred embodiment, a CRISPR-Cas system is employed, comprising a Cas protein (e.g., SpCas9) and a guide RNA (gRNA) designed to target a specific genomic sequence. The gRNA directs the Cas protein to cleave the DNA at a predetermined site, creating a dsDNA break that triggers NHEJ repair.

[0096] The NHEJ repair process is modulated by expression of TdT, which catalyzes the addition of random nucleotides to the 3' ends of the DNA strands at the break site. Unlike standard NHEJ repair, which predominantly results in deletions, TdT-mediated repair favors insertions, typically ranging from one to four base pairs in length. These insertions introduce significant sequence entropy, as each insertion event can incorporate any of the four nucleotides (A, T, C, or G), thereby increasing the diversity of the resulting indel population.

[0097] In one example, plasmids encoding SpCas9, a gRNA targeting a safe-harbor locus (e.g., AAVS1, hRosa26, or CCR5), and TdT are co-transfected into a human cell line, such as HEK293, HCT116, A549, HeLa cells, or other mammalian cell line, such as Chinese hamster ovary (CHO) cells. The cells are cultured under standard conditions, and after a period (e.g., 72 hours), genomic DNA is extracted from transfected cells. The targeted locus is amplified via polymerase chain reaction (PCR) using primers specific to the region surrounding the cleavage site. The amplified DNA is subjected to next-generation sequencing (NGS), such as Illumina MiSeq paired-end sequencing, to characterize the indel population.

[0098] The resulting indel profiles demonstrate a shift toward insertion-dominated repair when TdT is expressed, with insertions accounting for over 80% of genetic alterations compared to approximately 84% deletions in the absence of TdT. The increased entropy from insertion events eliminates the need for barcoding, simplifying the PUF generation process, reducing costs, and reducing production time from approximately three months to two weeks.III. Methods for Assessing PUFsA. PUF classification by removing common indels

[0099] To authenticate cell lines using genetic PUFs, the invention includes a set-selection method for creating authentication classifiers based on the frequency distributions of indels. This method employs a post-sequencing processing method which starts by omitting sequences that only appear a single time in each sample. Subsequently, we designate a set (list without duplicates) of indel sequences from a specific sample (e.g. seq-TdTi), as the reference set and intersect it with every other sample (e.g. seq-TdT2-e) and their technical replicates (e.g. seq- TdT2r-6r). The intersection represents all common indels across samples, which we exclude from subsequent analysis. Lastly, the top 25% most abundant indels are used to compute pairwise BCD

[0100] In one embodiment, the set-selection process involves the following steps:

[0101] Sequencing Replicates - Multiple replicate samples of a cell line are sequenced using NGS to determine the indel population at the targeted locus. The sequencing data are processed to fdter out low-quality reads and indels appearing only once in each sample, ensuring only reliable indel sequences are considered.

[0102] Frequency Distribution Analysis - The frequency distribution of indels is calculated for each replicate sample. This distribution represents the relative abundance of each unique indel sequence in the population.

[0103] Common Indel Removal - A combined list of the most common indels across all replicates of all cell lines is generated, ranked by frequency. A predetermined percentage (e.g., 1-50%, preferably 25%) of the most frequent indels is removed from each replicate’s indel list to create revised lists. This step eliminates indels that are common across different PDFs, enhancing uniqueness.

[0104] Statistical Dissimilarity Measurement - The Bray-Curtis dissimilarity (BCD) metric is used to measure the statistical dissimilarity between the revised frequency distributions of indels for each replicate. The BCD is calculated as follows:where the frequencies of indel (k) in samples (i) and (j), respectively; and n is the frequency of index k in each sample. This metric quantifies the difference in indel composition between samples.

[0105] Threshold Dissimilarity Designation - Based on the BCD calculations, a threshold dissimilarity value is established for each authentication classifier. This threshold determines the maximum allowable dissimilarity for a sample to be considered authentic. The methodology leverages standard next-generation sequencing (NGS) tools, such as Illumina MiSeq, and computational libraries routine in the art (e.g., Python's scipy. spatial. distance.braycurtis for BCD calculations), enabling a person of ordinary skill in biotechnology and bioinformatics to implement it without undue experimentation.

[0106] The set-selection methodology enhances the ability to distinguish between cell lines. For example, experimental results demonstrate that the fold change between intra-PUF and inter-PUF dissimilarities increases from approximately 2.338 to 3.217 for PUFs generated at the AAVS1 locus, and from 2.662 to 4.349 and 3.568 to 6.161 for PUFs at the hRosa26 and CCR5 loci, respectively, after applying set-selection.B. Classification of PUFs via logistic regression on TdT-mediated indels

[0107] To enable robust authentication of cell lines across large libraries of genetic physical unclonable functions (PUFs), the invention includes a method for classification of PUFs based on logistic regression applied to TdT-mediated indel profiles. Unlike heuristic feature filtering, such as common-indel removal, this approach leverages data-driven feature selection to identify subsets of indels most informative for distinguishing between unique PUFs, thereby improving scalability and classification accuracy.

[0108] In one embodiment, the method comprises the following steps:

[0109] Sequencing and Feature Extraction - Multiple replicate samples of PUF-engineered cell lines are sequenced to generate profiles of indel sequences and their respective frequencies at the targeted genomic loci. Indel distributions are compiled into feature vectors, where each unique indel corresponds to a model feature.

[0110] Multiclass Logistic Regression with Feature Selection - A multiclass logistic regression model is trained using the indel frequency vectors from replicate samples of multiple PUFs. An LI regularization penalty is applied to eliminate features with low discriminative power, thereby isolating the subset of indels most informative for classification. After model training on 12 experimentally derived PUFs, classification accuracy for intra-PUF replicates (e.g., seq-TdTl vs. seq-TdTlr) ranged from 96.2% to 99.9%, demonstrating that the model effectively distinguishes between intra- and inter-PUFs.

[0111] Scalability Across Cellular Contexts - To assess scalability, TdT-based PUFs were constructed in additional human cell lines (HeLa, HCT116, A549) and Chinese hamster ovary (CHO) cells using the same experimental pipeline. The AAVS1 locus was targeted in all human cell lines, while Rosa26 was used for CHO cells. Pairwise Bray-Curtis dissimilarity (BCD) analysis revealed significant separation between intra- and inter-PUFs, with average fold differences of 4.006 for HeLa, 5.039 for HCT116, 5.627 for A549, and 1.661 for CHO cells. Retraining the logistic regression model on the expanded dataset using an L2 penalty preservedclassification accuracy above 99% for all samples, demonstrating the scalability of the supervised learning approach to diverse cellular contexts.

[0112] Simulation-Based Performance Evaluation - To evaluate classifier performance on larger libraries, a synthetic dataset was generated by concatenating indel sequences and frequencies from all experimental data to construct a comprehensive probability distribution. Simulated PUFs were generated by random sampling with replacement from this distribution, and replicate samples were produced with 77% overlap in composition, reflecting experimentally observed averages. Logistic regression models trained on simulated libraries containing 25 to 300 PUFs maintained prediction accuracies above 95.2% across all tested scenarios.

[0113] Entropy Constraints and Classification Limits - To evaluate the effect of repair entropy on classification accuracy, simulated PUF libraries were generated under controlled constraints on insertion sizes. Since TdT-mediated insertions are the primary source of repair entropy, insertion lengths were limited to 3-7 base pairs in one simulation set. Under these constraints, a significant proportion of samples exhibited reduced classification performance, with accuracies falling below 95%, confirming that insertion size diversity directly impacts classification reliability.

[0114] Implementation and Utility - These results demonstrate that TdT-mediated indel distributions provide sufficient repair entropy to enable reliable classification of at least 100 distinct PUFs produced within a single batch while maintaining prediction accuracies above 95%. The methodology leverages standard next-generation sequencing (NGS) platforms, such as Illumina MiSeq, and readily available computational libraries (e.g., Python’s scikit-learn for logistic regression and scipy. spatial. distance. braycurtis for BCD calculations), enabling straightforward implementation by a person of ordinary skill in biotechnology and bioinformatics.C. Authentication of Cell Lines

[0115] The invention further provides a method for authenticating an unknown cell line using the generated genetic PUFs and authentication classifiers. The authentication process comprises the following steps:

[0116] Sequencing the Unknown Cell Line - Replicate samples of the unknown cell line are sequenced using NGS to characterize the indel population at the targeted locus.

[0117] Creating Indel Lists - A first list of indels and their corresponding frequency distributions is generated for each replicate sample. A second list is created by removing indels that are not part of the authentication classifier for the cell line being tested.

[0118] Dissimilarity Measurement - The BCD metric is used to measure the statistical dissimilarity between the frequency distributions of the second indel list and the authentication classifier.

[0119] Authentication Decision - The cell line is designated as authentic if the measured dissimilarity is below the threshold dissimilarity value associated with the authentication classifier.

[0120] In a preferred embodiment, the sequencing is performed using amplicon-based NGS techniques, and the BCD metric is employed for dissimilarity measurements due to its robustness in comparing frequency distributions. The authentication process ensures that only cell lines with indel profiles matching the authentication classifier are verified, confirming their provenance and authenticity.D. System for Authentication

[0121] The invention also encompasses a system for authenticating cell lines, comprising one or more of the following:

[0122] A Modified Cell Line - A cell line containing a genetic PUF, characterized by a unique set of random indels generated through TdT-mediated NHEJ repair at a targeted locus.

[0123] An Authentication Classifier - A dataset comprising the list of indels in the PUF and a corresponding threshold dissimilarity value, derived through the set-selection methodology.

[0124] Sequence Data - NGS data from replicate samples of an unknown cell line.

[0125] A Computer System - A computing device configured to perform the authentication steps, including creating indel lists, filtering indels based on the authentication classifier, measuring statistical dissimilarity using the BCD metric, and designating the cell line as authentic or not based on the threshold value.

[0126] The computer system may include software for processing NGS data, calculating frequency distributions, and performing BCD calculations. The system is designed to integrate seamlessly with standard laboratory workflows for DNA extraction, PCR amplification, and NGS sequencing.

[0127] Experimental results demonstrate the efficacy of TdT-mediated genetic PUFs. In experiments conducted with human cell lines (HEK293, HCT116, A549, HeLa) and Chinese Hamster Ovary cell line, co-transfection of plasmids encoding SpCas9, a gRNA targeting the AAVS1, hRosa26, or CCR5 locus, and TdT resulted in a significant shift toward insertion- dominated indel profiles. For example, at the AAVS1 locus, insertions accounted for over 80% of indels, with a preference for cytosine (C) and guanine (G) nucleotides, influenced by the cleavage site’s nucleotide composition (e.g., GC). Similar insertion biases were observed at the hRosa26 (AT) and CCR5 (CT) loci, with adenine (A) and thymine (T) insertions being more prevalent, respectively.

[0128] The consistency of indel patterns across replicates and the influence of cleavage site composition on insertion bias were confirmed through multiple independent transfections. The BCD metric revealed an average fold change of 3.408 between intra-PUF and inter-PUF dissimilarities at the AAVS1 locus, indicating robust separation between cell lines. The setselection methodology further improved this separation, enhancing the reliability of authentication.

[0129] Additional experiments showed that indel distributions were not significantly affected by the timing of transfection, suggesting that cleavage site composition and TdT activity are the primary determinants of indel entropy. These findings support the robustness and reproducibility of the TdT-based PUF generation process.

[0130] The present invention offers several advantages over prior genetic PUF methodologies:

[0131] Elimination of Barcoding - By leveraging TdT-mediated insertions, the invention removes the need for labor-intensive and costly barcoding steps, reducing the production time from a few months to days. This change also decreases the total cost of engineering each PUF cell line by 3 -fold, including reagents, consumables, and NGS service.

[0132] Increased Entropy - TdT-mediated insertions introduce greater sequence diversity compared to deletion-dominated NHEJ repair, enhancing the uniqueness and unclonability of PUFs.

[0133] Streamlined Authentication - The set-selection methodology optimizes the identification of indel signatures, improving the efficiency and accuracy of cell line authentication.

[0134] Commercial Viability - The reduced time and cost, combined with maintained PUF criteria (robustness, uniqueness, unclonability), make the invention a practical solution for commercial applications in cell line authentication and provenance verification.

[0135] The invention is particularly suited for applications in biotechnology, where ensuring the authenticity and provenance of cell lines is critical for research, therapeutic development, and biomanufacturing. The methods and systems described herein provide a scalable and cost- effective approach to generating secure genetic identifiers.E. Variations and Modifications

[0136] Those skilled in the art will recognize that various modifications can be made to the described embodiments without departing from the scope of the invention. For example, alternative endonucleases, such as Cas variants with broadened protospacer adjacent motif (PAM) compatibility, may be used to target additional genomic loci. Similarly, other Pol X family proteins or modified versions of TdT may be employed to further modulate indel patterns. The culture conditions, such as supplementation with deoxyribonucleosides, may be adjusted to manipulate indel entropy. Additionally, alternative statistical dissimilarity metrics or machine learning algorithms may be used in place of the BCD metric or set-selection methodology or the logistic regression to optimize authentication classifiers.

[0137] The invention is not limited to the specific cell lines (e.g., HEK293, HCT116) or loci (e.g., AAVS1, hRosa26, CCR5) described herein. Other cell types, including primary cells, stem cells, or non-human cell lines, and other genomic loci may be used, provided they are suitable for targeted DNA cleavage and repair. The sequencing platform, PCR conditions, and data analysis pipelines may also be adapted to suit specific laboratory capabilities.

[0138] The embodiments described herein are illustrative of the principles of the invention. Other embodiments, including combinations of the described methods and systems, are contemplated and fall within the scope of the invention as defined by the claims.IV. Examples

[0139] The following examples illustrate preferred embodiments of the invention. The techniques disclosed in these examples represent approaches the inventors have found effectivein practicing the invention. Those skilled in the art will appreciate that modifications to these embodiments may be made without departing from the invention’s scope.Example 1Biosecurity Primitive: Polymerase X-based Genetic Physical Unclonable FunctionsA. Results and discussion

[0140] Characterization of TdT-mediated indels. The Pol X family member TdT is known to catalyze the addition of nucleotides to the 3' ends of DNA strands (FIG. 2A)(Motea and Berdis, Biochim Biophys Acta 1804, 1151-66, 2010; Ramsden and Asagoshi, Environ Mol Mutagen 53, 741-51, 2012; Lutze et al., Methods Mol Biol 2394, 299-317, 2022; Loveless et al., Nat Chem Biol 17, 739-47, 2021). We hypothesized that delivering TdT and Cas9 together would yield significant differences in the resulting indel landscape compared to the nuclease alone. We co-transfected pCMV-SpCas9-U6-sgRNAmKate2 (pl-Cas9) with and without pCMV- TdT (pl-TdT) plasmids into HCT116 cells that harbor a pCMV-mKate2 cassette at the CCR5 locus. The guide RNA targets the open reading frame of mKate2 (FIG. 2B). 72 hours posttransfection, we harvested genomic DNA, amplified the target locus via PCR, and performed next-generation amplicon sequencing.Primers TableAmplicon Sequences

[0141] To assess possible off-target effects of the CRISPR-Cas9 system, we also conducted in parallel T7 endonuclease I-based mutation detection assay of top predicted off-targets identified by CasOFFinder (Bae et al., Bioinformatics 30, 1473-75, 2014). No cleavage was observed in these samples, indicating no detectable off-target effects. In addition, we examined possible cytotoxic effects of TdT transfection via Annexin V / PI-based apoptosis assay and confirmed that moderate TdT overexpression used to generate PUFs does not substantially affect cell morphology or compromise viability.

[0142] The sequencing results were then analyzed using CRISPResso223 to showcase the diversity of TdT-induced indel profiles (FIG. 2C). CRISPR / SpCas9 alone predominantly induced deletions (approximately 84% of the resulting genetic alterations), with an indel distribution enriched for deletions of one or two base pairs. The addition of pl-TdT significantly altered this distribution, with 80% of the events being insertions, most of which were one to four base pairs in length. To illustrate the impact of insertion-dominant repair pathway on the entropy, we calculated the theoretical expected contribution to diversity for each indel profile by multiplying the frequencies of indel size by the number of possible replacements (FIG. 2D). To quantify theimpact of TdT-mediated insertions on the entropy of the edited locus, we computed Shannon entropy directly, following previous studies that used this metric to characterize the CRISPR- induced mutation profiles (Zou et al., Nature Cell Biology 2022 24:9 24, 1433-44, 2022; Kalhor et al., Nature Methods 2016 14:2 14, 195-200, 2016). To account for the difference in the number of reads between each sample, we performed a resampling procedure wherein we selected 100,000 indel reads without replacement, repeated this procedure 100 times to calculate the mean Shannon entropy of the subsamples. The results show that addition of pl-TdT nearly doubles the mean entropy from 0.86 to 1.74 bits at the target locus (FIG. 2E). Moreover, the subsampled population with pl-TdT contains a larger variety of unique indels and a pronounced shift toward insertions. Collectively, our results demonstrate that following a SpCas9-induced double-strand break, TdT can introduce random nucleotides at the exposed double-strand break during NHEJ-mediated repair and increase the indel complexity.

[0143] To further characterize TdT operation, we performed additional experiments in HEK293 cells. Plasmids carrying SpCas9 / gRNAAAvsi / mKate2 and TdT were co-transfected. After 72 hours, we enriched for transfected cells using fluorescence-activated cell sorting (FACS) (FIG. 2F). We extracted the genomic DNA from the mKate2+ population and amplified the AAVS1 locus via PCR. The PCR products were then subjected to next-generation amplicon sequencing in two batches (seq-TdTi-4 and seq-TdTs-e), each with their respective technical replicates. Sequencing results were analyzed based on our previously established pipeline. Briefly, the sequencing data from FASTQ files filtered with a regular expression-based method to remove corrupted reads. The regular expression pattern targeted three key regions: a variable barcode region, a flexible indel region spanning 20 bp upstream and downstream of the cut site, and fixed regions matching the reference sequence. After filtration, wild-type and substitution mutations were excluded to isolate the indel -containing reads. For TdT-mediated indels, the regular expression pattern was altered to accommodate insertions up to 10 bp in length. The filtered reads were further processed to extract indels, identify inserted or deleted sequences by alignment to the reference sequences, and generate indel frequency tables. With Cas9 alone, the top 20 indels consisted predominantly of deletions (FIG. 2G), whereas co-expression with pl- TdT shifted the distribution entire toward insertions (FIG. 2H). Moreover, we observed that the top 20 indels, which capture more than 30% of the total population, have consistent nucleotide content across all independent experiments. More specifically, 85% of the most frequentinsertions contained nucleotides C, G, or both, consistent with previous findings that the two nucleotides flanking the cleavage site significantly influence indel distributions and nucleotide composition (Loveless et al., Nat Chem Biol 17, 739-47, 2021; Gisler et al., Nature Communications 2019 10:1 10, 1-14, 2019).

[0144] To mitigate noise from sequencing data, our data analysis pipeline incorporates strict regular expression filters to discard corrupted reads and to isolate short insertions and deletions. To explore whether any remaining noise could still interfere with our ability to distinguish between intra- and inter-PUFs, we also performed a principal component analysis (PCA) on the observed indel compositions of the TdT-PUF samples. In this PCA space, the samples and their replicates (intra-PUFs) cluster tightly together, whereas samples from different transfections (inter-PUFs) remain well separated, suggesting that any remaining noise from sequencing is not expected to affect classification performance.

[0145] To quantify the similarity between indel frequencies in a population of cells resulting from individual transfections (i.e. inter-PUF distance), we applied the pairwise Bray-Curtis dissimilarity (BCD) metric (Bray and Curtis, Ecol Monogr 27, 325-49, 1957) to the 12 samples, which include seq-TdTi-6 and their respective replicates (seq-TdTirto TdTer)( Li et al., Sci Adv 8, 4106, 2022). We first used seq-TdTi as the reference and examined the indels of the top 5% to 100%, calculating the difference between intra-PUF (defined as the variation between a specific PUF and its corresponding repeat or freeze-thaw counterparts) and inter-PUFs (defined as the variation between two different PUFs). We determined that the top 25% of indels provide good separation between intra-PUF and inter-PUFs. We further calculated BCD between all sample pairs using the top 25% of indels and calculated the difference in fold change using each individual sample as the reference. The results show an average 3.408-fold difference between intra- and inter-PUFs, providing an acceptable classification.

[0146] TdT-mediated insertion dependence on time and genomic context. To assess the possibility that transfections at different times could affect the TdT-mediated indel distribution (e.g., due to differences in nucleotide availability and cell cycle stage), we conducted 3 independent transfections at 3 different time points using the same constructs as the previous experiment (FIG. 3A). For each timepoint, we selected mKate2+ cells via FACS 72 hours posttransfection and processed them using the same pipeline as in previous experiments. Across all time points, the most common insertion was two cysteines (CC), and the top 5 most commoninsertions were CC, C, GC, GG, and T (FIG. 3B). BCDs were calculated pairwise using the top 25% of indels. The results show that the insertion distributions are not influenced by transfection time. We therefore conclude that the main determinant of insertion entropy and nucleotide distribution is the location of the editing (i.e., nucleotides that flank the cleavage site)

[0147] We also assessed the temporal stability of the TdT-PUF design by generating an additional PUFs in HEK293 (seq-TdTsi) cells using the same protocol and cultured it for 10 passages (approximately two days per passage). At each passage, we extracted genomic DNA and PCR-amplified the edited regions as before. Using the transfected population as a reference (P0), we calculated the BCD across all 10 passages. When we compared this trend against the BCD to two other PUFs generated in parallel (seq-TdTs2 and seq-TdTss), we found that the dissimilarity due to temporal instability remained below the lowest inter-PUF distance, i.e. minimum identification threshold. These results demonstrate that our design remains stable for at least 10 passages or >20 days of continuous culturing. Notably, when compared to the first- generation 2D barcode-indel PUFs reported by Li et al., which exceeded the minimum identification threshold by passage 6, the TdT-PUF consistently maintained a lower dissimilarity metric throughout the time course.

[0148] To investigate whether different targeting loci and cleavage site nucleotide composition affect the TdT-mediated indel pattern, we performed additional experiments with the hRosa26 and CCR5 safe harbor sites (Irion et al., Nat Biotechnol 25, 1477-82, 2007; Scharenberg et al., Nat Commun 11, 2020; Papapetrou and Schambach, Mol Ther 24, 678-84, 2016) of HEK293 cells (FIG. 3C), following the same experimental protocol as the experiment targeting the AAVS1 site. For each site, three transfections were performed: seq-TdT?-9 for the hRosa26 site and seq-TdTio-12 for the CCR5 site. We again observed cleavage site-dependent insertion bias at both hRosa26 (AT) and CCR5 (CT) loci, highlighting the difference in indel compositions (FIG. 3D). For these two sites, the frequencies of A and T insertions were higher than what was observed at the AAVS1 site. For the hRosa26 site, A was present in the two most dominant indels (FIG. 3D, left), while for the CCR5 site, T was the most dominant indel (FIG. 3D, right).

[0149] To quantitatively assess the differences among sequences, we computed BCDs for all sequences at both the hRosa26 and CCR5 sites, using seq-TdT? and seq-TdTio as references. The results showed that using 25% of the indels can provide good separation between intra-PUFs andinter-PUFs. Herein, pairwise BCDs were calculated for samples of seq-TdT?-9 and their repeats (seq-TdT7r-9r) and for samples of seq-TdTio-12 and their repeats (seq-TdTior-i2r). As was observed with the AAVS1 locus, average differences between intra PUFs and inter PUFs were 2.810-fold and 3.929-fold for hRosa26 and CCR5 sites, respectively.

[0150] Enhancing PUF classification by removing common indels. For the results presented thus far, we quantified similarity between samples using a set of indel frequencies in a population of cells. We observed that simple post-sequencing processing steps can improve classification performance. First, based on the analysis of indel compositions and frequencies among samples, we identified a set of indels appearing in most samples at high frequencies. Identical entries across all samples typically lead to increased similarity scores when comparing vectors, thereby resulting in reduced dissimilarity measures (e.g., Bray-Curtis). In other words, while these common entries contribute to the overall similarity, they simultaneously reduce the discriminatory power of the comparison and our ability to robustly classify cell populations as we incorporate more PUFs.

[0151] To address this feature, we developed a post-sequencing processing method which starts by omitting sequences that only appear a single time in each sample. Subsequently, we designate a set (list without duplicates) of indel sequences from a specific sample (e.g. seq- TdTi), as the reference set and intersect it with every other sample (e.g. seq-TdT2-e) and their technical replicates (e.g. seq-TdT2r-6r). The intersection represents all common indels across samples, which we exclude from subsequent analysis. Lastly, the top 25% most abundant indels are used to compute pairwise BCD (FIG. 4A-4B). The results show that removing common indels significantly impacts the differences in indel composition across various populations. This is evident in the indel frequency table, which initially featured many of the same sequences before filtering (FIG. 2G), but now contains a heterogenous set of sequences (FIG. 4C). For example, while the insertion GTTCC was the most dominant indel in seq-TdTi, it was found to be absent in seq-TdT2. Further, the majority of frequent indels observed in seq-TdTi were absent in seq-TdTe (FIG. 4C). We proceeded with the pairwise BCD calculation using the top 25% most frequent indels. The results show that this procedure enhances the ability to distinguish between PUFed cell populations, increasing the average BCD fold change between intra-PUF and inter- PUFs from 2.338 to 3.217 when using seq-TdTi as reference (FIG. 4D). Applying this method to PUFs at different genomic loci, seq-TdT7.9 (hRosa26) and seq-TdTio-12 (CCR5), resulted insimilar fold change increases: from 2.662-fold to 4.349-fold at hRosa26 and from 3.568-fold to 6.161-fold at CCR5.

[0152] Classification of PUFs via logistic regression on TdT-mediated indels. While heuristic methods like the removal of common indels are sufficient for distinguishing between a small number of unique PUFs, we aimed to develop a more robust, data-driven approach that efficiently identifies relevant features for provenance attestation in a large library of unique PUFs. To accomplish this, we employed a feature selection method commonly used in machine learning (Chen and Jeong, Proceedings - 6th International Conference on Machine Learning and Applications, ICMLA 2007 429-435, 2007): multiclass logistic regression with an LI penalty to eliminate less significant features while isolating features most informative for classification (Huerta et al., Lecture Notes in Computer Science 7996 LNAI, 244-51, 2013; Huang et al., Applied Intelligence 48, 594-607, 2018; Darst et al., BMC Genet 19, 2018; Bahl et al., NanoImpact 15, 100179, 2019). After training a logistic regression model on 12 PUFs generated across various loci, we evaluated its ability to classify the intra-PUF that were included in the training set (e.g., seq-TdTi and its corresponding repeat seq-TdTir) (FIG. 5A). We observed that the model correctly classified the repeat with probabilities ranging from 96.2% to 99.9% (FIG. 5B). This high level of accuracy indicates that the model can effectively distinguish between intra- and inter-PUFs.

[0153] Next, we systematically evaluated the scalability (i.e. the number of PUFs that we can produce without risk of misidentification) of the feature selection method. First, we implemented TdT-PUFs in alternative cellular contexts. We used the same experimental pipeline to construct PUFs in three additional human cell lines (HeLa, HCT116, A549) and Chinese hamster ovary (CHO) cells (FIG. 5C). TdT-mediated NHEJR was performed in the AA VS1 locus using the same sgRNA target sequence, except for the CHO cell line, where we used the Rosa26 locus. We then performed pairwise BCDs for all samples of the same cell lines and calculated the average differences between intra-PUFs and inter-PUFs. The average distances are 4.006-fold, 5.039- fold, 5.627-fold, and 1.661-fold for HeLa, HCT116, A549, and CHO cell lines separately. Retraining the model on the expanded PUF library with L2 penalty retained classification accuracy at >99 % for every sample, demonstrating that supervised learning method scales effectively to an expanding set of PUF signatures (FIG. 5D).

[0154] To introduce additional samples, we first concatenated all indel sequences and their respective frequencies from all experiments, creating a comprehensive probability distribution of observed indels. We then generated synthetic PUFs by randomly sampling (with replacement) from this distribution (FIG. 5E). To create replicates of these simulated PUFs, we generated corresponding subsets with a 77% overlap in composition, reflecting the average overlap observed in experimentally derived PUFs. We then trained logistic regression models for 25 to 300 simulated PUF samples and assessed the algorithm’s ability to accurately discriminate between each intra- and inter-PUF samples (FIG. 5F). With the simulated PUFs, we observed a minimum average prediction accuracy above 95.2%. To identify the factor most affecting classification performance, we repeated the same procedure but limited the size of insertions in our simulated PUFs to a range of 3 to 7 base pairs. Since TdT-based indel formation contributes to entropy primarily through insertions, we reasoned that constraints on the size of insertions would hamper the ability to classify PUFs in large populations. As expected, reducing the indel size results in a significant number of samples with classification accuracy below 95%. Taken together, these simulations demonstrate how factors that contribute to the overall repair entropy (i.e., indel sequence identity, size, and frequency) impact PUF classification. Importantly, the results show that the TdT-induced indel profiles yield sufficient entropy to afford robust classification of at least 100 PUFs (produced in a single batch) (FIG. 5F).

[0155] Cell line misidentification has been a persistent issue compromising the integrity of biomedical research since the 1960s (Harbut et al., SLAS Discovery 100194, 2024). Despite advances in biotechnology and genome engineering, a reliable and practical standard for cell line authentication remains elusive, leading to significant waste of resources and potential misinterpretation of experimental results. In this manuscript, we present a novel solution to this critical problem by developing genetic PUFs that leverage the intrinsic entropy of DNA lesion repair mediated by TdT, a member of the X-family of DNA polymerases. Our method exploits the natural variability introduced during TdT-mediated repair, generating robust, unique, and virtually unclonable genetic identifiers directly within human cells. For fast and accurate authentication process, we implemented a machine learning-assisted classifier that analyzes nextgeneration sequencing data to precisely discriminate between individual genetic PUFs.

[0156] A key property of genetic PUFs is the inherent randomness arising from variations in the manufacturing protocol. Accordingly, we investigated the impact of transfection timing andgenomic location selection for the PUF engineering. Our results show that the resulting indel nucleotide frequencies and distributions are independent of editing time. We also investigated the impact of TdT-mediated indel generation across various genetic loci, each with a distinct cut site. Specifically, we evaluated 3 different sgRNAs with cleavage sites CG, AT, and CT; the results indicated unique indel outcomes at each of the target sites. Both parameters can be further explored to produce new generations of genetic PUFs. For example, considering that supplementing cultured cells with deoxyribonucleosides can alter TdT-mediated indel outcomes (Callisto et al., bioRxiv, 2024), there is a direct opportunity to control the PUF entropy during manufacturing. Additionally, Cas variants with broadened compatibility for protospacer adjacent motifs (PAM)(Hu et al., Nature 2018 556:7699 556, 57-63, 2018; Xu and Li, Comput Struct Biotechnol J 18, 2401, 2020) can be utilized to target adjacent genome sites during the PUF engineering.

[0157] To demonstrate the versatility of TdT-based PUFs, we applied the system across several commonly used immortalized cell lines, including HEK293, HCT116, HeLa, A549, and CHO. In these cell lines, TdT overexpression was well-tolerated, with no detectable impact on cell viability or morphology. However, when applied to a primary fibroblast cell line (CRL- 2522), we observed both reduced editing efficiency and increased cytotoxicity upon coexpression of Cas9 and TdT. These results suggest that while TdT-PUFs are broadly compatible with standard cell culture models, their implementation in primary or stem cells will require further optimization.

[0158] The proposed method offers significant improvements in the PUF manufacturing protocol (FIG. 6). By eliminating the need for barcoding prior to indel generation, we have reduced the production time from a few months to days. This change also decreases the total cost of engineering each PUF cell line by 3 -fold, including reagents, consumables, and NGS service. Despite these reductions in time and cost, it is important to note that the TdT-based PUFs maintain similar levels of unclonability and robustness. Manual construction of a single Barcode- Indel PUF requires establishing 500 individual cell lines and mixing them properly, where the TdT PUF requires building 380 cell lines. Therefore, both methods are sufficiently secure to prevent unauthorized reproduction. Finally, incorporating logistic regression into our model ensures robust classification of TdT-based PUFs.

[0159] To conclude, we demonstrate that TdT can be effectively harnessed to enhance intrinsic entropy during DNA lesion repair, enabling the rapid production of genetic PUFs that meet the criteria of robustness, uniqueness, and unclonability. Our results not only provide novel insights into the function of TdT but also, looking forward, represent a major advancement toward the practical application of genetic PUFs as a biosecurity primitive for cell line authentication and provenance verification, addressing a longstanding challenge in biomedical research.B. Methods

[0160] Plasmid cloning. The plasmid of pCMV-TdT was obtained from Addgene (126450). The pCMV-spCas9-T2A-mKate2-U6-sgRNA_AAVS l / sgRNA_CCR5 / sgRNA_Rosa26 were cloned through standard molecular cloning protocol, briefly, the oligos containing the desired spacer sequences were amplified using PCR and insert into the vector plasmid using restriction enzyme digestion and ligated by T4 DNA ligases. Transformations were performed using NEB® 5-alpha Competent E. coli (NEB, catalog C2987H). The plasmids were harvested and purified using the QIAprep Spin Miniprep Kit (Qiagen, catalog 27104).

[0161] Cell culture and transfection. HEK293 cells were obtained from ATCC(CRL-1573) and were cultured in Dulbecco's Modified Eagle Medium - high glucose (Sigma-Aldrich, catalog.D5796), supplemented with 10% FBS, 1% Non-essential Amino Acid (Gibco, catalog. 11140076) and 1% Penicillin- Streptomycin (Gibco, catalog. 15070063), at 37 °C and 5% CO2. Transient transfections of HEK293 cells were performed using JetPrime (Polyplus, ref. 101000046). Transfections were performed on 6-Well plates, 450,000 cells were seeded into each well 24 hours before transfection. For each well, 400ng pCMV-TdT plasmid and 1600ng pCMV-SpCas9-T2A-mKate2-U6-sgRNA_AAVS l / sgRNA_CCR5 / sgRNA_Rosa26 plasmids, a total amount of 2000ng DNA, mixed with 200 pl buffer and 4 pl JetPrime reagent, were transfected into the cells.

[0162] Genomic DNA extraction from selected cells. 72 hours post -transfection, the media was removed, and 500 pl 0.25% trypsin (Gibco, catalog. 25200114) was added to each, the plate was incubated at 37 °C for 5 minutes. After incubation, 1.5 ml of complete media was added to each well to neutralize trypsin. Then the cell suspension was pelleted by centrifuge at 1500 rpm for 5 minutes. Removing the supernatant, the cell pellets were resuspended in 600 pl PBS(Coming, Catalog. 21040CV). Subsequently, the cell suspensions were conducted FACS analysis to collect around 200,000 cells that expressed mKate2. The gDNA were extracted using DNeasy Blood & Tissue Kits (Qiagen, Catalog. 69504), following the standard protocol.

[0163] NGS of the targeted region. The DNA fragments of the targeted regions were PCR- amplified using around 100 ng extracted gDNA as the template. The PCR conditions followed the standard protocol of Q5 Hot Start Master Mix (2X)(NEB, Catalog. M0494S), 98°C ~ 30s, (98°C ~10s, 63°C ~30s, 72°C ~30s) *35, 72°C ~ 2mins, 4 °C -forever. The size of PCR products was examined by gel electrophoresis, the correct products will be purified using HiBind® DNA Mini Columns (Omega Bio-TEK, catalog. DNACOL-01). The purified amplicons were sent to Genewiz, Inc. for Amplicon-EZ sequencing, where they were further processed for library preparation and then sequenced on an Illumina Miseq for 2*250 pair-end sequencing.

[0164] Sequencing data analysis. The raw data were initially analyzed based on our previously established pipeline with modification for TdT-mediated insertions (4), which included the following steps: 1. Extract DNA sequence from R1 fastq files. 2. Filter out corrupted reads following the rules, which allowed deletions to happen within 20bp around the predicted cutting site, insertions are equal or smaller than 10 bp, and the rest sequences should be matched to the reference sequence. 3. Extract indel sequence. 4. Remove the length of the indel was 40bp, which include the wild-type sequence and substitutions. 5. Identify insertion or deletion by aligning to the reference sequences Indel lists were further analyzed using Python.

[0165] Bray-Curtis dissimilarity calculation. To calculate Bray-Curtis dissimilarities between sample i and sample j, they were first sorted into the same order and then were calculated by the following equation:fc i where n is the frequency of index k in each sample.

[0166] Entropy calculation. Shannon Entropies are calculated using the following equation:where i is the index of each unique indel, is the frequency of indel i, m is the total number of unique indels.

[0167] Logistic regression model. A logistic regression model was trained using the Scikit- leam library in Python, with ‘liblinear’ as the solver for all models. First, all PUF indel profiles were preprocessed into a standardized format, with indel sequences and their normalized counts represented as feature vectors. The dataset was then divided into a training set (data from original PUF samples) and a test set (data from replicate PUF samples). After training, the test dataset was used to examine the performance of the model. Both LI (Lasso) and L2 (Ridge) regularization penalties were tested. Model performance was evaluated on the test set to determine classification accuracy. The full logistic regression training and evaluation script is available online.

Claims

CEAIMS1. A method for authenticating an unknown cell line comprising the steps of:(a) synthesizing a plurality of genetically engineered cell lines, wherein a set of random insertions and deletions (indels) is generated across the plurality of cell lines and wherein each individual cell comprises a unique subset of random indels, thereby defining a genetic physical unclonable function (PUF) for each cell;(b) creating a set of authentication classifiers and corresponding threshold dissimilarity values for each of the plurality of cell lines by statistically characterizing and selecting informative subsets of random indels to create unique authentication classifiers for each PUF and a corresponding threshold dissimilarity value for each PUF; and(c) authenticating an unknown cell line by measuring the statistical dissimilarity between the frequency distribution of indels in the sequenced unknown cell line and the set of authentication classifiers, and determining that the PUF from step (a) is present when the statistical dissimilarity is less than the threshold dissimilarity value associated with the authentication classifier.

2. The method of claim 1, wherein the unknown cell line is selected from the group consisting of mammalian cells, non-mammalian eukaryotic cells, prokaryotic cells, plant cells, viral particles, and synthetic or engineered cells.

3. The method of claim 1, wherein the step of synthesizing a plurality of cell lines is accomplished by:(i) introducing dsDNA lesions into the genome of each of a plurality of cell lines using at least one endonuclease targeted to a safe-harbor nucleic acid sequence; and(ii) using a nucleic acid-repair protein to facilitate noisy DNA lesion repair at each dsDNA lesion through a non-homologous end joining (NHEJ) mechanism, in each of a plurality of cell lines, creating the subset of random indels (PUF) in each of the plurality of cell lines, wherein each random indel is random in both sequence and length.

4. The method of claim 3, wherein the endonuclease is one of the following: a CRISPR-Cas endonuclease, a transcription activator-like effector nuclease (TALEN), or a zinc finger nuclease.

5. The method of claim 3, wherein the nucleic acid-repair protein is a member of the polymerase- X family of proteins.

6. The method of claim 5, wherein the member of the polymerase-X family of proteins is Terminal deoxynucleotidyl Transferase (TdT).

7. The method of claim 1, wherein the step of creating a set of authentication classifiers comprises:(i) sequencing replicate samples of each cell line to identify the subset of random indels in each replicate;(ii) removing from each replicate any indel sequences appearing only once;(iii) calculating a frequency distribution of indels for each replicate;(iv) repeating steps (i) through (iii) for each cell line to generate an initial set of indel sets and corresponding frequency distributions;(v) compiling a combined list of the most common indels across all cell lines and ranking them by frequency;(vi) removing a predetermined percentage of the most frequent indels from the initial sets to create revised indel sets and corresponding frequency distributions;(vii) selecting one revised indel set as a reference and calculating statistical dissimilarity between the reference and each other revised indel set for all replicates;(viii) iteratively repeating steps (vi) and (vii) until intra-PUF dissimilarity falls below a predetermined threshold and inter-PUF dissimilarity exceeds a predetermined threshold, resulting in a final indel set for each replicate;(ix) designating the final indel set and its frequency distribution as the authentication classifier for that replicate;(x) repeating step (ix) for each cell line to create the set of authentication classifiers; and(xi) determining a threshold dissimilarity value for each authentication classifier based on measured intra-PUF and inter-PUF dissimilarities.

8. The method of claim 1, wherein the step of creating a set of authentication classifiers is accomplished by:(i) training a multiclass logistic regression model using frequency distributions of indels derived from a plurality of PUFs generated across one or more genomic loci;(ii) applying an LI regularization penalty to eliminate indels with low discriminative power and select a subset of indels most informative for classification;(iii) generating, for each PUF, an authentication classifier based on the selected subset of indels; and(iv) retraining the model on an expanded PUF library using an L2 regularization penalty to preserve classification accuracy when scaling to large PUF populations.

9. The method of claim 7 or claim 8, wherein the sequencing of replicate samples of a cell line is accomplished using a next generation amplicon sequencing technique.

10. The method of claim 7 or claim 8, wherein the step of removing a pre-determined percentage of the most common random indels, the pre-determined percentage is in the range of 1-50%.

11. The method of claim 7, wherein the step of measuring a statistical dissimilarity is accomplished using a Bray-Curtis dissimilarity metric.

12. The method of claim 1, wherein the step of authenticating a cell line is accomplished by:(i) sequencing replicate samples of an unknown cell line to characterize the frequency distribution of indels in the cell line;(ii) creating a first list of random indels and corresponding frequency distributions for each replicate sample of the unknown cell line;(iii) creating a second list of random indels and corresponding frequency distributions in the unknown cell line by removing indels and corresponding frequency distributions that are not part of the authentication classifier for the cell line;(iv) measuring the statistical dissimilarity between the corresponding frequency distributions of the second list of random indels in the unknown cell line and the authentication classifier for the cell line; and(v) designating the cell line as authentic if the statistical dissimilarity between the second list of random indels in the cell line and the authentication classifier is below the threshold value.

13. The method of claim 12 wherein the sequencing of the cell line in replicates is accomplished using a next generation amplicon sequencing technique.

14. The method of claim 13, wherein the step of measuring the statistical dissimilarity between the second list of random indels in the cell line and the authentication classifier is accomplished using the Bray-Curtis dissimilarity metric.

15. A method of creating a genetic physical unclonable function (PUF) in a cell line, comprising:(i) introducing dsDNA lesions into the genome of each of a plurality of cell lines using at least one endonuclease targeted to a safe-harbor nucleic acid sequence; and(ii) using a nucleic acid-repair protein to facilitate noisy DNA lesion repair at each dsDNA lesion through a non-homologous end joining (NHEJ) mechanism, in each of a plurality of cell lines, creating the subset of random indels (PUF) in each of the plurality of cell lines, wherein each random indel is random in both sequence and length.

16. The method of claim 15 wherein the endonuclease is one of the following: a CRISPR-Cas endonuclease, a transcription activator-like effector nuclease (TALEN), or a zinc finger nuclease.

17. The method of claim 15 wherein the nucleic acid-repair protein is a member of the polymerase-X family of proteins.

18. The method of claim 16 wherein the member of the polymerase-X family of proteins is Terminal deoxynucleotidyl Transferase (TdT).

19. A system for authenticating a cell line, comprising:(a) a non-transitory computer-readable medium storing instructions; and(b) a processing device configured to execute the instructions, wherein the instructions, when executed, cause the processing device to:(i) receive sequencing data from one or more replicate samples of an unknown cell line and generate a first list of random insertions and deletions (indels) and corresponding frequency distributions for each replicate sample;(ii) generate a second list of random indels and corresponding frequency distributions for the unknown cell line by filtering out indels and corresponding frequency distributions that are not part of a previously stored authentication classifier;(iii) compute a statistical dissimilarity measure between the second list of random indels for the unknown cell line and the authentication classifier; and(iv) designate the unknown cell line as authentic when the computed statistical dissimilarity measure is below a threshold dissimilarity value associated with the authentication classifier.

20. The system of claim 19, wherein the instructions, when executed, further cause the processing device to:(i) train a multiclass logistic regression model using frequency distributions of indels derived from a plurality of genetic physical unclonable functions (PUFs) generated across one or more genomic loci;(ii) apply an LI regularization penalty to identify and retain indels with the highest discriminative power, thereby selecting a subset of indels most informative for classification;(iii) generate, for each PUF, an authentication classifier based on the selected subset of indels; and(iv) retrain the model on an expanded PUF library using an L2 regularization penalty to maintain classification accuracy when scaling to large PUF populations.

21. The system of claim 19 or claim 20, wherein the instructions, when executed, further cause the processing device to:(i) implement one or more machine learning models selected from the group consisting of random forests, support vector machines (SVMs), gradient-boosted decision trees, artificial neural networks, convolutional neural networks (CNNs), recurrent neural networks (RNNs), transformer-based models, and ensemble classifiers combining two or more of the foregoing;(ii) train the selected machine learning model(s) using indel frequency distributions derived from a plurality of PUFs;(iii) generate, for each PUF, an authentication classifier based on the trained model(s); and(iv) authenticate an unknown cell line by predicting classifier membership based on the sequenced indel frequency distribution and designating the cell line as authentic when the prediction confidence exceeds a predetermined probability threshold.

22. The system of claim 21, wherein the instructions, when executed, further cause the processing device to:(i) receive sequencing data generated from a plurality of genetically engineered cell lines, each comprising a unique set of random insertions and deletions (indels) created by: introducing targeted double-stranded DNA (dsDNA) lesions at one or more safe-harbor genomic loci using at least one endonuclease; and facilitating noisy DNA lesion repair through a non-homologous end joining (NHEJ) mechanism mediated by a nucleic acid-repair protein, thereby creating a physical unclonable function (PUF) for each cell line;(ii) preprocess the sequencing data to compute frequency distributions of indels for each PUF;(iii) train one or more machine learning models as described in claim 20 using the computed indel frequency distributions to generate authentication classifiers for each PUF;(iv) receive sequencing data from an unknown cell line and compute its indel frequency distribution; and(v) authenticate the unknown cell line by predicting classifier membership using the trained machine learning model(s) and designating the cell line as authentic when the prediction confidence exceeds a predetermined probability threshold.

23. The method of claim 1, wherein the genetic physical unclonable function (PUF) comprises one or more genomic or epigenomic features selected from the group consisting of insertions, deletions, substitutions, single-nucleotide polymorphisms (SNPs), copy number variations, structural rearrangements, targeted or random base editing events, epigenetic modifications including DNA methylation and histone marks, and RNA editing events, wherein each cell comprises a unique subset of such features, thereby defining a PUF signature.

Citation Information

Patent Citations

  • Use of terminal deoxynucleotidyl transferase for mutagenic DNA repair to generate variability, at a determined position in DNA

    US20130165347A1

  • Genetic physical unclonable functions and methods of use thereof

    US20230183749A1