Method of mapping transgene integration

WO2025240423A3PCT designated stage Publication Date: 2026-05-28TACONIC BIOSCIENCES INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
TACONIC BIOSCIENCES INC
Filing Date
2025-05-13
Publication Date
2026-05-28

AI Technical Summary

Technical Problem

Current transgene mapping techniques require animal sacrifice for DNA harvesting, increasing costs and complexity, and often necessitate transgene-specific anchor primers, which are predicated on known transgene locations.

Method used

A computer-implemented method aligns sequence reads of transgenic organisms to a custom genome, identifying candidate junction reads and categorizing them into clusters to determine transgene integration sites without the need for animal sacrifice, using a custom genome comprising a reference genome sequence and a transgene sequence.

Benefits of technology

This approach allows for accurate mapping of transgene integration sites in transgenic organisms without animal sacrifice, reducing costs and complexity, and enabling precise determination of transgene locations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025029070_28052026_PF_FP_ABST
    Figure US2025029070_28052026_PF_FP_ABST
Patent Text Reader

Abstract

Provided is a computer implemented method for mapping transgene integration into n organism by providing a custom genome comprising a reference genome sequence and a transgene sequence, aligning sequence reads of the transgenic organism to the custom genome, identifying candidate junction reads, wherein the candidate junction reads correspond to sequence reads comprising sequence regions that align to the transgene sequence and to the host sequence, categorizing the candidate junction reads into junction clusters, wherein each of the junction clusters comprises a plurality of the candidate junction reads each with an alignment segment terminus within a cluster proximity region of the custom genome, determining a transgene junction based on the junction clusters that possesses at least three of the candidate junction reads within the cluster proximity region of the custom genome and mapping the integration of the transgene into the transgenic organism based on the transgene junctions.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No.: 0303.052AWO METHOD OF MAPPING TRANSGENE INTEGRATION CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims benefit of priority from U.S. Provisional Patent Application No.63 / 647,464 filed May 14, 2024, and U.S. Provisional Patent Application No.63 / 696,109, filed September 18, 2024, the entire contents of which are incorporated herein by reference. BACKGROUND

[0002] The present disclosure relates to mapping transgene sequences within transgenic organisms. Knowledge of where one or more transgenes are integrated in a host genome is important for planning crosses between experimental animal models. Integration of a transgene can disrupt an endogenous gene, potentially interfering with interpretation of the transgenic phenotype or preventing the generation of homozygous transgenic animals due to embryonic lethality when the transgene is bred to homozygosity. Current transgene mapping techniques may require animal sacrifice during cell harvesting used to acquire DNA necessary for transgene mapping resulting in the loss of the animal for further research and potential animal breeding programs. For instance, techniques for transgene sequence mapping such as Targeted Locus Amplification (TLA) require the harvesting of splenocytes and the sacrifice of the transgenic animal increasing the cost and complexity of creating transgenic animals. TLA also requires transgene specific anchor primers which require some prediction of the location and orientation of the transgene sequence. The present disclosure is directed to overcoming these and other deficiencies in the art and additional advantages are provided through the provision of a computer-implemented method, computer system, and computer program product. SUMMARY

[0003] In an aspect, disclosed is a method of mapping one or more transgene integration into a transgenic organism including providing a custom genome comprising a reference genome sequence and a transgene sequence; aligning sequence reads of the transgenic organism to the custom genome; identifying candidate junction reads, wherein the candidate junction reads correspond to sequence reads comprising sequence regions that align to the transgene sequence and to the host sequence; categorizing the candidate junction reads into junction clusters, wherein each of the junction clusters comprises a plurality of the candidate junction reads each with an alignment segment terminus within a cluster proximity region of the custom genome; determining a transgene junction based on the junction clusters that possesses at least three of theAttorney Docket No.: 0303.052AWO candidate junction reads within the cluster proximity region of the custom genome; and mapping the integration of the transgene into the transgenic organism based on the transgene junctions.

[0004] In an example, the sequence reads are about 750 base pairs or greater. In a further example, the sequence reads include whole genome sequencing reads. In still a further example, the sequence reads include sequence reads derived from whole genome sequencing reads with at least a 10x coverage of genome of the transgenic organism.

[0005] In yet a further example, a candidate junction read includes a sequence region wherein a minimum of about 50 nucleotides align to a reference genome sequence of the custom genome and wherein a minimum of about 50 nucleotides align to a transgene sequence of the custom genome. In a further example, a candidate junction read includes a sequence region alignment to a reference genome sequence of a custom genome with a minimum pairwise alignment of about 90% in the sequence region alignment and wherein the candidate junction read includes a sequence region alignment to a transgene sequence of the custom genome with a minimum pairwise alignment of about 90% in the sequence region alignment.

[0006] In a still further example, a cluster proximity region includes a region of about 10-100 nucleotides of the reference genome sequence of the custom genome. In an example, a cluster proximity region includes a region of about 10-100 nucleotides between at least two or more alignment segment termini aligning to the same chromosome of the custom genome.

[0007] In a further example, mapping the integration of the transgene into the transgenic organism further includes de novo assembly of the transgene junction using the candidate junction clusters. In still a further example, mapping the integration of the transgene into the transgenic organism further includes mapping the transgene junctions to the custom genome to determine the junction sequence. In an example, mapping the integration of the transgene into the transgenic organism includes (iii) identifying candidate junction reads including storing candidate junction read alignments in Pairwise Mapping Format (PAF) files.

[0008] In an aspect, disclosed is a computer system for mapping one or more transgene integration into a transgenic organism, the computer system including memory and at least one processor, the computer system configured to execute program instructions to perform a method including: providing a custom genome including a reference genome sequence and a transgene sequence; aligning sequence reads of the transgenic organism to the custom genome; identifying candidate junction reads, wherein the candidate junction reads correlate to candidate junctionsAttorney Docket No.: 0303.052AWO comprising sequence regions that align to the transgene sequence and to the host sequence; categorizing the candidate junctions into junction clusters, wherein each of the junction clusters comprises a plurality of the candidate junctions each with an alignment segment terminus within a cluster proximity region within the custom genome; determining a transgene junction based on the junction clusters that possesses at least three of the candidate junctions within the cluster proximity region of the custom genome; and mapping the integration of the transgene into the transgenic organism based on the transgene junctions.

[0009] In an example, the system may include sequence reads that are about 750 base pairs or greater. In yet a further example, the system may include sequence reads derived from whole genome sequencing reads. In still a further example, the system may include the whole genome sequencing reads with at least a 10x coverage of genome of the transgenic organism.

[0010] In yet a further example, the system may include a candidate junction read which may include a sequence region wherein a minimum of about 50 nucleotides align to a reference genome sequence of the custom genome and wherein a minimum of about 50 nucleotides align to a transgene sequence of the custom genome. In a further example, the system may include a candidate junction read which may include a sequence region alignment to a reference genome sequence of a custom genome with a minimum pairwise alignment of about 90% in the sequence region alignment and wherein the candidate junction read comprises a sequence region alignment to a transgene sequence of the custom genome with a minimum pairwise alignment of about 90% in the sequence region alignment.

[0011] In a still further example, the system may include a cluster proximity region which may include a region of about 10-100 nucleotides. In an example, the system may include a cluster proximity region which may include a region of about 10-100 nucleotides between at least two or more alignment segment termini aligning to the same chromosome of the custom genome. In yet a further example, the system may include candidate junction clusters used for de novo assembly of the transgene junction. In still a further example, the system may include a cluster proximity region which may include a region of about 10-100 nucleotides of the custom genome as defined by at least one alignment segment terminus.

[0012] In yet a further example, the system may include mapping the integration of the transgene into the transgenic organism further including de novo assembly of the transgene junction using the candidate junction clusters. In a further example, the system may includeAttorney Docket No.: 0303.052AWO mapping the integration of the transgene into the transgenic organism further including mapping the transgene junctions to the custom genome to determine the junction sequence. In still a further example, the system may include (iii) identifying candidate junction reads including storing candidate junction read alignments in Pairwise Mapping Format (PAF) files.

[0013] In an aspect, disclosed is a computer program product for mapping one or more transgene integration into a transgenic organism, the computer program product including: a tangible storage medium storing program instructions for execution to perform a method including: providing a custom genome comprising a reference genome sequence and a transgene sequence; aligning sequence reads of the transgenic organism to the custom genome; identifying candidate junction reads, wherein the candidate junction reads correlate to candidate junctions comprising sequence regions that align to the transgene sequence and to the host sequence; categorizing the candidate junctions into junction clusters, wherein each of the junction clusters comprises a plurality of the candidate junctions each with an alignment segment terminus that maps to a cluster proximity region within the custom genome; determining a transgene junction based on the junction clusters that possesses at least three of the candidate junctions mapping to the cluster proximity region of the custom genome; and mapping the integration of the transgene into the transgenic organism based on the transgene junctions.

[0014] In an example, the computer program may include sequence reads that are about 750 base pairs or greater. In yet a further example, the computer program may include the sequence reads include whole genome sequencing reads. In still a further example, the computer program may include sequence reads including sequence reads derived from whole genome sequencing reads with at least a 10x coverage of genome of the transgenic organism.

[0015] In still a further example, the computer program may include a candidate junction read including a sequence region wherein a minimum of about 50 nucleotides align to a reference genome sequence of the custom genome and wherein a minimum of about 50 nucleotides align to a transgene sequence of the custom genome. In a further example, the computer program may include a candidate junction read including a sequence region alignment to a reference genome sequence of a custom genome with a minimum pairwise alignment of about 90% in the sequence region alignment and wherein the candidate junction read comprises a sequence region alignment to a transgene sequence of the custom genome with a minimum pairwise alignment of about 90% in the sequence region alignment.Attorney Docket No.: 0303.052AWO

[0016] In a still further example, the computer program may include a cluster proximity region including a region of about 10-100 nucleotides of the reference genome sequence of the custom genome. In an example, the computer program may include a cluster proximity region including a region of about 10-100 nucleotides between at least two or more alignment segment termini aligning to the same chromosome of the custom genome.

[0017] In yet a further example, the computer program may include mapping the integration of the transgene into the transgenic organism further including de novo assembly of the transgene junction using the candidate junction clusters. In still a further example, the computer program may include mapping the integration of the transgene into the transgenic organism further including mapping the transgene junctions to the custom genome to determine the junction sequence. In yet a further example, the computer program may include (iii) identifying candidate junction reads including storing candidate junction read alignments in Pairwise Mapping Format (PAF) files. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] These and other features, aspects, and advantages of the present disclosure will become better understood when the following detailed description is read with reference to the accompanying drawings, wherein

[0019] FIG.1 illustrates non-limiting embodiments of sequence read alignments.

[0020] FIG.2 illustrates non-limiting embodiments of categorizing candidate junction reads into junction clusters.

[0021] FIG.3 illustrates non-limiting embodiments of a method for mapping one or more transgene integrations sites into a transgenic organism.

[0022] FIGs.4A and 4B illustrate non-limiting embodiments of visualizations depicting sequence coverage surrounding transgene junctions of a transgene inserted into an organism’s chromosome in accordance with aspects of the present disclosure.

[0023] FIG.5 illustrates non-limiting embodiment of a visualization depicting annotations of an inserted transgene sequence in accordance with aspects of the present disclosure. DETAILED DESCRIPTION

[0024] This disclosure relates to methods of mapping one or more transgene sequence in transgenic organisms. Gene mapping refers to the process of determining the location of genes on chromosomes. An efficient approach for gene mapping involves sequencing a genome andAttorney Docket No.: 0303.052AWO then using computer programs to analyze the sequence to identify the location of genes. Genome sequencing allows genes and transgenes to be mapped to the physical locations where they reside in the genome.

[0025] The ability to introduce transgenes into the germline genome has rendered the mouse a powerful and indispensable experimental model in fundamental and medical research. The DNA sequences may be integrated into the genome randomly or into a specific locus by homologous recombination. Transgene integration may be used in order to delete or insert mutations into genes of interest to determine their function, introduce human genes into the genome of mice, to generate animal models enabling study of human-specific genes and diseases, developing mice susceptible to infections by human-specific pathogens of interest, introduce individual genes or genomes of pathogens and introduce reporter genes that allow monitoring in vivo or ex vivo the expression of genes of interest.

[0026] A transgene may include an exogenous nucleic acid sequence, which is not present in a genome of an organism into which the transgene is being inserted. In some cases, a transgene may include an endogenous nucleotide sequence, which is present in the genome of an organism into which the transgene is being inserted. Transgenes may include one or more exogenous sequence, one or more endogenous sequence, or one or more of both. Transgenes may include DNA sequences, RNA sequences or both. Transgenes may include linear nucleic acid, circular nucleic acid or both. Transgenes may be introduced into an organism by a variety of techniques to generate a transgenic organism.

[0027] The most prevalent technique for introducing genes into the mouse germ line is direct microinjection of cloned DNA into pronuclei of fertilized eggs. Transgenes may be injected into the cytoplasm, into the nuclei of two-cell embryos, and into the blastocoel cavity. Transgenes may be introduced by infection of preimplantation embryos with natural or genetically engineered retroviruses. Another method of introducing genes into the germ line involves introducing DNA into totipotent teratocarcinoma cells or into embryonic stem cells and then incorporating these cells into the blastocyst of developing embryos or aggregating them with eight-cell embryos. In some embodiments, pronuclear microinjection allows the large DNA constructs of virtually any size to be injected into an organism. In some embodiments a lentiviral vector can be used to insert about 10 kilobase pairs or less of nucleic acid. In some embodimentsAttorney Docket No.: 0303.052AWO plasmids are convenient and efficient cloning vectors for carrying out a variety of recombinant DNA procedures.

[0028] Plasmid-based transgenes may be relatively small and may have decreased efficiency over 25 kb. Cosmids are modified plasmids which contain the cos sequences from bacteriophage lambda. Cosmids can be handled and propagated in Escherichia coli similar to plasmid vectors and may accept DNA fragments up to 47 kb. Bacterial Artificial Chromosomes (BACs) clones are usually 100–300 kb long and can be used for generating transgenic mouse lines. An advantage of BACs is that they are more likely to contain all the genomic regulatory elements. The bacteriophage P1 cloning system can accept DNA inserts of 100 kb, and its modified form, PAC (P1 artificial chromosome) vector can typically carry 100–250 kb of DNA. For a small portion of mammalian genes, even these large vectors are not large enough to contain the entire gene or the flanking regions important for expression. Yeast artificial chromosomes (YACs) can accommodate a couple of megabases of DNA, and human artificial chromosomes (HACs) can carry more than 10 Mb. Both vector types have been successfully used to generate transgenic mice.

[0029] Mapping refers to the process of determining the relative locations of landmarks or markers (such as genes, variants, transgenes and other DNA sequences of interest) within a chromosome or genome. Mapping one or more transgene integration site into a transgenic organism may include sequencing the junctions of where one or more transgenes integrates into chromosomal DNA of a host organism. Sequencing the transgene junction where it contacts the host genome sequence provides clues about the events that occurred during transgene integration. Deletions, duplications, and translocations of chromosomal sequences may occur at the site of integration, and some junctions may contain short novel DNA sequences found neither in the injected DNA nor host organism prior to transgene integration. Foreign DNA may be stably transmitted for many generations with no evidence of rearrangement. Transgenes may be rearranged, partially deleted, or amplified. Transgenes may replicate extrachromosomally or chromosomally. The germ-line transformation of mice is discussed in Palmiter, R.D.; Brinster, R.L. Germ-line transformation of mice. Annu. Rev. Genet.1986, 20, 465–499 which is herein incorporated by reference in its entirety. Transgenic mice are discussed in Kumar TR, Larson M, Wang H, McDermott J, Bronshteyn I. Transgenic mouse technology: principles and methods. Methods Mol Biol.2009, which is herein incorporated by reference in its entirety.Attorney Docket No.: 0303.052AWO

[0030] Some transgenes are expressed exclusively in one cell type, others are expressed in only a few cell types, and some are expressed in most cells. Transgenes may be tissue specific and only be expressed in a specific tissue. Transgene expression may vary with individuals and may vary depending on the copy number of the transgene. Multiple copies of transgenes may be expressed leading to high transgene expression level. In some instances copy number does not correlate to expression transgene expression level.

[0031] The transgene may include a nucleotide sequence from the same species, a nucleotide sequence from a different species, or a nucleotide sequence from each of more than one species a combination of species, a synthetic nucleotide sequence not present in any species’ genomes, or any combination of two or more of the foregoing. A transgene may have some exons or other nucleotide sequences deleted relative to a genome, it may have exons or other sequences added relative to a genome, it may be modified by inserting or deleting nucleotides relative to a genome, or may include any combination of the foregoing. Transgenes may include hybrid genes in which the control elements of the gene of interest are used to direct the expression of a reporter gene. Transgenes may include one or more genes, alleles or both. Transgenes may include one or more reporter genes. Transgenes may include one or more enzymes. Transgenes may include one or more receptors. Transgenes may include endogenous genes. Transgenes may include host genome sequences, host genes, synthetic sequences, synthetic genes and combinations thereof. In some examples one or more transgene is integrated into a mouse genome sequence and the transgene includes human sequences and mouse sequences.

[0032] Transgenic mice may be used to express human receptors to make them susceptible to human viruses, such as hepatitis viruses, papillomavirus, poliovirus, human immunodeficiency virus-1 (HIV-1) and measles. Transgenic mice that express human receptors specific for such viruses have rendered those transgenic mice susceptible to infection by human viruses of interest and subsequently enabled to investigate their pathogenesis in in vivo models.

[0033] To fully benefit from the transgenic model organism, it is useful to understand where the transgene has integrated in the host genome by mapping transgene integration in the transgenic organism genome. Sequencing data may be used to map transgene integration sites. Sequence data in the form of sequence reads may be used to map the regions where a transgene contacts the host genome at these junctions can be used to create a consensus sequence at the transgene junctions confirming each chromosomal location a transgene has integrated. In someAttorney Docket No.: 0303.052AWO embodiments a consensus sequence is a sequence of nucleotides or amino acids that is most common in a set of related sequences. A consensus sequence may be the most common and / or most prevalent nucleotide sequence for a particular location in the custom genome. A consensus sequence may be the nucleotide sequence for a particular location with the lowest error probability. A consensus sequence (or canonical sequence) may be the calculated sequence of the most frequent residues, either nucleotide or amino acid, found at each position in a sequence alignment. A consensus sequence may represent multiple sequence alignments.

[0034] A variety of sequencing methods may produce sequence reads. In some embodiments sequencing data refers to sequence reads. In some embodiments sequence reads are from the whole genome of the transgenic organism. In some embodiments, the sequencing data is DNA sequencing data. In some embodiments, the sequencing data is RNA sequencing data.

[0035] Many sequencing methods rely on reference genomes to assemble sequencing reads into longer sequences that represent the genome of the sequenced organism. Many sequencing technologies will perform a sequence alignment to a reference genome. The reference genome may be user selected. The reference genome may be user selected from an online database of reference genomes. The reference genome may be downloaded into computer memory. The reference genome may include human, mouse, rat, rodent, pig, chicken, cat, dog, monkey, horse, sheep, bacteria, virus, plants, yeast and fungi. In some embodiments a reference genome sequence is a genome sequence from a specific strain of the host sequence. In some embodiments a reference genome sequence is a genome sequence from a specific strain of mouse. In some embodiments a reference genome sequence includes genome sequence of mouse strain C57BL / 6J, B10.RIII, NOD / SCID, BALB / c Nude, CAST / EiJ, NZO / HILtJ, WSB / EiJ, DBA / 2J, C3H / HeOuJ, BALB / cJ, 129S1 / SvImJ , BALB / cByJ, A / J, CC032, or NOD / ShiLtJ. The reference genome may be a digital sequence of a reference sequence that can be contained in a database of reference sequences. A database of reference sequences can include a plurality of genomes and / or proteomes, sequenced with a plurality of methods. A reference genome can be a haploid, a diploid, or a polyploid genome. In some embodiments, the reference genome is an annotated genome. In some embodiments, a system, a method, an algorithm, and / or a computer program product can map one or more transgene integration in a transgenic organism. As described herein a reference genome may include a host genome sequence. A host organism is an organism that may have a transgene integrated into its genome. The genome sequences of manyAttorney Docket No.: 0303.052AWO host organisms are available on databases in digital form and are termed reference genomes. Since reference genomes do not include transgene sequences, sequence reads corresponding to transgene sequence may be interpreted as mismatches or simply not aligned to the reference genome. As referred to herein, s custom genome is a sequence of nucleotides stored on one or more computer memory device that includes a host genome sequence, as a reference genome sequence, and a transgene sequence added thereto.

[0036] Nucleic acids may be sequenced using a variety of methods. Many commercial nucleic acid sequencing instruments are available and produce nucleic acid sequencing data that may be used with the methods, products and systems disclosed herein. In some embodiments Illumina sequencing data will be provided. In some embodiments Oxford Nanopore sequencing data will be provided. In some embodiments Pacific Bio sequencing may be provided. In some embodiments sequence reads will be provided. The sequence read length may depend on the method of sequencing. Some technologies produce short read lengths. Some sequencing technologies produce long read lengths.

[0037] Sequencing technologies may be distinguished between short and long-read sequencing. Next-generation sequencing (NGS) read length refers to the number of base pairs (bp) sequenced from a DNA fragment. In some embodiments after sequencing, the regions of overlap between sequence reads are used to assemble and align the sequence reads to a reference genome, reconstructing the full DNA sequence. Coverage depth refers to the average number of sequencing reads that align to, or cover, each base in a sequenced sample.

[0038] Short-read protocols generate reads of < 300 base pairs (bp), whereas long-read sequencing can provide uninterrupted reads ranging from 10 kbp to several megabases depending on the technology. Long-read sequencing may improve the sequence phasing and be the preferred method for solving larger haplotypes and detection of complex structural variants and repeats. Short reads may be used for counting the abundance of specific reads and expression analysis. Short-read whole genome sequencing (WGS) protocols routinely provide 10 times (10X) coverage of more than 95% of a genome and a median coverage of 30X in a single analysis, and this is generally considered sufficient for germline analysis. WGS is normally performed as paired-end sequencing, which enables more accurate read alignment and detection of structural rearrangements. In some embodiments sequence reads of long-read sequencing may be used for mapping one or more transgene integration into a transgenic organism. In someAttorney Docket No.: 0303.052AWO embodiments sequence reads of short-read sequencing may be used for mapping one or more transgene integration into a transgenic organism. In some embodiments WGS sequence reads are use for mapping one or more transgene integration in a transgenic organism. In some embodiments sequence reads refer to WGS sequence reads. In some embodiments a minimum read length may be used for mapping one or more transgene integration sites into a transgenic organism. In some embodiments a sequence read length is about 300 base pairs or greater, about 500 base pairs or greater, about 750 base pairs or greater, about 1000 base pairs or greater, about 2000 base pairs or greater, about 3000 base pairs or greater, about 4000 base pairs or greater, about 5000 base pairs or greater or about 10,000 base pairs or greater.

[0039] In some embodiments sequence reads are short-read WGS, which in non-limiting examples may be generated using Illumina technology, yielding paired-end ~150 bp reads. In some embodiments sequence reads are long-read WGS, which in non-limiting examples may be generated using single molecule technologies from Pacific Biosciences (PacBio) or Oxford Nanopore Technologies (ONT), yielding 10-100 kb reads or longer. In some embodiments sequence reads are linked-read WGS, which in non-limiting examples may be generated using technology from 10X Genomics, which generates barcoded Illumina short-reads from longer molecules (e.g., ~50 kb). In some embodiments deeper sequence coverage improves variant detection sensitivity, and also improves accuracy by allowing for more sophisticated filtering schemes. In some embodiments deeper coverage (>30x) is needed for mapping one or more transgene integration into a transgenic organism. In some embodiments lower coverage (>20x) is needed for mapping one or more transgene integration into a transgenic organism. In some embodiments sequence reads with about 0.5 x, about 1x, about 1.5 x, about 2x, about 3x, about 4x, 5x, about 10x, about 15x, about 20x, about 25x, about 30x, about 35x, about 40x, about 45x, about 50x, about 55x, about 60x, about 65x, about 70x, about 75x, about 80x, about 85x, about 90x, about 95x or about 100x coverage are used for mapping one or more transgene into a transgenic organism.

[0040] A custom genome may include a reference genome sequence or host genome sequence plus one or more transgene sequences. In some embodiments a transgene includes an exogenous nucleic acid sequence. In some embodiments a transgene includes an engineered nucleic acid sequence. In some embodiments a transgene includes one or more copies of an endogenous nucleic acid sequence. In some embodiments a transgene includes host genomeAttorney Docket No.: 0303.052AWO sequences. In some embodiments a providing a custom genome includes combining one or more transgene sequences with a reference genome sequence in silico. In some embodiments providing a custom genome includes creating an artificial chromosome from the transgene sequence in silico. In some embodiments creating an artificial chromosome includes assembling into a continuous sequence each one or more transgene sequence in silico. An artificial chromosome sequence, generated in silico, may be used for any alignments as disclosed herein, and referred to herein in general as an artificial chromosome. Artificial chromosomes may be used to align sequence reads. Artificial chromosomes may be used to identify candidate junction reads. In some embodiments a custom genome includes appending one or more transgene sequences to one or more chromosomes in silico. In some embodiments a custom genome includes creating a custom genome in-silico comprising one or more transgene sequences and a reference genome sequence. In some embodiments a custom genome may include an in-silico contiguous sequence of the reference genome sequence and the transgene sequence. In some embodiments creating an artificial chromosome in silico includes saving the genome sequence of the artificial chromosome in computer memory. In some embodiments an in-silico genome is stored in computer memory wherein at least one computer processor aligns sequence reads to the custom genome, identifies candidate junction reads, determines a transgene junction, maps the integration of the transgene into the transgenic organism based on the transgene junctions or any combination thereof.

[0041] In some embodiments a host organism includes a mouse and a mouse reference genome may be used with one or more transgene sequence to create a custom genome. In some embodiments a custom genome includes addition of the transgene sequence as an additional chromosome to the reference genome sequence. A transgene may include one or more gene sequences, untranslated regions (UTRs) and regulatory sequence. A transgene may include one or more poly adenylation signals, exon-intron splicing motifs, start codons, stop codons, promoter and vector sequence. In some embodiments a custom genome includes a reference genome sequence to a mouse and a transgene sequence derived from mostly human genome sequences. In some embodiments a transgene sequence may include sequence related to a vector. In some embodiments a reference genome sequence is to a first species and the transgene sequence is from a second species. In some embodiments a transgene is derived from one or more species. In some embodiments a transgene sequence is derived from two or more species.Attorney Docket No.: 0303.052AWO In some embodiments a transgene sequence is derived from three or more species. In some embodiments a transgene sequence is derived from four or more species.

[0042] As disclosed herein mapping one or more transgene integration site into a transgenic organism may include aligning sequence reads of the transgenic organism to the custom genome. In some embodiments sequence reads and reads will be used interchangeably and refer to nucleotide sequences derived from sequencing data. In some embodiments the raw sequence data is processed into a string of nucleotide sequences called a read or sequence read. In some embodiments sequence reads are obtained from one or more sequencing experiment. In some embodiments a sequence read is a nucleic acid sequence of a transgenic organism. In some embodiments a read is a sequence of a transgenic organism wherein the sequence is derived from whole genome sequencing techniques. In some embodiments sequence reads may number from the hundreds of thousands to tens of millions of nucleotides and be of various lengths depending on the sequencing technology used. In some embodiments a sequence read is stored in a FASTQ file. FASTQ is a text-based sequencing data file format that stores both raw sequence data and quality scores.

[0043] Sequence read alignments may include the alignment of a plurality of sequence reads to a reference genome sequence, a custom genome sequence, a transgene sequence, or any combination thereof. In some embodiments a computer system comprising memory and at least one processor configured to execute program instructions will perform alignments. Sequence read alignments may include the aligning of nucleotide or amino acid sequences in a sequence read to a custom genome, transgene sequence, reference genome sequence or any combination thereof. Sequence read alignments may include the aligning of matching nucleotide or amino acid sequences in a sequence read to a reference genome sequence of a custom genome. In some embodiments sequence reads will first be aligned to a reference genome sequence as part of the nucleic acid sequencing protocols. In other embodiments sequence reads will not be aligned to a reference genome sequence as part of the nucleic acid sequencing protocol. In some embodiments a first alignment will be performed on the sequence reads to the reference genome sequence and a second alignment will be performed on the sequence reads to a custom genome. In some embodiments aligned sequence reads are maintained in Binary Alignment Map (BAM) format. In some embodiments sequence read alignments to a reference genome are stored in one or more SAM files, BAM files, CRAM files, Pairwise Mapping Format (PAF) files andAttorney Docket No.: 0303.052AWO combinations thereof. In some embodiments a plurality of sequence reads are stored in computer memory wherein an alignment algorithms is executed using a processor and memory to align each sequence read of the plurality of the sequence reads to a reference genome sequence, custom genome or transgene sequence. In some embodiments computer memory store sequence reads that align to only the reference genome sequence of the custom genome, sequence reads that align only to transgene sequence of the custom genome and read sequences that align to both the reference genome sequence and the transgene sequence of the custom genome. Sequence alignment algorithms are discussed in Heng Li, Nils Homer, A survey of sequence alignment algorithms for next-generation sequencing, Briefings in Bioinformatics, Volume 11, Issue 5, September 2010, Pages 473–483 and J. Kim, M. Ji and G. Yi, "A Review on Sequence Alignment Algorithms for Short Reads Based on Next-Generation Sequencing," in IEEE Access, vol.8, pp.189811-189822, 2020 ,both of which are hereby incorporated by reference in their entirety.

[0044] Alignments of sequence reads may include alignments of sequence regions. Sequence regions may include portions of a nucleotides sequence of sequence read. In some embodiments a sequence region includes less than the entire nucleic acid sequence of a read sequence. In some embodiments a sequence region includes the entire nucleic acid sequence of a read sequence. In some embodiments a sequence region of a sequence read aligns to a reference genome sequence of a custom genome. In some embodiments a sequence region of a sequence read aligns to a transgene sequence of a custom genome. In some embodiments a sequence region of a sequence read aligns to a reference genome sequence of a custom genome and a second sequence region of the sequence read aligns to a transgene sequence of the custom genome. In some embodiments a sequence quality score is used to select which nucleotide sequences in a sequence read will be used for an alignment. In some embodiments a sequence region includes nucleic acid sequences of a predetermined quality score. In some embodiments only nucleotide sequences in a sequence read with an error rate of 1 in 10 are used for an alignment. In some embodiments only nucleotide sequences in a sequence read with an error rate of less than 1 in 100 are used for an alignment. In some embodiments only nucleotide sequences in a sequence read with error rates less than 1 in 1000 are used in an alignment. In some embodiments only nucleotide sequences in a sequence read with error rates of less than 1 in 10,000 are used in alignments. In some embodiments a q-score indicates the acceptable error rate for nucleotides in sequence reads. InAttorney Docket No.: 0303.052AWO some embodiments a q-score may be called a quality score or a Phred Q-score. In some embodiments nucleotides in sequence reads with a q-scores of about 10, also expressed as Q10, may be used in alignments. In some embodiments nucleotides in sequence reads with a quality score of about 20, also expressed a Q20 may be used in alignments. In some embodiments nucleotides in sequence reads with a quality score of about 30, also expressed as Q30 may be used in alignments.

[0045] Mapping one or more transgene integration sites into a transgenic organism may include locating the junctions of the transgene sequence and host genome sequence. Transgene junctions may include where a host genome sequence contacts a transgene sequence. Transgene junctions may include the location where the transgene sequence begins and the host genome sequence ends. Transgene junctions may include the location where the transgene sequence ends and the host genome sequence begins. Transgene junctions may be useful in mapping the locations where one or more transgene has integrated into one or more host chromosome by locating where the host genome sequence borders the transgene sequence. Transgene junctions may include a collection of sequence reads that align to a transgene junction and the sequence reads may be used to determine a transgene junction sequence.

[0046] In some embodiments determining a transgene junction requires identifying candidate junction reads. Candidate junction reads may include all sequence reads where at least a sequence region of the sequence read aligns to the reference genome sequence of the custom genome and at least a sequence region of the sequence read aligns to the transgene sequence of the custom genome. In some embodiments candidate junctions reads include a sequence read with both a sequence region that aligns to the reference genome sequence of a custom genome and a sequence region that aligns to a transgene sequence of a custom genome. In some embodiments candidate junction reads include sequence reads that include a sequence region that aligns with a transgene sequence of a custom genome and a sequence region that aligns to the reference genome sequence of a custom genome. In some embodiments candidate junction reads include a sequence region that aligns to a reference genome sequence of a custom genome and a sequence region that aligns to a transgene sequence of a custom genome and the two sequence regions are contiguous in the sequence read. In some embodiments candidate junction reads include a sequence region that aligns to a reference genome sequence of a custom genome and a sequence region that aligns to a transgene sequence of a custom genome and the two sequenceAttorney Docket No.: 0303.052AWO regions are non-contiguous in the sequence read. In some embodiments a candidate junction read includes a sequence read with a sequence region that aligns with the reference genome sequence of the custom genome and a sequence region that aligns with the transgene sequence of the custom genome wherein the two sequence regions are in contact in the sequence read. In some embodiments a computer system comprising memory and at least one processor configured to execute program instructions identify candidate junction reads.

[0047] Candidate junction reads may be determined by Pairwise Mapping Format (PAF) files. A PAF may include a text file format describing the approximate mapping positions between two sequences. A PAF may describe approximate mapping positions between a sequence read and a custom genome, a reference genome sequence, a transgene sequence and combinations thereof. A mapping position may include where a sequence read aligns to a custom genome. A mapping position may include where a sequence read aligns to a reference genome sequence. A mapping position may include where a sequence read aligns to a transgene sequence. A mapping position may include where a sequence read aligns to a reference genome sequence and where a sequence read aligns to a transgene sequence in a custom genome. A mapping position may include a gene location, a physical location on a chromosome, a sequence number position on chromosome and combinations thereof. A PAF may include a TAB- delimited file with each line consisting of predefined fields. A PAF generated from an alignment may indicate the number of sequence matches and a total alignment length. In some embodiments the number of sequence matches in an alignment may be found in column 10 of the PAF. In some embodiments candidate junction reads may be identified where a PAF indicates a sequence region of a sequence read aligns to a custom genome and a sequence region of a sequence read aligns to a transgene sequence. Candidate junction reads may be identified by identifying all PAFs that include a transgene sequence alignment and then selecting only the PAFs that also include a reference genome sequence alignment resulting in plurality of PAFs that include a sequence alignment to both a transgene sequence and a reference genome sequence alignment. Since a PAF describes an alignment between two sequences a PAF may include one or more alignment properties as described herein, such as minimum or maximum number of nucleotides aligned, number of nucleotides that mismatch, percent pairwise homology, percent identity, gap size between alignment and any combinations thereof.Attorney Docket No.: 0303.052AWO

[0048] Mapping one or more transgene integration into a transgenic organism may require filtering on the number of nucleotides in a candidate junction read that align to the transgene sequence of the custom genome, the reference genome sequence of the custom genome or both. In some embodiments filtering may include one or more selection criteria. In some embodiments a candidate junction read includes a minimum number of nucleotides that aligns to the reference genome sequence of the custom genome and a minimum number of nucleotides in the sequence read that align to the transgene sequence of the custom genome. In some embodiments, identification of a candidate junction read requires at least about 20, about 30, about 40, about 50, about 60, about 70, about 80, about 90, about 100, about 110, about 120, about 130 about 140, about 150, about 160, about 170, about 180, about 190, about 200, about 250 or about 300 nucleotides to align to the transgene sequence of the custom genome, the reference genome sequence of the custom genome or both. In some embodiments, a sequence region of a candidate junction read requires at least about 20, about 30 about 40, about 50, about 60, about 70, about 80, about 90, about 100, about 110, about 120, about 130 about 140, about 150, about 160, about 170, about 180, about 190, about 200, about 250 or about 300 nucleotides to align to the transgene sequence of the custom genome, the reference genome sequence of the custom genome or both. In some embodiments the minimum number of nucleotides for an alignment are in a sequence region. In some embodiments the minimum number of nucleotides that must align to the transgene sequence of the custom genome is different than the minimum number of nucleotides that must align in a reference genome sequence of the custom genome in a candidate junction read. In some embodiments the minimum number of nucleotides that must align to the transgene sequence of the custom genome is equal to the minimum number of nucleotides that must align in a refer genome sequence of the custom genome in a candidate junction read. In some embodiments a computer system comprising memory and at least one processor configured to execute program instructions determines the minimum number of nucleotides in a sequence read that align to the transgene sequence of the custom genome and the minimum number of nucleotides that align to the reference genome sequence of the custom genome.

[0049] Mapping one or more transgene integration into a transgenic organism may include categorizing two or more candidate junction reads into junction clusters based on a predetermined distance in base pairs of an alignment segment terminus of a first candidate junction read to an alignment segment terminus of a second candidate junction read. AAttorney Docket No.: 0303.052AWO predetermined distance in base pairs used to measure the distance apart of two sequence alignments may be called a cluster proximity region. A cluster proximity region may be measured in nucleotides or base pairs. A cluster proximity region may be determined from one alignment segment terminus to another alignment segment terminus.

[0050] An alignment segment terminus may include at least one end of a sequence read alignment. An alignment segment terminus may include at least one end of a candidate junction read alignment. An alignment segment terminus may include at least one end of an alignment of a sequence read to a reference genome sequence, host genome sequence, transgene sequence or custom genome sequence. An alignment segment terminus may include at least one end of an alignment of a candidate junction read to a reference genome sequence, host genome sequence, transgene sequence or custom genome sequence. In some embodiments an alignment segment terminus may include at least one end of an alignment to the reference genome sequence of the custom genome. In some embodiments an alignment segment terminus may include at least one end of an alignment to the transgene sequence of the custom genome. In some embodiments an alignment of a sequence read includes an alignment segment terminus at the 5 prime end of the sequence alignment and an alignment segment terminus at the 3 prime end of the sequence alignment. In some embodiments an alignment of a candidate junction read includes an alignment segment terminus at the 5 prime end of the sequence alignment and an alignment segment terminus at the 3 prime end of the sequence alignment. In some embodiments an alignment segment terminus corresponds to the 5 prime end of a sequence region alignment. In some embodiments an alignment segment terminus corresponds to the 3 prime end of a sequence region alignment. In some embodiments alignment segment termini may refer to a 5 prime and 3 prime alignment segment terminus. In some embodiments alignment segment termini refer to all the alignment segment terminus in an alignment. In some embodiments alignment segment termini may be used for clustering sequence reads alignments.

[0051] In some embodiments a predetermined distance in base pairs between an alignment segment terminus of a first candidate junction read and an alignment segment terminus of a second candidate junction read may be termed a cluster proximity region. A cluster proximity region may be a predetermined number of nucleotides in a reference genome sequence of a custom genome used to determine if candidate junction reads aligned to a reference genome sequence of a custom genome are close enough in proximity to be grouped or clustered into aAttorney Docket No.: 0303.052AWO junction cluster. A cluster proximity region may be a predetermined number of base pairs on a reference genome sequence used to categorize alignment segment termini of a plurality of candidate junction reads aligned to the reference genome sequence of the custom genome into junction clusters if candidate junction reads include alignment segment termini within a cluster proximity region. In some embodiments an algorithm determines a cluster proximity region to categorize candidate junction reads into junction clusters if at least one alignment segment terminus of a candidate junction read is equal or less than a predetermined number of base pairs from at least a second alignment segment terminus of a second candidate junction read. In some embodiments a cluster proximity region may be determined in a plurality of PAF files to categorize candidate junction reads into junction clusters. A cluster proximity region may be determined in a plurality of PAF, BAM, SAM, CRAM and any combination thereof.

[0052] In some embodiments at least one computer processor executing instructions in computer memory determines if an alignment segment terminus of a candidate junction read is within a cluster proximity region to the alignment segment termini of one or more candidate junction reads and if the alignment segment termini are within a cluster proximity region they are categorized as cluster junctions and saved into computer memory otherwise they are not categorized as cluster junctions. In some embodiments an algorithm stored in computer memory is executed by at least one computer processor to determine if the alignment segment terminus of a candidate junction read is within a cluster proximity region to the alignment segment termini of a plurality of candidate junction reads. In some embodiments an algorithm measures the distance from the alignment segment terminus of each candidate junction read to the alignment segment termini of a plurality of candidate junction reads and if any two or more alignment segment termini are within a cluster proximity region they are categorized into junction clusters otherwise they are not categorized as junction clusters. A clustering algorithm may measure one or more distances between alignment segment termini of two or more candidate junctions reads by determining a distance in nucleotides from the alignment segment termini of the two or more candidate junction reads as aligned to a reference genome sequence of a custom genome and categorizing each candidate junction read into a junction cluster if the distance between their segment termini is within a cluster proximity region and otherwise not categorize each candidate junction read in a junction cluster. A clustering algorithm may measure one or more distances between alignment segment termini of a plurality of candidate junctions reads and categorizeAttorney Docket No.: 0303.052AWO each candidate junction read in junction clusters when the distance between the segment termini of the candidate junction reads is within a predetermined cluster proximity region.

[0053] A clustering algorithm may measure one or more distances between alignment segment termini of candidate junctions reads where the alignment segment termini correspond to the sequence region of the candidate junction reads that align to the transgene sequence of the custom genome. A clustering algorithm may measure one or more distances between alignment segment termini of candidate junctions reads where the alignment segment termini correspond to the sequence region of the candidate junction reads that align to the reference genome sequence of the custom genome. A clustering algorithm may measure one or more distances between alignment segment termini of a plurality of candidate junction reads and cluster the candidate junctions reads that have alignment segment termini within a predetermined cluster proximity region defined in nucleotides or base pairs. If the distance between alignment segment termini of a plurality of candidate junction reads is within a cluster proximity region they may be clustered or categorized into the same junction cluster. If the distance between segment alignment termini of a plurality of candidate junction reads is not within a predetermined distance they may not be clustered or categorized together into the same junction cluster.

[0054] In some embodiments candidate junction reads with alignment segment termini within a cluster proximity region of about 50 base pairs or less may be categorized as a junction cluster and those candidate junction reads outside a proximity region of 50 base pairs or less are not categorized as a junction cluster. In some embodiments a first candidate junction read is categorized into a junction cluster with a second candidate junction read if an alignment segment terminus of the first candidate junction read is in a cluster proximity region of 50 base pair or less from the alignment segment terminus of the second candidate junction read. In some embodiments a predetermined number of base pairs for categorizing a plurality of candidate junction reads into junction clusters is about 5 base pairs, about 10 base pairs, about 15 base pairs, about 20 base pairs, about 25 base pairs, about 30 base pairs, about 35 base pairs, about 40 base pairs, about 45 base pairs, about 50 base pairs, about 60 base pairs, about 70 base pairs, about 80 base pairs, about 90 base pairs, about 100 base pairs, about 110 base pairs, about 115 base pairs, about 120 base pairs, about 125 base pairs, about 130 base pairs, about 135 base pairs, about 140 base pairs, about 145 base pairs, about 150 base pairs, about 200 base pairs,Attorney Docket No.: 0303.052AWO about 250 base pairs, about 300 base pairs, about 350 base pairs, about 400 base pairs, about 450 base pairs, about 500 base pairs, about 1000 base pairs, or about 2000 base pairs.

[0055] In some embodiments a cluster proximity region is about 5 base pairs, about 10 base pairs, about 15 base pairs, about 20 base pairs, about 25 base pairs, about 30 base pairs, about 35 base pairs, about 40 base pairs, about 45 base pairs, about 50 base pairs, about 60 base pairs, about 70 base pairs, about 80 base pairs, about 90 base pairs, about 100 base pairs, about 110 base pairs, about 115 base pairs, about 120 base pairs, about 125 base pairs, about 130 base pairs, about 135 base pairs, about 140 base pairs, about 145 base pairs, about 150 base pairs, about 200 base pairs, about 250 base pairs, about 300 base pairs, about 350 base pairs, about 400 base pairs, about 450 base pairs, about 500 base pairs, about 1000 base pairs, or about 2000 base pairs.

[0056] One or more set of instructions in computer memory may be executed by at least one computer processor to categorize candidate junction reads in junction clusters by measuring a distance in base pairs between an alignment segment terminus of a first candidate junction read aligned to a custom genome to an alignment segment terminus of a second candidate junction read aligned to a custom genome and if the distance is equal to or less than cluster proximity region the candidate junction reads will be categorized into junction clusters with each other and if not they will not be categorized into junction clusters with each other. A clustering algorithm may measure a plurality of distances between a plurality of alignment segment termini of a plurality of candidate junctions reads aligned to a reference genome sequence of a custom genome and categorize two or more candidate junction reads into a junction cluster if their alignment segment termini are within a cluster proximity region. In some embodiments one or more set of instructions in computer memory may be executed by at least one computer processor to categorize candidate junction reads into a first junction cluster, a second junction cluster, a third junction cluster, a fourth junction cluster, a fifth junction cluster, a sixth junction cluster, a seventh junction cluster, an eight junction cluster, a ninth junction cluster a tenth junction cluster, or any number of junction clusters. In some embodiments a first junction cluster may be mapped to a region of the reference genome sequence that is 5 prime in relation to the transgene integration site and a second junction cluster may be mapped to a region of the reference genome sequence that is 3 prime in relation to the transgene integration site. In some embodiments one or more set of instructions in computer memory may be executed by at leastAttorney Docket No.: 0303.052AWO one computer processor to categorize candidate junction reads into junction clusters based on a cluster proximity distance from one or more PAF, BAM, SAM, CRAM and any combination thereof. In some embodiments base pairs may be substituted with nucleotides. In some embodiments base pairs may include bases. In some embodiments sequences may include single stranded sequences. In some embodiments sequences may include double stranded sequences. In some embodiments sequences include DNA sequences. In some embodiments sequences include RNA sequences. In some embodiments a single strand of a double stranded sequence may be used for sequence alignments. In some embodiments one or more algorithm categorized candidate junction reads into junction clusters using a cluster proximity region.

[0057] In some embodiments a computer system comprising memory and at least one processor configured to execute program instructions determines a distance between each of a plurality of alignment segment termini of a plurality of candidate junction reads to categorize all candidate junction reads that have alignment segment termini that are within the cluster proximity region into junction clusters. In some embodiments each candidate junction read with alignment segment termini within a cluster proximity region of another candidate junction read alignment segment termini are clustered into the same junction cluster. In some embodiments a plurality of cluster proximity regions are determined by an algorithm for a plurality of alignment segment termini corresponding to a plurality of candidate junction reads. In some embodiments an algorithm categorizes each candidate junction read of a plurality candidate junction reads aligned to the reference genome sequence of the custom genome into junction clusters for any candidate junction reads that include an alignment segment termini within a cluster proximity region of another candidate junction read otherwise they are not categorized as junction clusters.

[0058] In some embodiments a computer system comprising memory and at least one processor configured to execute program instructions categorize candidate junction reads into junction clusters based on a cluster proximity region for all candidate junction reads that include a sequence region that aligns to the same chromosome of the custom genome. In some embodiments all candidate junction reads will be first aligned to the reference genome sequence of the custom genome and then clustered into cluster junctions based on a cluster proximity region.

[0059] In some embodiments a computer system comprising memory and at least one processor configured to execute program instructions clusters sequence reads or candidateAttorney Docket No.: 0303.052AWO junction reads that align to the same chromosome of the custom genome. A cluster may include a group, collection or selection of sequence reads or candidate junction reads stored in computer memory. A chromosome cluster may include all the sequence reads or candidate junction reads that are grouped together as they align to the same chromosome as each other. In some embodiments chromosome clustering includes clustering all sequence reads that align to the same chromosome of the reference genome sequence of the custom genome as each other. In some embodiments chromosome clustering includes clustering all candidate junction reads that align to the same chromosome of the reference genome sequence of the custom genome. In some embodiments sequence reads, sequence regions or candidate junction reads that align to chromosome 1 are clustered together. In some embodiments sequence reads, sequence regions or candidate junction reads that align to chromosome 2 are clustered together. In some embodiments, sequence reads, sequence regions or candidate junction reads that align to chromosome 3 are clustered together. In some embodiments, sequence reads, sequence regions or candidate junction reads that align to chromosome 4 are clustered together. In some embodiments, sequence reads, sequence regions or candidate junction reads that align to chromosome 5 are clustered together. In some embodiments, sequence reads, sequence regions or candidate junction reads that align to chromosome 6 are clustered together. In some embodiments, sequence reads, sequence regions or candidate junction reads that align to chromosome 7 are clustered together. In some embodiments, sequence reads, sequence regions or candidate junction reads that align to chromosome 8 are clustered together. In some embodiments, sequence reads, sequence regions or candidate junction reads that align to chromosome 9 are clustered together. In some embodiments, sequence reads, sequence regions or candidate junction reads that align to chromosome 10 are clustered together. In some embodiments, sequence reads, sequence regions or candidate junction reads that align to chromosome 11 are clustered together. In some embodiments, sequence reads, sequence regions or candidate junction reads that align to chromosome 12 are clustered together. In some embodiments, sequence reads, sequence regions or candidate junction reads that align to chromosome 13 are clustered together. In some embodiments, sequence reads, sequence regions or candidate junction reads that align to chromosome 14 are clustered together. In some embodiments, sequence reads, sequence regions or candidate junction reads that align to chromosome 15 are clustered together. In some embodiments, sequence reads, sequence regionsAttorney Docket No.: 0303.052AWO or candidate junction reads that align to chromosome 16 are clustered together. In some embodiments, sequence reads, sequence regions or candidate junction reads that align to chromosome 17 are clustered together. In some embodiments, sequence reads, sequence regions or candidate junction reads that align to chromosome 18 are clustered together. In some embodiments, sequence reads, sequence regions or candidate junction reads that align to chromosome 19 are clustered together. In some embodiments, sequence reads, sequence regions or candidate junction reads that align to chromosome X are clustered together. In some embodiments, sequence reads, sequence regions or candidate junction reads that align to chromosome Y are clustered together. In some embodiments sequences read or sequence regions that align to a transgene sequence are clustered together. In some embodiments sequence reads, sequence regions or candidate junction reads that align to the same chromosome and the transgene sequence are clustered together. In some embodiments candidate junction reads are chromosome clustered wherein a cluster proximity region for each candidate junction read of a chromosome cluster determines junction clusters for each chromosome cluster.

[0060] In some embodiments, junction clusters that include at least a predetermined number of candidate junction reads are determined to be transgene junctions. Transgene junctions may be used to map one or more transgene integration sites to the reference genome sequence. Candidate junction reads that make up transgene junctions each align in close proximity, within a cluster proximity region, of the reference genome sequence of the custom genome and using the candidate junction reads from each transgene junction to align together may allow the determination of a consensus sequence of the integration site while a sequence region of the candidate junction reads in the transgene junction may allow mapping of the integration site to the physical integration site of the transgene. The one or more transgene integration sites may be more accurately mapped to the reference genome sequence if a junction cluster has at least a minimum number of candidate junction reads. In a non-limiting example, a junction cluster that includes at least three candidate junction reads will produce an accurate consensus sequence. In a non-limiting example, a junction cluster with 5 candidate junction clusters may allow an even more accurate consensus sequence. A transgene junction including more candidate junction reads may result in fewer errors in a consensus sequence and more result in a consensus sequence with higher quality scores for each nucleotide of the sequence. Sequence reads that are entirely within the transgene sequence may be useful for determining a consensus sequence of the internalAttorney Docket No.: 0303.052AWO transgene sequence while transgene junctions are useful for mapping the location of integration and determining a consensus sequence of the integration site. In some embodiments a computer system comprising memory and at least one processor configured to execute program instructions categorizes candidate junction reads in junction clusters to determine transgene junctions. In some embodiments junction clusters that include about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about18, about 19 or about 20 candidate junction reads are determined to be transgene junctions.

[0061] Determining the transgene junctions may allow mapping the integration site of one or more transgene in the transgenic organism. In some embodiments a computer system comprising memory and at least one processor configured to execute program instructions maps the integration of one or more transgene by aligning transgene junctions to the custom genome. In some embodiments mapping the integration of the transgene in the transgenic organism includes aligning the transgene junctions to the reference genome sequence of the custom genome. In some embodiments mapping the integration of the transgene in the transgenic organism includes aligning the transgene junctions to the transgene sequence of the custom genome. In some embodiments mapping the integration of the transgene into the transgenic organism includes aligning the transgene junctions to the reference genome sequence of the custom genome and creating a consensus sequence of the transgene junction. In some embodiments mapping the integration of the transgene into the transgenic organism includes aligning the transgene junctions to the transgene sequence of the custom genome and creating a consensus sequence of the transgene junction. In some embodiments mapping the integration of the transgene into the transgenic organism includes aligning the transgene junctions to the reference genome sequence of the custom genome and creating a consensus sequence of the transgene junction and the transgene sequence as determined from the aligned transgene junctions.

[0062] Mapping one or more transgene integration sites into a transgenic organism may include additional alignment steps. In some embodiments after the transgene junctions are mapped to the reference genome sequence of the custom genome the sequence reads that only align to the transgene sequence of the custom genome are mapped to determine the transgene consensus sequence on the interior of the transgene junction boundary. In some embodiments after the transgene junctions are mapped to the reference genome sequence of the customAttorney Docket No.: 0303.052AWO genome the sequence reads that align fully or partially to the transgene sequence of the custom genome are selected and used to determine the consensus sequence of the transgene to genome and transgene to transgene boundaries. In some embodiments, pairwise mapping of reads is used to identify common junctions between the transgene and host genome.

[0063] Mapping a transgene may include de novo sequence assembly. In some embodiments de novo sequence assembly includes mapping the transgene junction and mapping sequence reads that only align to the transgene sequence of the custom genome. In some embodiments, de novo sequence assembly includes mapping the transgene junction and selecting sequence reads that only align to the transgene sequence of the custom genome to be used for de novo assembly. In some embodiments reads corresponding to junction clusters are selected and used for de novo assembly. In some embodiment candidate junction reads are used for de novo assembly of the custom genome. A de novo assembly of candidate junction reads may be further annotated to generate visualization of potential transgene integration sites. In some embodiments one or more algorithm may be used for de novo sequence assembly for mapping one or more transgene integration into a transgenic organism. De novo assembly is a method for constructing genomes from a plurality of sequence reads with no a priori knowledge of the correct sequence or order of those sequence reads. Contigs are a set of overlapping oriented reads. A single contig is constructed from two or more overlapping and oriented reads. The reads share a subset or all of their nucleotide base pairs. Scaffolds are a set of joined-oriented contigs. A single scaffold is constructed from two or more joined and oriented contigs. The contigs may have to be reversed to yield a matching orientation. The contigs may be overlapping or non-overlapping. Chromosomes are a set of joined-oriented scaffolds. A single chromosome is constructed from two or more joined and oriented scaffolds. The scaffolds may have to be reversed to yield a matching orientation. The scaffolds may be overlapping or non-overlapping. In some embodiments de novo assembly process includes partially, or fully, overlapping reads are assembled into one or more contigs, sets of overlapping or non-overlapping contigs are joined into one or more scaffolds and sets of overlapping or non-overlapping scaffolds are joined into a single chromosome. De novo assemblers are discussed in Khan AR, Pervez MT, Babar ME, Naveed N, Shoaib M. A Comprehensive Study of De Novo Genome Assemblers: Current Challenges and Future Prospective. Evol Bioinform Online.2018, and Chen, Y., Zhang, Y., Wang, A.Y. et al. Accurate long-read de novo assembly evaluation with Inspector. Genome BiolAttorney Docket No.: 0303.052AWO 22, 312 (2021), each of which is hereby incorporated by reference in its entirety. De novo sequence assembly is discussed in Chen, Y., Nie, F., Xie, SQ. et al. Efficient assembly of nanopore reads via highly accurate and intact error correction. Nat Commun 12, 60 (2021), which is hereby incorporated by reference in its entirety.

[0064] Transgene integration into the genome of the transgenic organism can be complex and produce sequences that are novel to the host organism, to the transgene sequence or both. Mapping a transgene creates a consensus sequence at the transgene junction allowing novel oligonucleotides to be constructed. In some embodiments after mapping the transgene integration in the transgenic organism consensus sequences at the transgene junctions are used as sequences of oligonucleotide primers syntheized for PCR reactions. In some embodiments designing primer or oligonucleotides includes generating primer derived from the transgene integration site to distinguish the introduced transgene from the host genome. In some embodiments oligonucleotides derived from a transgene junction may be used for genotyping.

[0065] In some embodiments mapping a transgene may include determining a copy number of one or more transgenes within one or more integration sites. In some embodiments the copy of number of one or more transgenes is determined based on the average read depth of the transgene compared to the mouse genome. Read depth refers to the average number of sequencing reads that align to, or cover, each base in a sequenced sample. In some embodiments the haploid copy number indicates the copies per chromosome based on the indicated heterozygous genotype. In some embodiments transgene copy number is determined by comparing the read depth of the transgene sequence to the average read depth across the host genome. In some embodiments, copy number is determined by the number of reads associated with internal transgene to transgene junctions. The number of unique transgene junctions are counted and read depth of these junctions is taken as an estimate of the number of copies of each transgene to transgene junction. In some embodiments one greater than the number of internal transgene to transgene junctions are used to estimate the total transgene copy number. Internal junctions are comprised of neighboring alignment segment termini that both align to the transgene sequence. In some embodiments copy number is determined by comparing the read depth of the transgene sequence to the average read depth of the reference genome sequence. In some embodiments copy number is determined by an additional one or more algorithm. Copy number determination is discussed in Singh, A.K., Olsen, M.F., Lavik, L.A.S. et al. DetectingAttorney Docket No.: 0303.052AWO copy number variation in next generation sequencing data from diagnostic gene panels. BMC Med Genomics 14, 214 (2021) which is hereby incorporated by reference in its entirety.

[0066] Mapping transgene integration in transgenic organisms allows for an informed breeding strategy as the integration sites are known and can be compared to integration sites on potential organism crosses. Additionally, oligonucleotides synthesis from the transgene junctions allows rapid confirmation of the presence of the transgene in successive generations. In some embodiments a transgene integration site in a mouse genome can be used to select an appropriate mouse for a breeding program. In some embodiments knowledge of the transgene integration site allows breeding of two or more strains together to generate cohorts of experimental mice. In some embodiments transgene positive mice will be bred together to reach homozygosity wherein the litter may be screened using unique transgene primers to confirm homozygosity. In some embodiments a transgenic positive mouse will be bred to produce a chimeric mouse with both transgenic and normal cells. In some embodiments mapping a transgene integration site will allow for the selection of founder breeders for developing suitable mice strains. In some embodiments mapping one or more transgene integration sites allows determination of a mouse genotype wherein the transgenic mouse genotype is used in a breeding strategy to develop mice strains homozygous or heterozygous for a specific genotype. In some embodiments mapping one or more transgene integration site includes mapping the genomic location, copy number and arrangement of transgene insertions in existing transgenic organisms. In an example, mapping a transgene integration site may include mapping a new transgene added to a transgenic organism. In some embodiments mapping one or more transgene insertion sites includes sequencing and mapping the genomic location, copy number and arrangement of transgene insertions in the germline of transgenic founder mice or rats. In some embodiments mapping one or more transgene insertion sites includes assessing the mosaicism of different transgenic insertions in a given founder’s germline and use this information to predict germline transmission success and outcomes. Transgene mosaicism is a condition where a single individual, such as a transgenic mouse or rat, has two or more cell lines with different genetic compositions. Genetic mosaicism is defined as the presence of two or more cell lineages with different genotypes arising from a single zygote in a single individual. Due to mosaicism, the alleles seen from a tail snip or ear punch of founder mice or rats may not necessarily be the same alleles found in sperm or oocytes that can be transmitted from founder to subsequent generations. In some embodiments mappingAttorney Docket No.: 0303.052AWO one or more transgene insertion sites includes determining the sequence and structural changes in a transgene over generations. In some embodiments mapping one or more transgene insertion sites includes determining sequence and structural changes in a transgene after an intervention such as gene editing therapeutics.

[0067] In some embodiments collection of transgenic mouse nucleic acid for mapping one or more transgene integration sites does not require harvesting or sacrificing the mouse or rat allowing the mouse or rat to be used for further breeding strategies. In some embodiments a mouse or rat tail clipping or biopsy may be used to obtain DNA for mapping one or more transgene integration sites. In some embodiments mouse or rat somatic tissue may be used to obtain DNA for mapping one or more transgene integration sites. In some embodiments mouse or rat sperm may be used to obtain DNA for mapping a transgene integration site. In some embodiments founder mice or rat DNA is used for mapping transgene integration sites. In some embodiments germline mice or rat DNA is used for mapping transgene integration sites. Mapping a transgene integration site is discussed in Peter K Nicholls, Daniel W Bellott, Ting-Jan Cho, Tatyana Pyntikova, David C Page, Locating and Characterizing a Transgene Integration Site by Nanopore Sequencing, G3 Genes|Genomes|Genetics, Volume 9, Issue 5, 1 May 2019, Pages 1481–1486, which is hereby incorporated by reference in its entirety. Mapping a transgene integration site is also discussed in Carol Cain-Hom, Erik Splinter, Max van Min, Marieke Simonis, Monique van de Heijning, Maria Martinez, Vida Asghari, J. Colin Cox, Søren Warming, Efficient mapping of transgene integration sites and local structural changes in Cre transgenic mice using targeted locus amplification, Nucleic Acids Research, Volume 45, Issue 8, 5 May 2017, Page e62, which is hereby incorporated by reference in its entirety.

[0068] One or more sequencing read, candidate junction read or both may be aligned to a reference genome sequence, transgene sequence, custom genome or any combination thereof with a degree of homology. A degree of homology may include a percent homology. A degree of homology may be referred to as a percent pairwise homology. A first sequencing read can be aligned to a second sequencing read with a degree of homology. A sequence region may share a percent homology to a reference genome sequence, a transgene sequence, a custom genome or combinations thereof. The entirety of a sequencing read, candidate junction read or both may share a percent homology to a reference genome sequence, transgene sequence, custom genome or combinations thereof. The statistical significance of an alignment may be determined based onAttorney Docket No.: 0303.052AWO a percent of shared homology between two sequences. A sequencing read, candidate junction read and sequence region may share protein, nucleotide and / or base-pair (“pairwise”) homology to a reference genome sequence, transgene sequence, custom genome or combinations thereof. Homology can be measured over an entire sequence or a sequence region of a sequence. Homology can be measured by a computer-executable algorithm, for example, the Basic Local Alignment Search Tool (BLAST).

[0069] A sequencing read and a reference genome sequence can share about 1% pairwise homology, about 2% pairwise homology, about 3% pairwise homology, about 4% pairwise homology, about 5% pairwise homology, about 6% pairwise homology, about 7% pairwise homology, about 8% pairwise homology, about 9% pairwise homology, about 10% pairwise homology, about 11% pairwise homology, about 12% pairwise homology, about 13% pairwise homology, about 14% pairwise homology, about 15% pairwise homology, about 16% pairwise homology, about 17% pairwise homology, about 18% pairwise homology, about 19% pairwise homology, about 20% pairwise homology, about 21% pairwise homology, about 22% pairwise homology, about 23% pairwise homology, about 24% pairwise homology, about 25% pairwise homology, about 26% pairwise homology, about 27% pairwise homology, about 28% pairwise homology, about 29% pairwise homology, about 30% pairwise homology, about 31% pairwise homology, about 32% pairwise homology, about 33% pairwise homology, about 34% pairwise homology, about 35% pairwise homology, about 36% pairwise homology, about 37% pairwise homology, about 38% pairwise homology, about 39% pairwise homology, about 40% pairwise homology, about 41% pairwise homology, about 41% pairwise homology, about 42% pairwise homology, about 43% pairwise homology, about 44% pairwise homology, about 45% pairwise homology, about 46% pairwise homology, about 47% pairwise homology, about 48% pairwise homology, about 49% pairwise homology, about 50% pairwise homology, about 51% pairwise homology, about 52% pairwise homology, about 53% pairwise homology, about 54% pairwise homology, about 55% pairwise homology, about 56% pairwise homology, about 57% pairwise homology, about 58% pairwise homology, about 59% pairwise homology, about 60% pairwise homology, about 61% pairwise homology, about 62% pairwise homology, about 63% pairwise homology, about 64% pairwise homology, about 65% pairwise homology, about 66% pairwise homology, about 67% pairwise homology, about 68% pairwise homology, about 69% pairwise homology, about 70% pairwise homology, about 71% pairwise homology, about 72% pairwiseAttorney Docket No.: 0303.052AWO homology, about 73% pairwise homology, about 74% pairwise homology, about 75% pairwise homology, about 76% pairwise homology, about 77% pairwise homology, about 78% pairwise homology, about 79% pairwise homology, about 80% pairwise homology, about 81% pairwise homology, about 82% pairwise homology, about 83% pairwise homology, about 84% pairwise homology, about 85% pairwise homology, about 86% pairwise homology, about 87% pairwise homology, about 88% pairwise homology, about 89% pairwise homology, about 90% pairwise homology, about 91% pairwise homology, about 92% pairwise homology, about 93% pairwise homology, about 94% pairwise homology, about 95% pairwise homology, about 96% pairwise homology, about 97% pairwise homology, about 98% pairwise homology, about 99% pairwise homology, or about 100% pairwise homology to a reference sequence or a second sequencing read.

[0070] An algorithm can allow for gaps in an alignment. An alignment gap can be of about 1 nucleotide, about 2 nucleotides, about 3 nucleotides, about 4 nucleotides, about 5 nucleotides, about 6 nucleotides, about 7 nucleotides, about 8 nucleotides, about 9 nucleotides, about 10 nucleotides, about 11 nucleotides, about 12 nucleotides, about 13 nucleotides, about 14 nucleotides, about 15 nucleotides, about 16 nucleotides, about 17 nucleotides, about 18 nucleotides, about 19 nucleotides, about 20 nucleotides, about 21 nucleotides, about 22 nucleotides, about 23 nucleotides, about 24 nucleotides, about 25 nucleotides, about 26 nucleotides, about 27 nucleotides, about 28 nucleotides, about 29 nucleotides, about 30 nucleotides, about 31 nucleotides, about 32 nucleotides, about 33 nucleotides, about 34 nucleotides, about 35 nucleotides, about 36 nucleotides, about 37 nucleotides, about 38 nucleotides, about 39 nucleotides, about 40 nucleotides, about 41 nucleotides, about 42 nucleotides, about 43 nucleotides, about 44 nucleotides, about 45 nucleotides, about 46 nucleotides, about 47 nucleotides, about 48 nucleotides, about 49 nucleotides, or about 50 nucleotides. In some embodiments, an alignment gap can be more than 50 nucleotides.

[0071] Embodiments of mapping one or more transgene into a transgenic organism may include one or more computers, computer implemented method, computer system, computer program product with processors for processing and executing instructions and memory for storing sequence reads, candidate junction reads, a cluster proximity region, pairwise homology of alignments, alignment gaps, minimum alignments, cluster junctions and transgene junctions. A processor and memory may be used to implement embodiments disclosed herein. MultipleAttorney Docket No.: 0303.052AWO threads of execution can be used for parallel processing. In some embodiments, multiple processors or processors with multiple cores can be used, whether in a single computer system, in a cluster, or distributed across systems over a network comprising a plurality of computers, cell phones, and / or personal data assistant devices.

[0072] Processes and methods described herein may be performed singly or collectively by one or more computer systems. A computer system may also be referred to herein as a data processing device / system or computing device / system / node, or simply a computer. A computer system and method may be implemented as one or more of a personal computer system, server computer system, thin client, thick client, hand-held or laptop device, mobile device, multiprocessor system, microprocessor-based system, set top box, programmable consumer electronic, network PC, minicomputer system, mainframe computer system, and / or distributed cloud computing environment that includes any of the above systems or devices, and the like.

[0073] The computer system, computer implemented method and computer program product may include one or more processors or processing units and a memory that includes volatile memory (e.g. Random Access Memory, RAM) and non-volatile memory. Memory may further include removable / non-removable, volatile / non-volatile computer system storage media. Further, memory may include one or more readers for reading from and writing to a non-removable, non- volatile magnetic media, such as a hard drive, a magnetic disk drive for reading from and writing to a removable, non-volatile magnetic disk, and / or an optical disk drive for reading from or writing to a removable, non-volatile optical disk such as a CD-ROM, DVD-ROM. The system may also include a variety of computer readable tangible storage media. Such media may be any available media, such as volatile and non-volatile media, and removable and non-removable media.

[0074] Memory may include at least one program product having a set (e.g., at least one) of program modules implemented as executable instructions that, when executed, carry out functions described herein. Executable instructions may include an operating system, one or more application programs, other program modules, and program data or other types of software. Generally, program modules may include routines, programs, objects, components, logic, data structures, and so on, that perform particular tasks or implement particular abstract data types. Program modules may carry out functions, processes, algorithms, methods and the like described herein including, but not limited to, providing a custom genome, aligning sequence reads of theAttorney Docket No.: 0303.052AWO transgenic organism to the custom genome, identifying candidate junction reads, categorizing the candidate junction reads into junction clusters, alignments, aligning sequence reads, aligning candidate junction reads, determining a transgene junction, mapping the integration of the transgene into the transgenic organism based on the transgene junctions, clustering and determining a cluster proximity region. In some embodiments one or more clustering algorithm, algorithm, or combinations thereof are included in computer program and computer program product for mapping transgene integration into a transgenic organism.

[0075] Components of the computer system may be coupled by an internal bus that may be implemented as one or more of any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of a variety of bus architectures.

[0076] A Computer system may also communicate with one or more external devices such as a keyboard, a pointing device, a display and / or any devices (e.g., network card, modem, etc.) that enable computer system to communicate with one or more other computer systems, such as a server or other system hosted in a cloud computing environment. Such communication can occur via I / O interfaces, which may include a network interface to interface to one or more networks such as a local area network (LAN), a general wide area network (WAN), and / or a public network (e.g., the Internet) via a suitable network adapter.

[0077] As used herein, the term "cloud" or "cloud computing environment" may refer to various evolving arrangements, infrastructure, networks, and the like that will typically be based upon the Internet. The term may refer to any type of cloud, including client clouds, application clouds, platform clouds, infrastructure clouds, server clouds, and so forth. As will be appreciated by those skilled in the art, such arrangements will generally allow for use by owners or users of sequencing devices, provide software as a service (SaaS), provide various aspects of computing platforms as a service (PaaS), provide various network infrastructures as a service (IaaS) and so forth. Moreover, included in this term should be various types and business arrangements for these products and services, including public clouds, community clouds, hybrid clouds, and private clouds. Any or all of these may be serviced by third party entities. However, in certain embodiments, private clouds or hybrid clouds may allow for sharing of sequence data and services among authorized users.Attorney Docket No.: 0303.052AWO

[0078] A cloud facility may include a plurality of computer systems / nodes. The computing resources of the nodes may be pooled to serve multiple consumers, with different physical and virtual resources dynamically assigned and reassigned according to consumer demand. Examples of resources include storage, processing, memory, network bandwidth, and virtual machines. The nodes may communicate with one another to distribute resources, and such communication and management of distribution of resources may be controlled by a cloud management module residing in one or more nodes. The nodes may communicate via any suitable arrangement and protocol. Further, the nodes may include servers associated with one or more providers. For example, certain programs or software platforms may be accessed via a set of nodes provided by the owner of the programs while other nodes are provided by data storage companies. Communication with the cloud facility may include communication via a local area network (LAN), a general wide area network (WAN), and / or a public network (e.g., the Internet) via the communications link.

[0079] Results of a sequencing run and various analyses can be stored in files taking the form of FASTQ files, binary alignment files (bam), Pairwise mApping Format (PAF) *.bcl, *.vcf, *.paf, and / or *.csv files, as examples. The output files may be in formats that are compatible with sequence data viewing, modification, annotation, manipulation, alignment, and realignment software. Accordingly, the accessible sequence alignment dataset provided herein may be in the form of raw data, partially processed or processed data, and / or data files compatible with particular software programs. In this regard, a computer system, such as a computer system of or in communication with a sequencing device, or a cloud facility computer system, as examples, can obtain a bam or other sequencing alignment dataset and process the file by, for instance, reading its data and performing operations to carrying out aspects described herein. The computer system can then output a file having sequencing alignment data, for instance another bam file. Further, the output files may be compatible with other data sharing platforms or third-party software.

[0080] The terms “nucleic acid” or “nucleotide” refer to a deoxyribonucleotide or ribonucleotide polymer, in either single- or double-stranded form, and unless otherwise limited, can encompass known analogs of natural nucleotides that can function in a similar manner as naturally occurring nucleotides.Attorney Docket No.: 0303.052AWO

[0081] All literature and similar materials cited in this application, including but not limited to, patents, patent applications, articles, books, treatises, and internet web pages are expressly incorporated by reference in their entirety for any purpose. Further, unless otherwise required by context, singular terms shall include pluralities and plural terms shall include the singular.

[0082] As used herein, the term “about” means that the numerical value is approximate and small variations would not significantly affect the practice of the disclosed embodiments. Where a numerical limitation is used, unless indicated otherwise by the context, “about” means the numerical value can vary by ±1 or ±10%, or any point therein, and remain within the scope of the disclosed embodiments.

[0083] It should be appreciated that all combinations of the foregoing concepts and additional concepts discussed in greater detail herein (provided such concepts are not mutually inconsistent) are contemplated as being part of the inventive subject matter disclosed herein. In particular, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as being part of the inventive subject matter disclosed herein and may be used to achieve the benefits and advantages described herein. EXAMPLES

[0084] The following examples are intended to illustrate particular embodiments of the present disclosure but are by no means intended to limit the scope thereof.

[0085] Although some non-limiting examples have been depicted and described in detail herein, it will be apparent to those skilled in the relevant art that various modifications, additions, substitutions, and the like can be made without departing from the spirit of the present disclosure and these are therefore considered to be within the scope of the present disclosure as defined in the claims that follow.

[0086] In some embodiments a method for mapping one or more transgene insertion sites into a transgenic organism includes an algorithms that identifies candidate junction reads by aligning a plurality of sequence reads to a custom genome wherein sequence reads that align in part to the reference genome sequence portion of the custom genome and align in part to the transgene sequence portion of the custom genome are identified as candidate junction reads; candidate junction reads are aligned to a reference genome sequence portion of the custom genome and candidate junction reads with alignment segment termini within a predetermined cluster proximity region as defined by a predetermined number of nucleotides are categorizedAttorney Docket No.: 0303.052AWO into junction clusters; junction clusters including a predetermined number of candidate junction reads are determined to be transgene junctions; transgene junctions are then used for mapping the integration of the one or more transgenes into the transgenic organism by determining the location of the transgene sequence to host genome sequence boundaries. Transgene to host genome boundaries may identify the regions where one or more transgene is integrated into the host genome sequence. De novo sequence assembly of transgene junctions may produce a nucleotide sequence or consensus sequence of the transgene sequence to host genome sequence boundaries, transgene integration sites. Creating a consensus sequence using the sequence reads that belong to the transgene junctions may provide a consensus sequence or one or more transgene integration sites and the consensus sequence region that aligns to the reference genome sequence may allow the mapping of one or more transgene integration sites to a specific location in the genome of the transgenic organism.

[0087] FIG.1A illustrates non-limiting embodiments of a method of mapping the integration of one or more transgene integration into a transgenic organism as disclosed herein. In some embodiments a custom genome 150 includes a reference genome sequence 105 wherein sequence reads 110, 115, 120 and 125 align to sequence regions of the custom genome sequence as shown. Sequence reads 110, 115, 120 and 125 are shown to include a sequence region that does not align to the reference genome sequence 105 of the custom genome shown with a checkered pattern. In some embodiments a custom genome 150 includes a reference genome sequence 105 wherein sequence reads 110, 115, 120 and 125 align to a portion of the reference genome sequence 105 in custom genome. In some embodiments only a sequence region of the sequence reads 110, 115, 120 and 125 align to reference genome sequence 105 of the custom genome 150.

[0088] FIG.1B illustrates non-limiting embodiments of a method of mapping the integration of one or more transgene integration into a transgenic organism as disclosed herein. In some examples a custom genome 150 includes a transgene sequence 130. In some embodiments sequence reads 110, 115, 120 and 125 align to the transgene sequence 130 in custom genome 150 as shown with the checkered pattern of the transgene sequence 130 aligned with the sequence reads. In some embodiments a sequence region of the sequence reads 110, 115, 120 and 125 align to a portion of the transgene sequence 130 in custom genome 150.Attorney Docket No.: 0303.052AWO

[0089] In some embodiments sequence reads 110, 115, 120 and 125 that align in part to the reference genome sequence 105 of a custom genome 150 and sequence reads 110, 115, 120 and 125 align to a portion of the transgene sequence 130 of a custom genome 150 are termed candidate junction reads. In some embodiments sequence regions of sequence reads 110, 115, 120 and 125 that align to a portion of the reference genome sequence 105 in custom genome and sequence regions of sequence reads 110, 115, 120 and 125 align to a portion of the transgene sequence 130 in custom genome 150 are termed candidate junction reads.

[0090] FIG.2 illustrates non-limiting embodiments of categorizing candidate junction reads into junction cluster. A method for mapping one or more transgene integration sites in a transgenic organism may include aligning candidate junctions reads 225, 230, 235 and 240 to a reference genome sequence 205 of a custom genome 280. After alignment of candidate junction reads to the reference genome sequence 205 of the custom genome 280 each candidate junction read with an alignment segment terminus within a cluster proximity region of an alignment segment terminus of another candidate junction read are categorized as junction clusters. In an example, candidate junction read 225 includes a 5 prime alignment segment terminus 1 and a 3 prime alignment segment terminus 2 and candidate junction read 230 includes a 5 prime alignment segment terminus 5 and a 3 prime alignment segment terminus 6. A first candidate junction read may be categorized into a junction cluster with a second candidate junction read if the first candidate junction read includes an alignment segment terminus within a predetermined cluster proximity region of an alignment segment terminus of the second candidate junction read. In an example, alignment segment terminus 2 of candidate junction read 225 and alignment segment terminus 6 of candidate junction read 230 are within a predetermined cluster proximity region 245 and are categorized as a junction cluster. The cluster proximity region 245 may be about 10 to about 100 nucleotides as measured between the alignment segment termini as aligned to the reference genome sequence 205 of the custom genome 280. A cluster proximity region 245 may include a predetermined number of base pairs or nucleotides counted as the number of nucleotides on the reference genome sequence 205 between alignment segment termini 2, 6, 10 that are aligned to reference genome sequence 205 of a custom genome 280. In an example, a method of mapping one more transgene integration sites includes an algorithm categorizing candidate junction reads as junction clusters first aligns candidate junction reads to a reference genome sequence 205 of a custom genome 280 and each alignment segment termini 1, 2, 5, 6, 7,Attorney Docket No.: 0303.052AWO 8, 9 and 10 of the candidate junction read alignments that are within a cluster proximity region 245 of any other alignment segment termini are categorized into a cluster junction otherwise they are not categorized into a cluster junction. In an example, candidate junction reads 225, 230, and 240 include alignment segment termini 2, 6 and 10 within a cluster proximity region 245 and will be categorized as a junction cluster while candidate junction read 235 with alignment segment termini 7 and 8 will not be included in the junction cluster as alignment segment termini 7 and 8 are not within a cluster proximity region 245 of another alignment segment termini. Categorizing candidate junction reads may produce a plurality of junction clusters. In some embodiments candidate junctions reads 225, 230, 235 and 240 are aligned to the same chromosome of the reference genome sequence 205 of the custom genome 280 and each candidate junction read that includes an alignment segment terminus within a cluster proximity region 245 is categorized into a junction cluster otherwise they are excluded from the cluster junction.

[0091] After junction clusters are categorized, each junction cluster that contains a predetermined number of candidate junction reads will be determined to be a transgene junction. In an example, junction clusters with at least three candidate junction reads are determined to be transgene junctions. Transgene junctions include a collection of candidate read junctions that may be used for determining a consensus sequence of the transgene integration site of the transgenic organism. The candidate junction reads determined as transgene junction align to both the reference genome sequence and the transgene sequence of the custom genome and using the transgene junctions to assemble a consensus sequence of both the host genome sequence and transgene sequence at the transgene integration site allows a de novo assembly of the important, an often unique, nucleotide sequence at the transgene integration site. The consensus sequence of each transgene integration site may be used to map the location of each transgene integration site by aligning the host genome sequence region of the consensus sequence to the reference genome sequence and mapping the transgene integration site to the genome of the transgenic organism. Once an in-silico sequence assembly of the consensus sequence of the transgene integration site is complete the sequence of the integration site may be used to develop oligonucleotides to further sequence or to develop polymerase chair reaction assays to test for the presence of the specific transgene integration site in mice and rats, such as in breeding programs. Oligonucleotides may be designed complementary to the transgene integration site sequence andAttorney Docket No.: 0303.052AWO may be utilized for quantifying each transgene integration site in mice and rats used in breeding programs. Transgene junction may be used for consensus sequence assembly and gene mapping the location of the one or more transgene integration sites in the transgenic organism.

[0092] FIG.3 illustrates non-limiting embodiments of a method for mapping one or more transgene integrations sites into a transgenic organism. A method for mapping one or more transgene integrations sites may include collecting and saving into computer memory sequence reads 305 from a transgenic organism. The sequence reads 305 may be long read whole genome sequence reads. The sequence reads 305 may require a minimum quality score of Q10 (as a non- limiting example). The sequence reads 305 may be at least 750 base pairs in length. The sequence reads 305 that do not meet the quality score or minimum length may be excluded from further analysis. The sequence reads 305 that pass the filtering criteria may then be aligned to the custom genome 310. The custom genome includes both the reference genome sequence and the transgene sequence.

[0093] After the sequence reads are aligned to the custom genome 305 the candidate junction reads are identified 315. Identifying candidate junction reads 315 may include identifying all sequence reads that include a sequence region that aligns to a portion of the reference genome sequence of the custom genome and a sequence region that aligns to the transgene sequence of the custom genome. If a sequence read includes a sequence region that aligns to a reference genome it is then evaluated for alignment to the transgene sequence; if the sequence read includes a sequence region that aligns to the reference genome sequence of the custom genome and a different sequence region that aligns to a portion of the transgene sequence of the custom genome it will be saved in computer memory as a candidate junction read otherwise it will not be saved in memory as a candidate junction read. In some embodiments sequence reads are first aligned to the transgene sequence of the custom genome and all sequence reads that include a first sequence region that aligns to the transgene sequence of the custom genome are then aligned to the reference genome sequence of the custom genome, this may be more efficient as fewer sequence reads are predicted to include sequence regions that align to the transgene sequence of the custom genome; if the sequence read contains a second sequence region that aligns to the reference genome sequence the read is identified as a candidate junction read 315 and saved in computer memory as a candidate junction read, otherwise it is not saved in computer memory as a candidate junction read. Identifying a candidate junction read 315 may include aligning aAttorney Docket No.: 0303.052AWO sequence read to a transgene sequence of a custom genome and if a minimum of 50 nucleotides (as a non-limiting example) align to the transgene sequence then aligning the sequence read to the reference genome sequence of the custom genome and if a minimum of 50 nucleotides align to the reference genome sequence then identifying the sequence read as a candidate junction read and saving into computer memory as a candidate junction read otherwise if they do not meet this criteria they are not saved in computer memory as candidate junction reads.

[0094] Categorizing candidate junction reads into junction clusters 320 may include defining a cluster proximity region which may be described as a distance in base pairs or nucleotides on the reference genome sequence separating alignment segment termini aligned to the reference genome sequence. A cluster proximity region may be described as a sliding window on the reference genome sequence of a predetermined length in nucleotides such that after candidate junction reads are identified 315 each candidate junction reads is then aligned to the reference genome sequence of the custom genome and each alignment segment terminus that is within a cluster proximity region of another alignment segment terminus of another candidate junction read will result in the two candidate junction reads being categorized into junction clusters 320 otherwise if the alignment segment termini are further in distance than the predetermined cluster proximity region the candidate junction reads will not be categorized as junction clusters.

[0095] Determining a transgene junction 325 may include analysis of the junction clusters and each junction cluster that includes a predetermined number of candidate junction reads will be determined as a transgene junction otherwise a transgene junction will not be determined. In an example, junction clusters with at least three candidate junction reads are determined transgene junctions and saved in computer memory and junction clusters with less than three (as a non-limiting example) candidate junction reads are not determined to be junction clusters and not saved in computer memory.

[0096] Mapping the integration site of the one or more transgenes into the transgenic organism based on the transgene junctions 330 may include analyzing the alignment location in the reference genome sequence for each candidate junction read in a transgene junction to map the transgene junction to the reference genome sequence. In an example, a transgene junction includes ten candidate junction reads that each align to a chromosomal locus and the transgene junction can be mapped to that chromosomal locus and the integration site of the transgene can be confidently predicted as mapping to that chromosomal locus. Mapping the location of eachAttorney Docket No.: 0303.052AWO transgene integration site allows a determination of where in the genome sequence of the transgenic organism the transgene is integrated. Location information can be useful for determining if a transgene integration site disrupts a gene in the transgenic organism, if the transgene is a full sequence, truncated, inverted, duplicated, translocated or any combination thereof. In an example, transgene integration sites might indicate a partial transgene integration in one or more chromosomes, full integration into one or more chromosomes or any combination thereof.

[0097] Assemble transgene junction reads and create consensus sequence 335 may include aligning the candidate junction reads making up each transgene junctions to other candidate junction reads in the same transgene junction to determine a consensus sequence. The candidate junction reads in a transgene junction will typically be in close proximity to each other, based on the selection criteria of the cluster proximity region, and can be used to derive a consensus sequence for each transgene integration site into the transgenic organism. A consensus of the transgene integration site will reveal the location where each transgene is integrated into the transgenic organism and may indicate the orientation of the transgene relative to other genes in the integration site. The consensus sequence at the transgene integration site may be used to develop PCR assays to further sequence the integration site to perform mutation analysis, develop PCR assays to determine a copy number of a transgene as compared to an endogenous gene of the transgenic organism, and develop PCR assays that confirm the presence of the transgene in mouse and rat breeding and genetic crosses to confirm the presence or absence of a specific transgene.

[0098] Evaluating assembly alignments 340 may include further sequencing using sequencing primers designed from the consensus sequences to confirm one or more transgene integration sites and to confirm one or more transgene consensus sequences. Evaluation of the assembly alignments may include confirmation FISH assays, sequencing or PCR assays.

[0099] In some embodiments a method for mapping one or more transgene integration sites includes steps that are repeated one more times, steps may be in any order and steps of the method may include any steps or methods disclosed herein.

[0100] In some embodiments sequencing is conducted using the Oxford Nanopore PromethION prepared with the Ligation Sequencing Kit. Reads are aligned to a mouse genome (GRCm38) containing appended reference transgene sequences, that were provided for eachAttorney Docket No.: 0303.052AWO transgene, using minimap2 to create a custom genome. Transgene coordinates are based on the transgene reference file. Structural variants are called with a long-read variant caller, sniffles2. Read depth visualization may be performed using samplot and copy number variants are called using CNVkit. The sequences of the integration site junctions are prepared by assembling informative reads such as candidate junction reads from their associated structural variant using shasta.

[0101] FIG.4A illustrates a non-limiting embodiment of a visualization depicting sequence coverage surrounding transgene junctions, such as on a screen or monitor or printer to which a computer system performing a method as disclosed herein is performed, for depicting the visualization thereon or thereby. The uppermost graph depicts a non-limiting example of chromosome 1, abbreviated on the X-axis as chr1, and the physical location of the chromosome sequence with sequence coverage on the Y-axis where sequence coverage is determined using whole genome sequencing. In this non-limiting example, the sequence coverage in the central region of chromosome 1 decreases by approximately 50% indicating transgene integration may have deleted one copy of the host genome sequence. The lower graph is a non-limiting example of an illustration of a sequence coverage map of sequence reads that have an alignment that maps to the transgene and also maps to the reference genome sequence. In some embodiments immediate and complete drops in the coverage map of transgene-associated reads may illustrate transgene borders. In some embodiments, vertical faces facing inward may suggest transgene insertion caused a genomic deletion. Illustrations may include genes annotation and gene orientation.

[0102] FIG.4B illustrates a non-limiting embodiment of a visualization depicting sequence coverage surrounding transgene junctions, such as on a screen or monitor or printer to which a computer system performing a method as disclosed herein is performed, for depicting the visualization thereon or thereby. The uppermost graph depicts a non-limiting example of chromosome 16, abbreviated on the X-axis as chr16, and the physical location of the chromosome sequence with sequence coverage on the Y-axis where sequence coverage is determined using whole genome sequencing. In this non-limiting example, the sequence coverage in the central region of chromosome 16 increases by approximately 50% indicating transgene integration may have duplicated one copy of the host genome sequence. The lower graph is a non-limiting illustration of a sequence coverage map of sequence reads that have anAttorney Docket No.: 0303.052AWO alignment that maps to the transgene and also map to the reference genome sequence. In this non-limiting example, the lower graph illustrates that transgene insertion caused a duplication of endogenous chromosomal sequence. In some embodiments immediate and complete drops in the coverage map of transgene-associated reads may illustrate transgene borders. In some embodiments vertical faces facing outward may suggest transgene insertion caused a genomic duplication. Illustrations may include genes annotation and gene orientation.

[0103] FIG.5 illustrates a non-limiting embodiment of a visualization depicting transgene sequence annotation, such as on a screen or monitor or printer to which a computer system performing a method as disclosed herein is performed, for depicting the visualization thereon or thereby. In some embodiments transgene annotation is used to generate an illustration of transgene sequences and the surrounding reference genome sequence. In this non-limiting example chromosome 16 is depicted next to an annotated transgene which includes approximately two copies of hCSF2 and one copy of hIL3. An illustration of one or more transgene annotations may be an automated step in mapping transgene integration. In some examples transgene annotations may include visual indications permitting differentiation of different sequences (e.g., color-coding, different font, etc.) where an entire transgene sequence is provided with different visualizable indications indicating reference genome sequence and a transgene sequence. Additional colors may be used to further illustrate specific genes within the transgene. In some examples, color coded sequence annotations may be useful for designing oligonucleotide primers of transgene junctions.

Claims

Attorney Docket No.: 0303.052AWO WHAT IS CLAIMED IS:

1. A method of mapping one or more transgene integration into a transgenic organism, comprising: (i) providing a custom genome comprising a reference genome sequence and a transgene sequence; (ii) aligning sequence reads of the transgenic organism to the custom genome; (iii) identifying candidate junction reads, wherein the candidate junction reads correspond to sequence reads comprising sequence regions that align to the transgene sequence and to the host sequence; (iv) categorizing the candidate junction reads into junction clusters, wherein each of the junction clusters comprises a plurality of the candidate junction reads each with an alignment segment terminus within a cluster proximity region of the custom genome; (v) determining a transgene junction based on the junction clusters that possess at least three of the candidate junction reads within the cluster proximity region of the custom genome; and (vi) mapping the integration of the transgene into the transgenic organism based on the transgene junctions.

2. The method of claim 1, wherein the sequence reads are about 750 base pairs or greater.

3. The method of any one of claims 1 or 2, wherein the sequence reads comprise whole genome sequencing reads.

4. The method of any one of claims 1 to 3, wherein the sequence reads comprise sequence reads derived from whole genome sequencing reads with at least a 10x coverage of genome of the transgenic organism.Attorney Docket No.: 0303.052AWO 5. The method of any one of claims 1 to 4, wherein a candidate junction read comprises a sequence region wherein a minimum of about 50 nucleotides align to a reference genome sequence of the custom genome and wherein a minimum of about 50 nucleotides align to a transgene sequence of the custom genome.

6. The method of any one of claims 1 to 5, wherein a candidate junction read comprises a sequence region alignment to a reference genome sequence of a custom genome with a minimum pairwise alignment of about 90% in the sequence region alignment and wherein the candidate junction read comprises a sequence region alignment to a transgene sequence of the custom genome with a minimum pairwise alignment of about 90% in the sequence region alignment.

7. The method of any one of claims 1 to 6, wherein a cluster proximity region comprises a region of about 10-100 nucleotides of the reference genome sequence of the custom genome.

8. The method of any one of claims 1 to 7, wherein a cluster proximity region comprises a region of about 10-100 nucleotides between at least two or more alignment segment termini aligning to the same chromosome of the custom genome.

9. The method of any one of claims 1 to 8, wherein the mapping the integration of the transgene into the transgenic organism further comprises de novo assembly of the transgene junction using the candidate junction clusters.

10. The method of any one of claims 1 to 9, wherein the mapping the integration of the transgene into the transgenic organism further comprises mapping the transgene junctions to the custom genome to determine the junction sequence.

11. The method of any one of claims 1 to 10, wherein the (iii) identifying candidate junction reads comprises storing candidate junction read alignments in Pairwise Mapping Format (PAF) files.Attorney Docket No.: 0303.052AWO 12. A computer system for mapping one or more transgene integration into a transgenic organism, the computer system comprising memory and at least one processor, the computer system configured to execute program instructions to perform a method comprising: (i) providing a custom genome comprising a reference genome sequence and a transgene sequence; (ii) aligning sequence reads of the transgenic organism to the custom genome; (iii) identifying candidate junction reads, wherein the candidate junction reads correlate to candidate junctions comprising sequence regions that align to the transgene sequence and to the host sequence; (iv) categorizing the candidate junctions into junction clusters, wherein each of the junction clusters comprises a plurality of the candidate junctions each with an alignment segment terminus within a cluster proximity region within the custom genome; (v) determining a transgene junction based on the junction clusters that possesses at least three of the candidate junctions within the cluster proximity region of the custom genome; and (vi) mapping the integration of the transgene into the transgenic organism based on the transgene junctions.

13. The system of claim 12, wherein the sequence reads are about 750 base pairs or greater.

14. The system of any one of claims 12 or 13, wherein the sequence reads comprise sequence reads derived from whole genome sequencing reads.Attorney Docket No.: 0303.052AWO 15. The system of any one of claims 12 to 14, wherein the sequence reads comprise sequence reads derived from whole genome sequencing reads with at least a 10x coverage of genome of the transgenic organism.

16. The system of any one of claims 12 to 15, wherein a candidate junction read comprises a sequence region wherein a minimum of about 50 nucleotides align to a reference genome sequence of the custom genome and wherein a minimum of about 50 nucleotides align to a transgene sequence of the custom genome.

17. The system of any one of claims 12 to 16, wherein a candidate junction read comprises a sequence region alignment to a reference genome sequence of a custom genome with a minimum pairwise alignment of about 90% in the sequence region alignment and wherein the candidate junction read comprises a sequence region alignment to a transgene sequence of the custom genome with a minimum pairwise alignment of about 90% in the sequence region alignment.

18. The system of any one of claims 12 to 17, wherein a cluster proximity region comprises a region of about 10-100 nucleotides.

19. The system of any one of claims 12 to 18, wherein a cluster proximity region comprises a region of about 10-100 nucleotides between at least two or more alignment segment termini aligning to the same chromosome of the custom genome.

20. The system of any one of claims 12 to 19, wherein the candidate junction clusters are used for de novo assembly of the transgene junction.

21. The system of any one of claims 12 to 20, wherein a cluster proximity region comprises a region of about 10-100 nucleotides of the custom genome as defined by at least one alignment segment terminus.Attorney Docket No.: 0303.052AWO 22. The system of any one of claims 12 to 21, wherein the mapping the integration of the transgene into the transgenic organism further comprises de novo assembly of the transgene junction using the candidate junction clusters.

23. The system of any one of claims 12 to 22, wherein the mapping the integration of the transgene into the transgenic organism further comprises mapping the transgene junctions to the custom genome to determine the junction sequence.

24. The system of any one of claims 12 to 23, wherein the (iii) identifying candidate junction reads comprises storing candidate junction read alignments in Pairwise Mapping Format (PAF) files.

25. A computer program for mapping one or more transgene integration into a transgenic organism, the computer program comprising: a tangible storage medium storing program instructions for execution to perform a method comprising: (i) providing a custom genome comprising a reference genome sequence and a transgene sequence; (ii) aligning sequence reads of the transgenic organism to the custom genome; (iii) identifying candidate junction reads, wherein the candidate junction reads correlate to candidate junctions comprising sequence regions that align to the transgene sequence and to the host sequence; (iv) categorizing the candidate junctions into junction clusters, wherein each of the junction clusters comprises a plurality of the candidate junctions each with an alignment segment terminus that maps to a cluster proximity region within the custom genome; (v) determining a transgene junction based on the junction clusters that possesses at least three of the candidate junctions mapping to the cluster proximity region of the custom genome; andAttorney Docket No.: 0303.052AWO (vi) mapping the integration of the transgene into the transgenic organism based on the transgene junctions.

26. The computer program of claim 25, wherein the sequence reads are about 750 base pairs or greater.

27. The computer program of any one of claims 25 or 26, wherein the sequence reads comprise whole genome sequencing reads.

28. The computer program of any one of claims 25 to 27, wherein the sequence reads comprise sequence reads derived from whole genome sequencing reads with at least a 10x coverage of genome of the transgenic organism.

29. The computer program of any one of claims 25 to 28, wherein a candidate junction read comprises a sequence region wherein a minimum of about 50 nucleotides align to a reference genome sequence of the custom genome and wherein a minimum of about 50 nucleotides align to a transgene sequence of the custom genome.

30. The computer program of any one of claims 25 to 29, wherein a candidate junction read comprises a sequence region alignment to a reference genome sequence of a custom genome with a minimum pairwise alignment of about 90% in the sequence region alignment and wherein the candidate junction read comprises a sequence region alignment to a transgene sequence of the custom genome with a minimum pairwise alignment of about 90% in the sequence region alignment.

31. The computer program of any one of claims 25 to 30, wherein a cluster proximity region comprises a region of about 10-100 nucleotides of the reference genome sequence of the custom genome.Attorney Docket No.: 0303.052AWO 32. The computer program of any one of claims 25 to 31, wherein a cluster proximity region comprises a region of about 10-100 nucleotides between at least two or more alignment segment termini aligning to the same chromosome of the custom genome.

33. The computer program of any one of claims 25 to 32, wherein the mapping the integration of the transgene into the transgenic organism further comprises de novo assembly of the transgene junction using the candidate junction clusters.

34. The computer program of any one of claims 25 to 33, wherein the mapping the integration of the transgene into the transgenic organism further comprises mapping the transgene junctions to the custom genome to determine the junction sequence.

35. The computer program of any one of claims 25 to 34, wherein the (iii) identifying candidate junction reads comprises storing candidate junction read alignments in Pairwise Mapping Format (PAF) files.