Model for selecting guide RNA molecules

A system using machine learning and AI modeling predicts guide RNA performance and INDEL effects in plants, addressing inefficiencies in current methods to achieve precise genetic editing in crops like maize, soybean, and wheat.

WO2026055608A1PCT designated stage Publication Date: 2026-03-12INARI AGRICULTURE TECHNOLOGY INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-08
Publication Date
2026-03-12

AI Technical Summary

Technical Problem

Current techniques for selecting guide RNAs and predicting the effect of INDELS in genetic modification of plants are inefficient and lead to unpredictable results, limiting the efficiency and effectiveness of genetic editing.

Method used

A system and method using machine learning and artificial intelligence modeling to predict target gene editing frequency scores and functional effects of INDELS, enabling the selection of preferred guide RNAs for precise genetic modifications in plants.

Benefits of technology

The system efficiently and accurately predicts the phenotypic effects of INDELS, allowing for the selection of guide RNAs that achieve desired genetic outcomes in plants, such as maize, soybean, wheat, and rice, with improved speed, safety, and cost-effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025045365_12032026_PF_FP_ABST
    Figure US2025045365_12032026_PF_FP_ABST
Patent Text Reader

Abstract

System(s) and / or method(s) provide for selection of a preferred guide RNA and for prediction of functional effect(s) of one or more INDELS. Such system(s) and / or method(s) can, at least partially, be computer-implemented. The system(s) and / or method(s) can include operations comprising predicting a target gene editing frequency score for each of a plurality of candidate guide RNA, predicting a functional effect of an INDELS obtained in a target plant gene, and selecting a preferred guide RNA based on the target gene editing frequency score and / or the functional effect of the INDELS.
Need to check novelty before this filing date? Find Prior Art

Description

Agent Ref. No. P14786WOOOTITLE : MODEL FOR SELECTING GUIDE RNA MOLECULESCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority under 35 U.S.C. § 119 to provisional patent application U.S. Serial No. 63 / 692,600, filed September 9, 2024. The provisional patent application is herein incorporated by reference in its entirety , including without limitation, the specification, claims, and abstract, as well as any figures, tables, appendices, or drawings thereof.INCORPORATION OF SEQUENCE LISTING

[0002] The sequence listing contained in the file named “P14786WOOO.XML” which is 27 ,567 bytes (measured in MS-Windows®), comprises biological sequences, and was created on August 15, 2025, is electronically filed herewith and is incorporated herein by reference in its entirety.TECHNICAL FIELD

[0003] The present disclosure relates generally to system(s) and / or method(s) having applications in at least the genetic editing, genetic modification, plant modification, and agricultural industries. More particularly, but not exclusively, the present disclosure relates to system(s) and / or method(s) for selecting guide RNAs, predicting guide RNA performance, and predicting the effect of one or more INDELS in planta.BACKGROUND

[0004] The background description provided herein gives context for the present disclosure. Work of the presently named inventors, as well as aspects of the description that may not otherwise qualify as prior art at the time of filing, are neither expressly nor impliedly admitted as prior art.

[0005] A key aspect of some approaches to genetic modification of plants involves using guide RNAs (gRNAs) and CAS nucleases to edit genes. Different types of guide RNA are more or less effective at editing genes with CAS nucleases depending on a variety of factors. Current techniques in selecting effective guide RNA to edit genes are slow and inefficient. Additionally, current techniques of selecting guide RNA often lead to unpredictable results which limits the efficiency and effectiveness of the genetic editing.Agent Ref. No. P14786WOOO

[0006] Another aspect of some approaches to genetic modification of plants with gRNAs and Cas nucleases involves generating an insertion, deletion, and / or substitution (INDELS) at a particular location of the plant's DNA in order to modify the plant so that it exhibits an intended phenotype. Current approaches to generating one or more INDELS with gRNAs and Cas nucleases often lead to unpredictable results for any given gRNA and Cas nuclease, which limits recovery' of plant with intended phenotypes and one or more INDELS.

[0007] Thus, there exists a need in the art for systems and / or methods that efficiently select a preferred guide RNA that is effective to provide a desired outcome. There also exists a need in the art for systems and / or methods that can efficiently and effectively predict the phenotypic effect of one or more INDELS generated by a given gRNA and Cas nuclease to allow selection of gRNAs more likely to provide one or more INDELS with the intended phenotype.SUMMARY

[0008] The following objects, features, advantages, aspects, and / or embodiments, are not exhaustive and do not limit the overall disclosure. No single embodiment need provide each and every' object, feature, or advantage. Any of the objects, features, advantages, aspects, and / or embodiments disclosed herein can be integrated with one another, either in full or in part.

[0009] It is a primary object, feature, and / or advantage of the present disclosure to improve on or overcome the deficiencies in the art.

[0010] It is a further object, feature, and / or advantage of aspects and / or embodiments shown and / or described in the present disclosure to predict a target gene editing frequency score for each of a plurality of distinct candidate guide RNAs (gRNAs).

[0011] It is a further object, feature, and / or advantage of aspects and / or embodiments shown and / or described in the present disclosure to predict functional effect(s) of one or more INDELS obtained in a target plant gene for each of the plurality' of distinct candidate gRNAs.

[0012] It is still yet a further object, feature, and / or advantage of aspects and / or embodiments shown and / or described in the present disclosure to select the preferred gRNA based on the target gene editing frequency score for each of the candidate gRNAs and the predicted functional effect(s) of the one or more INDELS for each of the candidate gRNAs.Agent Ref. No. P14786WOOO

[0013] The system(s) and / or method(s) disclosed herein can be used in a wide variety7of applications. For example, the system(s) and / or method(s) described herein can be used in plants including maize, soybean, wheat, rice, sorghum, cotton, and the like. Additionally, the system(s) and / or method(s) described herein can be used to analyze many different candidate gRNAs.

[0014] It is preferred the system(s) and / or method(s) be safe, effective, cost-effective, efficient, and speedy. For example, the system(s) and / or method(s) can analyze large amounts of data and make determinations and / or predictions of target gene editing frequency and functional effect(s) of one or more INDELS generated by candidate gRNAs and Cas nucleases quickly and efficiently. Additionally, the system(s) and / or method(s) can utilize machine learning and / or artificial intelligence modeling techniques. The use of such machine learning and / or artificial intelligence modeling techniques helps to increase efficiency, speed, effectiveness, and cost-effectiveness by continuously improving over time.

[0015] Methods can be practiced which facilitate use, manufacture, assembly, maintenance, and repair of system(s) described herein which accomplish some or all of the previously stated objectives.

[0016] The method(s) described herein can be incorporated into system(s) which accomplish some or all of the previously stated objectives.

[0017] The system(s) described herein can be incorporated into larger design(s) which accomplish some or all of the previously stated objectives.

[0018] According to some aspects of the present disclosure, a system for selection of a preferred guide RNA for obtaining a plant or plant part comprising a desired DNA insertion, deletion, and / or substitution (INDELS) in a target plant gene via Cas nuclease mediated gene editing comprises: a processing unit: a non-transitory computer-readable medium that stores executable instructions that, when executed by the processing unit, perform operations, the operations comprising: (i) predicting a target gene editing frequency score for each of a plurality of distinct candidate guide RNAs (gRNAs) independently introduced into a plant cell with a Cas nuclease encoding DNA vector, wherein the candidate gRNAs are not associated with the Cas nuclease and the Cas nuclease is not present in the plant cell prior to their introduction; and (ii) selecting the preferred gRNA from the plurality of distinct candidate gRNAs for use in obtaining the plant or plant part comprising the desired INDELS in a target plant gene based on the target gene editing frequency score.Agent Ref. No. P14786WOOO

[0019] According to some aspects of the present disclosure, a system for selection of a preferred guide RNA for obtaining a plant or plant part comprising a desired DNA insertion, deletion, and / or substitution (INDELS) in a target plant gene encoding a protein via Cas nuclease mediated gene editing comprises: a processing unit; a non-transitory computer- readable medium that stores executable instructions that, when executed by the processing unit, perform operations, the operations comprising: (i) predicting a functional effect of an INDELS obtained in a target plant gene via Cas nuclease mediated gene editing for each of a plurality of distinct candidate guide RNAs (gRNAs) independently introduced into a plant cell with a Cas nuclease encoding DNA vector, wherein the candidate gRNAs are not associated with the Cas nuclease and the Cas nuclease is not present in the plant cell prior to their introduction; and (ii) selecting the preferred gRNA based on the predicted functional effect of the INDELS.

[0020] According to some aspects of the present disclosure, a system for selection of a preferred guide RNA for obtaining a plant or plant part comprising a desired DNA insertion, deletion, and / or substitution (INDELS) in a target plant gene encoding a protein via Cas nuclease mediated gene editing comprises: a processing unit; a non-transitory computer- readable medium that stores executable instructions that, when executed by the processing unit, perform operations, the operations comprising: (i) predicting a target gene editing frequency score for each of a plurality of distinct candidate guide RNAs (gRNAs) independently introduced into a plant cell with a Cas nuclease encoding DNA vector, wherein the candidate gRNAs are not associated with the Cas nuclease and the Cas nuclease is not present in the plant cell prior to their introduction; (ii) predicting a functional effect of an INDELS obtained in a target plant gene via Cas nuclease mediated gene editing for each of the plurality of distinct candidate gRNAs; and (iii) selecting the preferred gRNA based on the target gene editing frequency score and the predicted functional effect of the INDELS.

[0021] According to some aspects of the present disclosure, the candidate gRNAs and the Cas nuclease encoding DNA vector are introduced into at least one plant protoplast cell via PEG-mediated transfection and / or electroporation.

[0022] According to some aspects of the present disclosure, the predicting a target gene editing frequency score step further comprises filtering a set of alleles to remove allele(s) outside an expected editing window and / or to remove false positive(s).Agent Ref. No. P14786WOOO

[0023] According to some aspects of the present disclosure, the predicting a target gene editing frequency score step further comprises determining: (i) an average editing efficiency in the plant cell for a subset of test gRNAs directed to the target gene; and (ii) an actual editing outcome in the plant or plant part for the subset of test gRNAs.

[0024] According to some aspects of the present disclosure, the predicting a target gene editing frequency score step further comprises applying a linear model to correlate the average editing efficiency in the plant cell with the actual editing outcome in the plant or plant part for the subset of test gRNAs.

[0025] According to some aspects of the present disclosure, the linear model is configured to:(i) apply least squares regression to the correlation of the average editing efficiency in the plant cell and the actual editing outcome in the plant or plant part for each test gRNA; (ii) derive an equation to determine a line of best fit of the correlation of the average editing efficiency in the plant cell and the actual editing outcome in the plant or plant part of each test gRNA based, at least in part, on the least squares regression; and (iii) predict editing efficiency of a candidate gRNA using the derived equation.

[0026] According to some aspects of the present disclosure, the selecting step comprises: (i) ranking a plurality of gRNAs based on the predicted target gene editing frequency score to obtain a subset of top-ranked gRNAs with the highest target gene editing frequency scores;(ii) identifying a desired allele among gene editing alleles produced in the plant cell by the top-ranked gRNA; (iii) selecting a gRNA which is predicted to provide the desired allele.

[0027] According to some aspects of the present disclosure, the number of top-ranked gRNAs is at least five.

[0028] According to some aspects of the present disclosure, the number of top-ranked gRNAs is at least ten.

[0029] According to some aspects of the present disclosure, the predicting of the functional effect of the INDELS step further comprises selecting a subset of gRNAs based on their target gene editing frequency score.

[0030] According to some aspects of the present disclosure, the predicting of the functional effect of the INDELS step further comprises: inputting a set of alleles produced by the subset of gRNAs into a protein language model; processing the set of alleles via the model; outputting a ratio from the model wherein the ratio can be used to evaluate the predicted functional effect of each allele of the set of alleles, and selecting one or more gRNAs fromAgent Ref. No. P14786WOOO the subset of gRNAs based on the predicted functional effect of the alleles produced by the gRNA.

[0031] According to some aspects of the present disclosure, the predicting of the functional effect of the INDELS step comprises using a model that utilizes machine learning and is trained in a self-supervised manner.

[0032] According to some aspects of the present disclosure, evolutionary information is integrated into parameters of the model.

[0033] According to some aspects of the present disclosure, the predicting of the functional effect of the INDELS step further comprises inputting first and second protein sequence variants into the model.

[0034] According to some aspects of the present disclosure, the predicting of the functional effect of the INDELS step further comprises: (i) masking a single amino acid position of the first sequence variant; (ii) predicting a probability distribution over all possible amino acids at the single amino acid position based on surrounding sequence context and the evolutionary information; (iii) determining a conditional probability of the single amino acid position based on the probability distribution; (iv) performing steps (i)-(iii) for each amino acid position of a length of the first sequence variant in a one-by-one manner; and (v) calculating a product of each conditional probability of each amino acid position of the length to produce an approximated joint probability of the first sequence variant.

[0035] According to some aspects of the present disclosure, the predicting of the functional effect of the INDELS step further comprises: (i) masking a single amino acid position of the second sequence variant; (ii) predicting a probability distribution over all possible amino acids at the single amino acid position of the second sequence variant based on surrounding sequence context and the evolutionary information; (iii) determining a conditional probability of the single amino acid position of the second sequence vanant based on the probability distribution; (iv) performing steps (i)-(iii) for each amino acid position of a length of the second sequence variant in a one-by-one manner; and (v) calculating a product of each conditional probability of each amino acid position of the length of the second sequence variant to produce an approximated joint probability of the second sequence variant.

[0036] According to some aspects of the present disclosure, the predicting of the functional effect of the INDELS step further comprises calculating a quotient of the approximated jointAgent Ref. No. P14786WOOO probability of the first sequence variant divided by the approximated joint probability of the second sequence variant.

[0037] According to some aspects of the present disclosure, the second sequence variant differs from the first sequence variant by an in-frame INDEL such that the length of the second sequence variant differs from the length of the first sequence variant, and the predicting of the functional effect of the INDELS step further comprises calculating a quotient of a geometric mean of the approximated joint probability of the first sequence variant divided by a geometric mean of the approximated joint probability of the second sequence variant.

[0038] According to some aspects of the present disclosure, the predicting of the functional effect of the INDELS step further comprises: (i) partitioning the first sequence variant into one or more groups of amino acid positions; (ii) masking a group of amino acid positions of the one or more groups of amino acid positions of the first sequence variant; (iii) predicting a probability distribution associated with the group of amino acid positions based on surrounding sequence context and the evolutionary information; (iv) simultaneously determining a conditional probability of all amino acid positions of the group of amino acid positions based on the probability distribution; (v) performing steps (ii)-(iv) for each group of the one or more groups of amino acid positions of a length of the first sequence variant in a group-by -group manner; and (vi) calculating a product of each conditional probability of each group of the one or more groups of amino acid positions of the length to produce an approximated joint probability of the first sequence variant.

[0039] According to some aspects of the present disclosure, the predicting of the functional effect of the INDELS step further comprises: (i) partitioning the second sequence variant into one or more groups of amino acid positions; (ii) masking a group of amino acid positions of the one or more groups of amino acid positions of the second sequence variant; (iii) predicting a probability distribution associated with the group of amino acid positions of the second sequence variant based on surrounding sequence context and the evolutionary information; (iv) simultaneously determining a conditional probability of all amino acid positions of the group of amino acid positions of the second sequence variant based on the probability distribution associated with the second sequence variant; (v) performing steps (ii)- (iv) for each group of the one or more groups of amino acid positions of a length of the second sequence variant in a group-by-group manner; and (vi) calculating a product of eachAgent Ref. No. P14786WOOO conditional probability of each group of the one or more groups of amino acid positions of the length of the second sequence variant to produce an approximated joint probability of the second sequence variant.

[0040] According to some aspects of the present disclosure, the predicting of the functional effect of the INDELS step further comprises calculating a quotient of the approximated joint probability of the first sequence variant divided by the approximated joint probability of the second sequence variant.

[0041] According to some aspects of the present disclosure, the second sequence variant differs from the first sequence variant by an in-frame INDEL such that the length of the second sequence variant differs from the length of the first sequence variant, and wherein the predicting of the functional effect of the INDELS step further comprises calculating a quotient of a geometric mean of the approximated joint probability of the first sequence variant divided by a geometric mean of the approximated joint probability of the second sequence variant.

[0042] According to some aspects of the present disclosure, each of the one or more groups of amino acid positions of the first sequence variant and each of the one or more groups of amino acid positions of the second sequence variant comprise random amino acid positions.

[0043] According to some aspects of the present disclosure, the predicting of the functional effect of the INDELS step further comprises: (i) predicting a probability' distribution over all possible amino acids at each amino acid position of the first sequence variant based on surrounding sequence context and the evolutionary information; (ii) determining a conditional probability' of each amino acid position of the first sequence variant based on the probability distribution of each amino acid position; and (iii) calculating a product of each conditional probability of each amino acid position of the first sequence variant to produce an approximated joint probability of the first sequence variant.

[0044] According to some aspects of the present disclosure, the predicting of the functional effect of the INDELS step further comprises: (i) predicting a probability distribution over all possible amino acids at each amino acid position of the second sequence variant based on surrounding sequence context and the evolutionary information; (ii) determining a conditional probability of each amino acid position of the second sequence variant based on the probability distribution of each amino acid position of the second sequence variant; and (iii) calculating a product of each conditional probability' of each amino acid position of theAgent Ref. No. P14786WOOO second sequence variant to produce an approximated joint probability of the second sequence variant; wherein the second sequence variant and the first sequence variant have the same length.

[0045] According to some aspects of the present disclosure, the predicting of the functional effect of the INDELS step further comprises calculating a quotient of the approximated joint probability of the first sequence variant divided by the approximated joint probability' of the second sequence variant.

[0046] According to some aspects of the present disclosure, the predicting of the functional effect of the INDELS step further comprises using a model that utilizes machine learning and is trained in a self-supervised manner.

[0047] According to some aspects of the present disclosure, the predicting of the functional effect of the INDELS step further comprises inputting first and second protein sequence variants into the model.

[0048] According to some aspects of the present disclosure, the predicting of the functional effect of the INDELS step further comprises using the model to compute an alignment-based dimensionless score of relative fitness regarding the first and second sequence variants.

[0049] According to some aspects of the present disclosure, the predicted functional effect of the INDELS and / or the preferred gRNA can be used to select and / or predict in-planta information.

[0050] According to some aspects of the present disclosure, the in-planta information comprises phenotype information.

[0051] According to some aspects of the present disclosure, the system further comprises an experimental plant, experimental plant part, or experimental plant cell lacking the desired INDELS wherein: (a) (i) the selected preferred gRNA or a polynucleotide encoding the selected preferred gRNA; and (ii) a Cas nuclease which binds the selected preferred gRNA or a polynucleotide encoding the Cas nuclease are introduced into the experimental plant, experimental plant part, or experimental plant cell lacking the desired INDELS; or (b) a ribonucleoprotein complex comprising the Cas nuclease bound to the selected preferred gRNA is introduced into the experimental plant, experimental plant part, or experimental plant cell lacking the desired INDELS.Agent Ref. No. P14786WOOO

[0052] According to some aspects of the present disclosure, the plant or plant part comprising the desired INDELS is obtained from the experimental plant, experimental plant part, or experimental plant cell lacking the desired INDELS.

[0053] According to some aspects of the present disclosure, (i) the plant, plant part, experimental plant, and / or experimental plant part is and / or comprises a seed, leaf, stem, flower, meristem, or root; or (ii) the plant cell and / or experimental plant cell comprises plant callus tissue, embryogenic callus tissue, or a protoplast.

[0054] According to some aspects of the present disclosure, the system further comprises one or more DNA extraction, DNA amplification, and / or DNA sequencing device(s) configured to isolate, amplify, and / or sequence genomic DNA sequences containing the target plant gene from an experimental plant cell that had received the distinct candidate guide RNA and the Cas nuclease encoding DNA vector, wherein the isolated, amplified, and / or sequenced genomic DNA sequences can be used in the step of predicting the target gene editing frequency score and / or in the step of predicting the functional effect of the INDELS.

[0055] According to some aspects of the present disclosure, a method for selection of a preferred guide RNA for obtaining a plant or plant part comprising a desired DNA insertion, deletion, and / or substitution (INDELS) in a target plant gene via Cas nuclease mediated gene editing comprises: (i) predicting a target gene editing frequency score for each of a plurality of distinct candidate guide RNAs (gRNAs) independently introduced into a plant cell with a Cas nuclease encoding DNA vector, wherein the candidate gRNAs are not associated with the Cas nuclease and the Cas nuclease is not present in the plant cell prior to their introduction; and (ii) selecting the preferred gRNA from the plurality of distinct candidate gRNAs for use in obtaining the plant or plant part comprising the desired INDELS in a target plant gene based on the target gene editing frequency score.

[0056] According to some aspects of the present disclosure, a method for selection of a preferred guide RNA for obtaining a plant or plant part comprising a desired DNA insertion, deletion, and / or substitution (INDELS) in a target plant gene encoding a protein via Cas nuclease mediated gene editing comprises: (i) predicting a functional effect of an INDELS obtained in a target plant gene via Cas nuclease mediated gene editing for each of a plurality of distinct candidate guide RNAs (gRNAs) independently introduced into a plant cell with a Cas nuclease encoding DNA vector, wherein the candidate gRNAs are not associated with the Cas nuclease and the Cas nuclease is not present in the plant cell prior to theirAgent Ref. No. P14786WOOO introduction; and (ii) selecting the preferred gRNA based on the predicted functional effect of the INDELS.

[0057] According to some aspects of the present disclosure, a method for selection of a preferred guide RNA for obtaining a plant or plant part comprising a desired DNA insertion, deletion, and / or substitution (INDELS) in a target plant gene encoding a protein via Cas nuclease mediated gene editing comprises: (i) predicting a target gene editing frequency score for each of a plurality of distinct candidate guide RNAs (gRNAs) independently introduced into a plant cell with a Cas nuclease encoding DNA vector, wherein the candidate gRNAs are not associated with the Cas nuclease and the Cas nuclease is not present in the plant cell prior to their introduction; (ii) predicting a functional effect of an INDELS obtained in a target plant gene via Cas nuclease mediated gene editing for each of the plurality of distinct candidate gRNAs; and (iii) selecting the preferred gRNA based on the target gene editing frequency score and the predicted functional effect of the INDELS.

[0058] According to some aspects of the present disclosure, the candidate gRNAs and the Cas nuclease encoding DNA vector are introduced into at least one plant protoplast cell via PEG-mediated transfection and / or electroporation.

[0059] According to some aspects of the present disclosure, the predicting a target gene editing frequency score step further comprises filtering a set of alleles to remove allele(s) outside an expected editing window and / or to remove false positive(s).

[0060] According to some aspects of the present disclosure, the predicting a target gene editing frequency score step further comprises determining: (i) an average editing efficiency in the plant cell for a subset of test gRNAs directed to the target gene; and (ii) an actual editing outcome in the plant or plant part for the subset of test gRNAs.

[0061] According to some aspects of the present disclosure, the predicting a target gene editing frequency score step further comprises applying a linear model to correlate the average editing efficiency in the plant cell with the actual editing outcome in the plant or plant part for the subset of test gRNAs.

[0062] According to some aspects of the present disclosure, the linear model is configured to: (i) apply least squares regression to the correlation of the average editing efficiency in the plant cell and the actual editing outcome in the plant or plant part for each test gRNA; (ii) derive an equation to determine a line of best fit of the correlation of the average editing efficiency in the plant cell and the actual editing outcome in the plant or plant part of eachAgent Ref. No. P14786WOOO test gRNA based, at least in part, on the least squares regression; and (iii) predict editing efficiency of a candidate gRNA using the derived equation.

[0063] According to some aspects of the present disclosure, the selecting step comprises: (i) ranking a plurality of gRNAs based on the predicted target gene editing frequency score to obtain a subset of top-ranked gRNAs with the highest target gene editing frequency scores; (ii) identifying a desired allele among gene editing alleles produced in the plant cell by the top-ranked gRNA; and (iii) selecting a gRNA which is predicted to provide the desired allele.

[0064] According to some aspects of the present disclosure, the number of top-ranked gRNAs is at least five.

[0065] According to some aspects of the present disclosure, the number of top-ranked gRNAs is at least ten.

[0066] According to some aspects of the present disclosure, the predicting of the functional effect of the INDELS step further comprises selecting a subset of gRNAs based on their target gene editing frequency score.

[0067] According to some aspects of the present disclosure, the predicting of the functional effect of the INDELS step further comprises: inputting a set of alleles produced by the subset of gRNAs into a protein language model; processing the set of alleles via the model; outputting a ratio from the model wherein the ratio can be used to evaluate the predicted functional effect of each allele of the set of alleles, and selecting one or more gRNAs from the subset of gRNAs based on the predicted functional effect of the alleles produced by the gRNA.

[0068] According to some aspects of the present disclosure, the predicting of the functional effect of the INDELS step comprises using a model that utilizes machine learning and is trained in a self-supervised manner.

[0069] According to some aspects of the present disclosure, evolutionary information is integrated into parameters of the model.

[0070] According to some aspects of the present disclosure, the predicting of the functional effect of the INDELS step further comprises inputting first and second protein sequence variants into the model.

[0071] According to some aspects of the present disclosure, the predicting of the functional effect of the INDELS step further comprises: (i) masking a single amino acid position of the first sequence variant; (ii) predicting a probability distribution over all possible amino acidsAgent Ref. No. P14786WOOO at the single amino acid position based on surrounding sequence context and the evolutionary information; (iii) determining a conditional probability of the single amino acid position based on the probability distribution; (iv) performing steps (i)-(iii) for each amino acid position of a length of the first sequence variant in a one-by-one manner; and (v) calculating a product of each conditional probability of each amino acid position of the length to produce an approximated joint probability of the first sequence variant.

[0072] According to some aspects of the present disclosure, the predicting of the functional effect of the INDELS step further comprises: (i) masking a single amino acid position of the second sequence variant; (ii) predicting a probability distribution over all possible amino acids at the single amino acid position of the second sequence variant based on surrounding sequence context and the evolutionary information; (iii) determining a conditional probability of the single amino acid position of the second sequence vanant based on the probability distribution; (iv) performing steps (i)-(iii) for each amino acid position of a length of the second sequence variant in a one-by-one manner; and (v) calculating a product of each conditional probability of each amino acid position of the length of the second sequence variant to produce an approximated joint probability of the second sequence variant.

[0073] According to some aspects of the present disclosure, the predicting of the functional effect of the INDELS step further comprises calculating a quotient of the approximated joint probability of the first sequence variant divided by the approximated joint probability of the second sequence variant.

[0074] According to some aspects of the present disclosure, the second sequence variant differs from the first sequence variant by an in-frame INDEL such that the length of the second sequence variant differs from the length of the first sequence variant, and the predicting of the functional effect of the INDELS step further comprises calculating a quotient of a geometric mean of the approximated joint probability of the first sequence variant divided by a geometric mean of the approximated joint probability of the second sequence variant.

[0075] According to some aspects of the present disclosure, the predicting of the functional effect of the INDELS step further comprises: (i) partitioning the first sequence variant into one or more groups of amino acid positions; (ii) masking a group of amino acid positions of the one or more groups of amino acid positions of the first sequence variant; (iii) predicting a probability distribution associated with the group of amino acid positions based onAgent Ref. No. P14786WOOO surrounding sequence context and the evolutionary7information; (iv) simultaneously determining a conditional probability of all amino acid positions of the group of amino acid positions based on the probability distribution; (v) performing steps (ii)-(iv) for each group of the one or more groups of amino acid positions of a length of the first sequence variant in a group-by -group manner; and (vi) calculating a product of each conditional probability of each group of the one or more groups of amino acid positions of the length to produce an approximated joint probability of the first sequence variant.

[0076] According to some aspects of the present disclosure, the predicting of the functional effect of the INDELS step further comprises: (i) partitioning the second sequence variant into one or more groups of amino acid positions; (ii) masking a group of amino acid positions of the one or more groups of amino acid positions of the second sequence variant; (iii) predicting a probability distribution associated with the group of ammo acid positions of the second sequence variant based on surrounding sequence context and the evolutionary information; (iv) simultaneously determining a conditional probability of all amino acid positions of the group of amino acid positions of the second sequence variant based on the probability distribution associated with the second sequence variant; (v) performing steps (ii)- (iv) for each group of the one or more groups of amino acid positions of a length of the second sequence variant in a group-by-group manner; and (vi) calculating a product of each conditional probability of each group of the one or more groups of amino acid positions of the length of the second sequence variant to produce an approximated joint probability of the second sequence variant.

[0077] According to some aspects of the present disclosure, the predicting of the functional effect of the INDELS step further comprises calculating a quotient of the approximated joint probability of the first sequence variant divided by the approximated joint probability of the second sequence variant.

[0078] According to some aspects of the present disclosure, the second sequence variant differs from the first sequence variant by an in-frame INDEL such that the length of the second sequence variant differs from the length of the first sequence variant, and wherein the predicting of the functional effect of the INDELS step further comprises calculating a quotient of a geometric mean of the approximated joint probability of the first sequence variant divided by a geometric mean of the approximated joint probability of the second sequence variant.Agent Ref. No. P14786WOOO

[0079] According to some aspects of the present disclosure, each of the one or more groups of amino acid positions of the first sequence variant and each of the one or more groups of amino acid positions of the second sequence variant comprise random ammo acid positions.

[0080] According to some aspects of the present disclosure, the predicting of the functional effect of the INDELS step further comprises: (i) predicting a probability distribution over all possible amino acids at each amino acid position of the first sequence variant based on surrounding sequence context and the evolutionary information; (ii) determining a conditional probability of each amino acid position of the first sequence variant based on the probability distribution of each amino acid position; and (iii) calculating a product of each conditional probability of each amino acid position of the first sequence variant to produce an approximated joint probability of the first sequence variant.

[0081] According to some aspects of the present disclosure, the predicting of the functional effect of the INDELS step further comprises: (i) predicting a probability distribution over all possible amino acids at each amino acid position of the second sequence variant based on surrounding sequence context and the evolutionary information; (ii) determining a conditional probability of each amino acid position of the second sequence variant based on the probability distribution of each amino acid position of the second sequence variant; and (iii) calculating a product of each conditional probability of each amino acid position of the second sequence variant to produce an approximated joint probability of the second sequence variant; wherein the second sequence variant and the first sequence variant have the same length.

[0082] According to some aspects of the present disclosure, the predicting of the functional effect of the INDELS step further comprises calculating a quotient of the approximated joint probability of the first sequence variant divided by the approximated joint probability of the second sequence variant.

[0083] According to some aspects of the present disclosure, the predicting of the functional effect of the INDELS step further comprises using a model that utilizes machine learning and is trained in a self-supervised manner.

[0084] According to some aspects of the present disclosure, the predicting of the functional effect of the INDELS step further comprises inputting first and second protein sequence variants into the model.Agent Ref. No. P14786WOOO

[0085] According to some aspects of the present disclosure, the predicting of the functional effect of the INDELS step further comprises using the model to compute an alignment-based dimensionless score of relative fitness regarding the first and second sequence variants.

[0086] According to some aspects of the present disclosure, the predicted functional effect of the INDELS and / or the preferred guide RNA can be used to select and / or predict in-planta information.

[0087] According to some aspects of the present disclosure, the in-planta information comprises phenotype information.

[0088] According to some aspects of the present disclosure, the method further comprises introducing into an experimental plant, experimental plant part, or experimental plant cell lacking the desired INDELS: (a) (i) the selected preferred gRNA or a polynucleotide encoding the selected preferred gRNA; and (li) a Cas nuclease which binds the selected preferred gRNA or a polynucleotide encoding the Cas nuclease; or (b) a ribonucleoprotein complex comprising the Cas nuclease bound to the selected preferred gRNA.

[0089] According to some aspects of the present disclosure, the method further comprises obtaining the plant or plant part comprising the desired INDELS from the experimental plant, experimental plant part, or experimental plant cell lacking the desired INDELS.

[0090] According to some aspects of the present disclosure, (i) the plant, plant part, experimental plant, and / or experimental plant part is and / or comprises a seed, leaf, stem, flower, meristem, or root; or (ii) the plant cell and / or experimental plant cell comprises plant callus tissue, embryogenic callus tissue, or a protoplast.

[0091] According to some aspects of the present disclosure, the method further comprises isolating, amplifying, and / or sequencing genomic DNA sequences containing the target plant gene from an experimental plant cell that had received the distinct candidate guide RNA and the Cas nuclease encoding DNA vector, wherein the isolated, amplified, and sequenced genomic DNA sequences can be used in the step of predicting the target gene editing frequency score and / or in the step of predicting the functional effect of the INDELS.

[0092] These and / or other objects, features, advantages, aspects, and / or embodiments will become apparent to those skilled in the art after reviewing the following brief and detailed descriptions of the drawings. Furthermore, the present disclosure encompasses aspects and / or embodiments not expressly disclosed but which can be understood from a reading of theAgent Ref. No. P14786WOOO present disclosure, including at least: (a) combinations of disclosed aspects and / or embodiments and / or (b) reasonable modifications not shown or described.BRIEF DESCRIPTION OF THE DRAWINGS

[0093] Several embodiments in which the present disclosure can be practiced are illustrated and described in detail, wherein like reference characters represent like components throughout the several views. The drawings are presented for exemplary purposes and may not be to scale unless otherwise indicated.

[0094] As will be understood, any of the aspects of any of the embodiments shown and / or described herein could be combined with one another to form any number of embodiments, whether expressly disclosed or not, which would be understood by one skilled in the art.

[0095] Figure 1 shows a block diagram of a system configured to perform operations related to genetic modification of plants.

[0096] Figure 2 shows a flow chart of a method for genetic modification of plants.

[0097] Figure 3 shows a flow chart of a method for genetic modification of plants.

[0098] Figure 4 shows a flow chart of a method for genetic modification of plants.

[0099] Figures 5A-G show a flow chart of a method for genetic modification of plants.

[0100] Figure 6 shows a chart comparing aspects of three different approaches for evaluating guide RNA performance in plant cells or plants. The left column is deliver}' of RNPs comprising Cas nucleases and gRNAs to protoplasts. The middle column is delivery of a plasmid encoding the Cas nuclease and a gRNA molecule (i.e., hybrid delivery assay). The right column is delivery of DNA encoding the Cas nuclease and DNA encoding the guide RNA to plant cells which are regenerated into plants.

[0101] Figure 7 shows a graphical representation comparing gene editing outcomes (combined INDELS) when the three different approaches to Cas nuclease and guide RNA delivery described in Figure 6 are used with eight different guide RNAs (CR1808, CR1810, CR1811, CR1816, CR1819, CR1822, CR1825, and CR1826).

[0102] Figure 8A shows a plot of the proportion of plants with high quality gene editing alleles obtained with in planta plasmid transformation with eight different gRNAs (CR1808. CR 1810, CR1811, CR1816, CR1819, CR1822, CR1825, and CR1826) versus the Cambridge mean combined INDELS.Agent Ref. No. P14786WOOO

[0103] Figure 8B shows a plot of the most frequent alleles obtained with the CR1808 guide RNA obtained with in planta plasmid transformation-based editing where the alleles were either False or True in editing results obtained with that same guide RNA in the hybrid delivery assay.

[0104] Figure 9A shows a plot of the percentage of high quality edited plants obtained with in planta plasmid transformation-based editing versus the mean FPF INDELS % for different gRNAs hybrid delivery assay.

[0105] Figure 9B shows a plot of the percentage of high quality edited plants obtained with in planta plasmid transformation-based editing versus the percentage of high quality edited explants obtained with in planta plasmid transformation-based editing for different gRNAs.

[0106] Figure 10 shows a comparison of editing efficiency results obtained with hybrid delivery (Hybrid Mean FPF INDELS), with in planta plasmid transformation-based editing (Plant pct HQE), and with explants obtained from plasmid transformation-based editing of cultured plant cells (Explant pct HQE) with three different gRNAs (CR1820, CR1821, and CR1933). The CR1933 gRNA would be selected for in planta plasmid transformation-based editing based on its higher editing efficiency in the hybrid delivery assay.

[0107] Figure 11 compares the percentage of instances where the correct gRNA is selected in a hybrid delivery assay versus the percentage of instances where the correct gRNA is selected by assaying explants obtained from plasmid transformation-based editing of cultured plant cells.

[0108] Figure 12 shows a table comparing gene editing results obtained using unoptimized gRNAs with results obtained using a gRNA selected by use of hybrid delivery assays and other methods disclosed herein.

[0109] Figure 13 shows a graphical representation of population level editing efficiency obtained in whole plant gene editing experiments (e.g, in planta plasmid transformationbased editing) versus results obtained in a hybrid delivery assay with plant protoplasts.

[0110] Figures 14A-C show graphical representations comparing alleles obtained with hybrid delivery based editing in protoplasts and alleles obtained with in planta plasmid transformation-based editing with the CR1808, CR1946, and CR1820 gRNAs.

[0111] Figure 15 shows a graphical representation of the percentage of overlap of alleles obtained with hybrid delivery and obtained with in planta plasmid transformation-based editing with different gRNAs (CR1808, CR1820, CR1821, CR1906, CR1907, CR1908,Agent Ref. No. P14786WOOOCR1933, CR1945, CR1947, CR1948, CR1949, CR1950, CR1954, CR1955, CR1956, CR1957, CR1958, CR1959, CR1961, CR1962, CR1963, CR1964, CR1965, CR1967, CR1968. CR1971. CR1972. and CR1973).

[0112] Figure 16 shows a graphical representation of the percentage of overlap of the top 5 alleles obtained with hybrid delivery' and the top 10 alleles obtained with in planta plasmid transformation-based editing with different gRNAs (CR1808, CR1820, CR1821, CR1906, CR1907. CR1908, CR1933, CR1945, CR1947, CR1948, CR1949, CR1950, CR1954, CR1955, CR1956, CR1957, CR1958, CR1959, CR1961, CR1962, CR1963, CR1965, CR1967, CR1968, CR1971, CR1972, and CR1973).

[0113] Figure 17 shows a depiction of a bioinformatics pipeline related to qualitative prediction(s) of gene editing outcomes. Frequently edited alleles are shown in a summary table denoting if the edit resulted in a premature stop codon, frameshift mutation, and the resulting protein length. SEQ ID NOs for each variant coding region (cds) and protein translation are also shown.

[0114] Figure 18 shows a depiction of predicting an edited allele w ith an in frame deletion’s (SEQ ID NO: 18) effect on function of the encoded protein using a particular protein language model by comparing the edited allele to the wild-type allele (SEQ ID NO: 17) using masking (SEQ ID NOs: 19, 20 and 21) to determine the likelihood of its residues.

[0115] Figure 19 show s a graphical representation of prediction(s) of gene editing outcomes based on the protein language model of Figure 18 using data on the effect of 314 single AA deletions in the PTEN tumor suppressor gene in humanized yeast cells.

[0116] Figures 20A-D show' various graphical representations of predictions regarding the effect(s) of edited alleles using the protein language model of Figure 18 using data on the effect of 314 single AA deletions in the PTEN tumor suppressor gene in humanized yeast cells.

[0117] Figure 21 shows a depiction of a pipeline for predicting editing effect(s) of one or more INDELS, comparing a wild-ty pe allele (SEQ ID NO: 17) with an edited allele with an in frame deletion (SEQ ID NO: 18).

[0118] Figure 22 shows the schematic of a gene and various locations for editing a gene to modify gene expression.Agent Ref. No. P14786WOOO

[0119] Figure 23 shows a diagrammatic view of the architecture of a protein language model comparing an edited allele with an in-frame deletion (SEQ ID NO: 18) to a wild-type allele (SEQ ID NO: 17) according to at least some aspects of the present disclosure.

[0120] Figure 24 shows a diagrammatic depiction of multiple masking strategies capable of being used in conjunction with the protein language model of Figure 23. An example edited allele with an in-frame deletion (SEQ ID NO: 18) is shown after applying various masking strategies (SEQ ID NOs: 19. 22. and 23).

[0121] Figure 25 shows a workflow diagram incorporating use of the protein language model of Figure 23.

[0122] Figure 26 shows a workflow diagram incorporating use of a model and / or algorithm according to at least some aspects of the present disclosure.

[0123] An artisan of ordinary skill in the art need not view, within isolated figure(s), the near infinite number of distinct permutations of features described in the following detailed description to facilitate an understanding of the present disclosure.DETAILED DESCRIPTION

[0124] The present disclosure is not to be limited to that described herein. Mechanical, electrical, chemical, procedural, and / or other changes can be made without departing from the spirit and scope of the present disclosure. No features shown or described are essential to permit basic operation of the present disclosure unless otherwise indicated.

[0125] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which embodiments of the present disclosure pertain.

[0126] The terms “a.” “an.” and “the” include both singular and plural referents.

[0127] The term “or” is synonymous with “and / or” and means any one member or combination of members of a particular list.

[0128] The term “and / or” where used herein is to be taken as specific disclosure of each of the two specified features or components with or without the other. Thus, the term “and / or” as used in a phrase such as “A and / or B” herein is intended to include “A and B,” “A or B.” “A” (alone), and “B” (alone). Likewise, the term “and / or” as used in a phrase such as “A, B, and / or C” is intended to encompass each of the following embodiments: A, B, and C; A, B,Agent Ref. No. P14786WOOO or C; A or C; A or B; B or C; A and C; A and B; B and C; A (alone); B (alone); and C (alone).

[0129] The terms “invention” or “present invention” are not intended to refer to any single embodiment of the particular invention but encompass all possible embodiments as described in the specification and the claims.

[0130] The term “about” as used herein refers to slight variations in numerical quantities with respect to any quantifiable variable. Inadvertent error can occur, for example, through use of typical measuring techniques or equipment or from differences in the manufacture, source, or purity of components.

[0131] The term “substantially” refers to a great or significant extent. “Substantially” can thus refer to a plurality, majority, and / or a supermajority of said quantifiable variable, given proper context.

[0132] The term “generally” encompasses both “about” and “substantially.”

[0133] The term “configured” describes structure capable of performing a task or adopting a particular configuration. The term “configured” can be used interchangeably with other similar phrases, such as “constructed”, “arranged”, “adapted”, “manufactured”, and the like.

[0134] Terms characterizing sequential order, a position, and / or an orientation are not limiting and are only referenced according to the views presented.

[0135] The “scope” of the present disclosure is defined by the appended claims, along with the full scope of equivalents to which such claims are entitled. The scope of the disclosure is further qualified as including any possible modification to any of the aspects and / or embodiments disclosed herein which would result in other embodiments, combinations, subcombinations, or the like that would be obvious to those skilled in the art.

[0136] As used herein, the phrases “hybrid delivery” and “hybrid delivery assay” refer to transfer of a DNA molecule encoding a Cas nuclease and an RNA molecule comprising a guide RNA (i.e., gRNA or “guide”) to a cell (e g., a plant cell).

[0137] As used herein, the phrase “amorphic allele” refers to an allele of a gene having no gene activity in comparison to the wild-type allele of the gene. Amorphic alleles are also known as null alleles or “knockout” alleles.

[0138] As used herein, the term “expression” refers to the production of a functional endproduct (e.g, an mRNA, guide RNA, or a protein) in either precursor or mature form.Agent Ref. No. P14786WOOO

[0139] As used herein, the phrase “hypomorphic allele” refers to an allele of a gene with less gene activity than a wild-type allele but more gene activity than an amorphic allele.

[0140] As used herein, the terms “include,” “includes,” and “including” are to be construed as at least having the features to which they refer while not excluding any additional unspecified features.

[0141] As used herein, the term “isomorphic allele” refers to an allele of a gene having wildtype gene activity.

[0142] The term “isolated” as used herein means having been removed from its natural environment.

[0143] As used herein, the term “introduced” means providing a nucleic acid (e.g., expression construct) or protein into a cell. “Introduced” includes reference to the incorporation of a nucleic acid into a eukaryotic or prokaryotic cell where the nucleic acid may be incorporated into the genome of the cell and includes reference to the transient provision of a nucleic acid or protein to the cell. “Introduced” includes reference to stable or transient transformation methods. Thus, “introduced” in the context of inserting a nucleic acid fragment (e.g., a recombinant DNA construct / expression construct) into a cell, means “transfection” or “transformation” or “transduction” and includes reference to the incorporation of a nucleic acid fragment into a eukaryotic or prokaryotic cell where the nucleic acid fragment may be incorporated into the genome of the cell (e.g., nuclear chromosome, plasmid, plastid, chloroplast, or mitochondrial DNA), converted into an autonomous replicon, or transiently expressed (e.g., transfected mRNA).

[0144] As used herein, a “loss-of-function allele” can include an amorphic allele or a hypomorphic allele of a gene.

[0145] As used herein, the term “plant” includes a whole plant and any descendant, cell, tissue, part, or parts of the plant. The term “plant” thus includes reference to an immature or mature whole plant, including a plant from which seed or grain or anthers have been removed.

[0146] The term “plant part” includes any part(s) of a plant, including, for example and without limitation: seed (including mature seed and immature seed); grain; stover; a plant cutting; a plant cell; a plant cell culture; or a plant organ (e.g, pollen, embryos, pods; flowers, fruits, shoots, leaves, roots, stems, and explants). A plant tissue or plant organ may be a seed, protoplast, callus, or any other group of plant cells that is organized into a structural orAgent Ref. No. P14786WOOO functional unit. A plant cell or tissue culture may be capable of regenerating a plant having the physiological and morphological characteristics of the plant from which the cell or tissue was obtained, and of regenerating a plant having substantially the same genotype as the plant. Regenerable cells in a plant cell or tissue culture may be embryos, protoplasts, meristematic cells, callus, pollen, leaves, anthers, roots, root tips, flowers, or stalks. In contrast, some plant cells are not capable of being regenerated to produce plants and are referred to herein as “non-regenerable” plant cells.

[0147] As used herein, the term '‘exemplary” refers to an example, an instance, or an illustration, and does not indicate a most preferred embodiment unless otherwise stated.

[0148] As used herein, the term “INDELS” refers to a DNA insertion, deletion, and / or substitution.

[0149] To the extent to which any of the preceding definitions is inconsistent with definitions provided in any patent or non-patent reference incorporated herein by reference, any patent or non-patent reference cited herein, or in any patent or non-patent reference found elsewhere, it is understood that the preceding definition will be used herein.

[0150] It is to be understood that any guide RNA specifically identified herein with a letter and number combination (such as CR1808, CR1954, etc.) refers to a unique guide RNA having a unique spacer RNA sequence. For example, CR1808 and CR1954 are different guide RNAs having different spacer RNA sequences.

[0151] Referring now to the figures. Figure 1 shows a block diagram of a system 10.According to some embodiments, the system 10 can comprise a cyberinfrastructure 100. The cyberinfrastructure 100 can comprise a memory unit 102, executable instructions 104, a processing unit 106, a database 108, a communication module 110, and / or a human-machine interface (HMI) 112. According to some embodiments, the system 10 can further comprise an experimental plant, expenmental plant part, and / or an experimental plant cell 114. According to some embodiments, the system 10 can further comprise DNA extraction, DNA amplification, and / or DNA sequencing device(s) 116.

[0152] The memory unit 102 can be and / or comprise any suitable computer memory and / or storage unit. The memory unit 102 can include, according to some embodiments, a program storage area and / or data storage area. The memory unit 102 can comprise read-only memory (“ROM”, an example of non-volatile memory, meaning it does not lose data when it is not connected to a power source) and / or random-access memory (“RAM”, an example of volatileAgent Ref. No. P14786WOOO memory, meaning it will lose its data when not connected to a power source). Nonlimiting examples of volatile memory include static RAM (“SRAM”), dynamic RAM (“DRAM”), synchronous DRAM (“SDRAM”), etc. Examples of non-volatile memory include electrically erasable programmable read only memory (“EEPROM”), flash memory, hard disks, SD cards, etc.

[0153] According to some embodiments, the memory unit 102 can be and / or comprise at least one non-transitory computer-readable medium. In communications and computing, a computer readable medium is a medium capable of storing data in a format readable by a mechanical device. The term “non-transitory” is used herein to refer to computer readable media (“CRM”) that store data for short periods or in the presence of power such as a memory device. According to some embodiments, the non-transitory computer readable medium can be a tangible non- transitory computer readable medium.

[0154] The memory unit 102 can be used to store executable instructions 104. When executed, the executable instructions 104 cause the system 10 and / or the cyberinfrastructure 100 described herein to perform any method(s) and / or methodolog(ies) described herein. The instructions 104 can be stored, completely or at least partially, within the memory unit 102 and / or any other aspect of the cyberinfrastructure 100. When executed, the executable instructions 104 can comprise steps capable of performing any of the methods described herein including, but not limited to, at least some aspect(s) of the method 200, the method 300, the method 400, and / or the method 500. For example, the linear model and / or the protein language model and / or algorithm of the method 500 described herein can be part of the cyberinfrastructure 100 such that the cyberinfrastructure 100 can execute the linear model and / or the protein language model and / or algorithm.

[0155] As shown in Figure 1, the cyberinfrastructure 100 further includes a processing unit 106. The processing unit 106 can be operatively connected to the memory unit 102 and can execute software instructions that are capable of being stored in a RAM of the memory (e.g., during execution), a ROM of the memory (<?.g., on a generally permanent basis), or any sort of non-transitory computer readable medium such as another memory or a disc. For example, the processing unit 106 can be configured to perform, run, carry out, and / or otherwise execute the executable instructions 104.

[0156] The processing unit 106 can be an electronic circuit which performs operations on some external data source, usually memory or some other data stream. The processing unit 106 canAgent Ref. No. P14786WOOO be and / or comprise any number of processors ranging from 1 to N where N is any number greater than 1. Non-limiting examples of processors include a processor, microprocessor, a controller, a microcontroller, an arithmetic logic unit (‘‘ALU7’), a graphics processing unit (‘’GPU”), and most notably, a central processing unit (’‘CPU”). A CPU, also called a central processor or main processor, is the electronic circuitry within a computer that carries out the instructions of a computer program by performing the basic arithmetic, logic, controlling, and input / output (“I / O”) operations specified by the instructions. Processing units are common in tablets, telephones, handheld devices, laptops, user displays, smart devices (TV, speaker, watch, etc.), and other computing devices. The processing unit 106 can further include components for establishing communications. The processing unit 106 can also include other components and can be implemented partially or entirely on a semiconductor (e g., a field- programmable gate array (“FPGA”)) chip, such as a chip developed through a register transfer level (“RTL”) design process.

[0157] According to some embodiments, GPU-based computing can be used in one or more aspects. GPU-based computing refers to the practice of using a GPU simultaneously with one or more central processing units (CPUs) and / or GPUs. GPU-based computing allows for a sort of parallel processing between the GPU and the one or more CPUs and / or GPUs such that the GPU can take on some of the computational load to increase speed and efficiency. Additionally, GPUs commonly have a much higher number of processing cores than a traditional CPU, which allows a GPU to be able to process pictures, images, and / or graphical data faster than a traditional CPU.

[0158] According to some embodiments, the non-transitory computer readable medium operates under control of an operating system stored in a memory, such as the memory 102. The non-transitory computer readable medium implements a compiler which allows a software application written in a programming language such as COBOL. C++. FORTRAN, or any other known programming language to be translated into code readable by the central processing unit. After completion, a central processing unit, such as the processing unit 106, accesses and manipulates data stored in the memory of the non-transitory computer readable medium using the relationships and logic dictated by a software application and generated using the compiler.

[0159] According to some embodiments, the software application and the compiler are tangibly embodied in the computer-readable medium. When instructions, such as the executable instructions 104, are read and executed by the non-transitory computer readableAgent Ref. No. P14786WOOO medium, the non-transitory computer readable medium performs the steps necessary to implement and / or use aspects of the present disclosure. A software application, operating instructions, and / or firmware (semi-permanent software programmed into read-only memory) may also be tangibly embodied in the memory and / or data communication devices, thereby making the software application a product or article of manufacture according to the present disclosure.

[0160] As shown in Figure 1. the cyberinfrastructure 100 can further include a database 108. The database 108 can be a structured set of data held and / or stored such that it can be accessed by the cyberinfrastructure 100. The database 108, as well as data and information contained therein, need not reside in a single physical or electronic location. For example, the database 108 may reside, at least in part, on a local storage device, in an external hard drive, on a database server connected to a network, on a cloud-based storage system, in a distributed ledger (such as those commonly used with blockchain technology), and the like.

[0161] According to some embodiments, the database 108 can store information which includes data related to target gene editing frequencies, proportions of explants and / or plants with intended or high quality gene edits, off-target editing frequencies, and / or gene-editing outcomes, and the like which are obtained with candidate gRNAs and Cas nucleases in protoplasts, callus, embryogenic callus, plants, and the like.

[0162] As shown in Figure 1, the cyberinfrastructure 100 can further include a communications module 110. The communications module 110 can be configured to be able to send data to and / or receive data from various components within the cyberinfrastructure 100. Additionally, the communications module 110 can be configured to be able to send data to and / or receive data from entities external to the cyberinfrastructure 100. For example, the communications module 110 can connect to a third-party entity to receive and / or send data from / to said entity.

[0163] The communications module 110 can include any combination of modem(s), router(s), access point(s), bridge(s), gateway(s), hub(s), repeater(s), switch(es), transceiver(s), and the like in order to facilitate communication. The communications module 110 can be configured to perform data communication wirelessly and / or in a wired fashion. The communications module 110 can include one or more communications ports such as Ethernet, serial advanced technology attachment (“SATA”), universal serial bus (“USB”), or integrated drive electronics (“IDE”), for transferring, sending, receiving, and / or or storing data.Agent Ref. No. P14786WOOO

[0164] According to some embodiments, the communications module 110 and / or other components of the cyberinfrastructure 100 are able to perform data communication either within the cyberinfrastructure 100 and / or externally of the cyberinfrastructure 100 in a wireless fashion using any sort of wireless connection device and / or protocol. This can include, but is not limited to, Bluetooth, Wi-Fi, cellular data, radio waves, satellite, and / or generally any other form of wireless connection. Therefore, the communications module 110 and / or any other component(s) of the cyberinfrastructure 100 will include generally any electronic components necessary to allow for such wireless communication.

[0165] According to some embodiments, the communications module 110 and / or other components of the cyberinfrastructure 100 are able to perform data communication either within the cyberinfrastructure 100 and / or externally of the cyberinfrastructure 100 via a wired connection. Wired communication can take the form of CAN bus, Ethernet, co-axial cable, fiber optic line, and / or generally any other device and / or protocol which will allow for wired communication. Therefore, the communications module 110 and / or any other component(s) of the cyberinfrastructure 100 will include generally any electronic components necessary’ to allow for such wired communication.

[0166] According to some embodiments, the communications module 110 and / or other components of the cyberinfrastructure 100 are able to perform data communication either within the cyberinfrastructure 100 and / or externally of the cyberinfrastructure 100 via a network. According to some embodiments, the network is, by way of example only, a wide area network (“WAN”) such as a TCP / IP based network or a cellular network, a local area network (“LAN”), a neighborhood area network (“NAN”), a home area network (“HAN”), or a personal area network (“PAN”) employing any of a variety of communication protocols, such as Wi-Fi. Bluetooth, ZigBee, near field communication (“NFC”), etc., although other types of networks are possible and are contemplated herein. Communications through the network can be protected using one or more encryption techniques, such as those techniques provided by the Advanced Encry ption Standard (AES), which superseded the Data Encry ption Standard (DES), the IEEE 802. 1 standard for port-based network security, pre-shared key, Extensible Authentication Protocol (“EAP”), Wired Equivalent Privacy (“WEP”), Temporal Key Integrity Protocol (“TKIP”), Wi-Fi Protected Access (“WPA”), and the like.

[0167] Ethernet is a family of computer networking technologies commonly used in local area networks (“LAN”), metropolitan area networks (“MAN”), and wide area networks (“WAN”).Agent Ref. No. P14786WOOOSystems communicating over Ethernet divide a stream of data into shorter pieces called frames. Each frame contains source and destination addresses, and error-checking data so that damaged frames can be detected and discarded; most often, higher-layer protocols trigger retransmission of lost frames. As per the OSI model, Ethernet provides services up to and including the data link layer. Ethernet was first standardized under the Institute of Electrical and Electronics Engineers (“IEEE"’) 802.3 working group / collection of IEEE standards produced by the working group defining the physical layer and data link layer’s media access control (" AC”) of wired Ethernet. Ethernet has since been refined to support higher bit rates, a greater number of nodes, and longer link distances, but retains much backward compatibility. Ethernet has industrial application and interworks well with Wi-Fi. The Internet Protocol (“IP”) is commonly carried over Ethernet and so it is considered one of the key technologies that make up the Internet.

[0168] The Internet Protocol (“IP”) is the principal communications protocol in the Internet protocol suite for relay ing datagrams across network boundaries. Its routing function enables internetworking, and essentially establishes the Internet. IP has the task of delivering packets from the source host to the destination host solely based on the IP addresses in the packet headers. For this purpose, IP defines packet structures that encapsulate the data to be delivered. It also defines addressing methods that are used to label the datagram with source and destination information.

[0169] The Transmission Control Protocol (“TCP”) is one of the main protocols of the Internet protocol suite. It originated in the initial network implementation in which it complemented the IP. Therefore, the entire suite is commonly referred to as TCP / IP. TCP provides reliable, ordered, and error-checked delivery of a stream of octets (bytes) between applications running on hosts communicating via an IP network. Major internet applications such as the World Wide Web. email, remote administration, and file transfer rely on TCP, which is part of the Transport Layer of the TCP / IP suite.

[0170] Transport Layer Security, and its predecessor Secure Sockets Layer (“SSL / TLS”), often runs on top of TCP. SSL / TLS are cryptographic protocols designed to provide communications security over a computer network. Several versions of the protocols find widespread use in applications such as web browsing, email, instant messaging, and voice over IP (“VoIP”). Websites can use TLS to secure all communications between their servers and web browsers.Agent Ref. No. P14786WOOO

[0171] As shown in Figure 1, the cyberinfrastructure 100 can further comprise a human machine interface (HMI) 112. The HMI 112, which can also be referred to as a user interface and / or a visualization portal, is how a user interacts with a machine. The HMI 112 can be a digital interface, a command-line interface, a graphical user interface ('‘GUI”), oral interface, virtual reality interface, or any other way a user can interact with a machine (user-machine interface). For example, the HMI 112 can include a combination of digital and analog input and / or output devices or any other type of user interface input / output device required to achieve a desired level of control of, interaction with, and / or monitoring of a device. Nonlimiting examples of input and / or output devices include computer mice, keyboards, touchscreens, knobs, dials, toggles, levers, sliders, switches, buttons, speakers, microphones, printers, LIDAR, RADAR, etc.

[0172] The human machine interface 112 can include a display, which can act as an input and / or output device. More particularly, the display can be a liquid crystal display (“LCD”), a light-emitting diode (“LED”) display, an organic LED (“OLED”) display, an electroluminescent display (“ELD”), a surface-conduction electron emitter display (“SED”), a field-emission display (“FED”), a thin-film transistor (“TFT”) LCD. a bistable cholesteric reflective display (z.e., e-paper), a touch-screen display, etc.

[0173] A user can use the HMI 112 to input information and / or data into the cyberinfrastructure 100. Inputting information into the cyberinfrastructure 100 via the HMI 112 can include modifying information and / or data displayed via the HMI 112 and / or entering new information and / or data. Input(s) received via the HMI 112 can be processed via the cyberinfrastructure 100 and / or components thereof such as the processing unit 106. The HMI 112 can be used by the cyberinfrastructure 100, and / or any components thereof, to output and / or display information, data, text, graphics, graphs, charts, toggles, levers, sliders, tabs, and the like. For example, according to some embodiments, a user can make decisions regarding predictions of target gene editing frequency score and functional effect(s) of one or more INDELS using the HMI 112. Additionally or alternatively, according to some embodiments, a user can make selections regarding guide RNAs using the HMI 112. Additionally or alternatively, the HMI 112 can be used to display any information and / or data that is the same and / or similar to that shown in any of Figures 6-27.

[0174] According to some embodiments, the cyberinfrastructure 100 can be implemented and / or accessed as a downloadable computer application to be stored on a device such as aAgent Ref. No. P14786WOOO computer, laptop, phone, tablet, smart device, and the like. According to some embodiments, the cyberinfrastructure 100 can be implemented and / or accessed online via cloud computing. According to some embodiments, the cyberinfrastructure 100 can be implemented via cloud computing as a Software as a Service (SaaS), Platform as a Service (PaaS), and / or Infrastructure as a Service (laaS). The cyberinfrastructure 100 could utilize any sort of cloud computing deployment model such as a private cloud, a community cloud, a public cloud, and / or a hybrid cloud.

[0175] Cloud computing is a model of sendee delivery for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and sen ices) that can be rapidly provisioned and released with minimal management effort or interaction with a provider of the service.

[0176] A cloud computing environment is service oriented with a focus on statelessness, low coupling, modularity, and semantic interoperability. At the heart of cloud computing is an infrastructure comprising a network of interconnected nodes.

[0177] All components of the cyberinfrastructure 100 including, but not limited to. the memory unit 102, the executable instructions 104, the processing unit 106, the database 108, the communication module 110, and / or the HMI 112 can be operatively connected, by a common bus and / or any other suitable connection element, such that all components of the cyberinfrastructure 100 can be in communication with each other.

[0178] As shown in Figure 1, the system 10 can further comprise an experimental plant, an experimental plant part, and / or an experimental plant cell 114. As noted below with regard to the method 500, the experimental plant, experimental plant part, and / or experimental plant cell 114 can lack a particular and / or desired INDELS. The experimental plant, experimental plant part, and / or experimental plant cell 114 can be manipulated such that (a) (i) a selected preferred gRNA or a polynucleotide encoding a selected preferred gRNA; and (ii) a Cas nuclease which binds the selected preferred gRNA or a polynucleotide encoding the Cas nuclease are introduced into the experimental plant, experimental plant part, and / or experimental plant cell 114; or (b) a ribonucleoprotein complex comprising the Cas nuclease bound to the selected preferred gRNA is introduced into the experimental plant, experimental plant part, and / or experimental plant cell 114. A plant or plant part comprising the particularAgent Ref. No. P14786WOOO and / or desired INDELS can then be obtained via the experimental plant, experimental plant part, and / or experimental plant cell 114.

[0179] As shown in Figure 1. the system 10 can include DNA extraction, DNA amplification, and / or DNA sequencing device(s) 116. The DNA extraction, DNA amplification, and / or DNA sequencing device(s) 116 can be any off-the-shelf device(s) known in the art. The DNA extraction, DNA amplification, and / or DNA sequencing device(s) 116 can be configured to isolate, amplify, and / or sequence genomic DNA sequences containing a target plant gene from an experimental plant cell that had received a distinct candidate guide RNA and a Cas nuclease encoding DNA vector, wherein the isolated, amplified, and / or sequenced genomic DNA sequences can be used for predicting target gene editing frequency score and / or for predicting functional effect(s) of one or more INDELS.

[0180] According to some embodiments, the experimental plant, experimental plant part, and / or an experimental plant cell 114 and / or the DNA extraction, DNA amplification, and / or DNA sequencing device(s) 116 can be used for at least one step of the method 500 described herein.

[0181] Figure 2 shows a flow chart depicting a method 200 that involves predicting aspect(s) of gene editing and also involves selection of guide RNA based on said prediction(s). For example, according to some embodiments, the method 200 can be a method for selection of a preferred guide RNA for obtaining a plant or plant part comprising a desired DNA insertion, deletion, and / or substitution (INDELS) in a target plant gene via Cas nuclease mediated gene editing. As shown in Figure 2, the method 200 includes the step 202 of predicting a target gene editing frequency score for each of a plurality of distinct candidate guide RNAs (gRNAs) independently introduced into a plant cell with a Cas nuclease encoding DNA vector, wherein the candidate gRNAs are not associated with the Cas nuclease and the Cas nuclease is not present in the plant cell prior to their introduction.

[0182] The method 200 shown in Figure 2 further includes the step 204 of selecting a preferred gRNA from the plurality of distinct candidate gRNAs for use in obtaining a plant or plant part comprising a particular and / or desired INDELS in a target plant gene based on the target gene editing efficiency score.

[0183] According to some embodiments, at least one of the steps of the method 200 can be performed, at least partially, using the system 10 and / or any component(s) thereof.Agent Ref. No. P14786WOOO

[0184] Figure 3 shows a flow chart depicting a method 300 that involves predicting aspect(s) of gene editing and also involves selection of guide RNA based on said predict! on(s). For example, according to some embodiments, the method 300 can be a method for selection of a preferred guide RNA for obtaining a plant or plant part comprising a desired DNA insertion, deletion, and / or substitution (INDELS) in a target plant gene encoding a protein via Cas nuclease mediated gene editing. As shown in Figure 3, the method 300 includes the step 302 of predicting a functional effect of an INDELS obtained in a target plant gene via Cas nuclease mediated gene editing for each of a plurality of distinct candidate guide RNAs (gRNAs) independently introduced into a plant cell with a Cas nuclease encoding DNA vector, optionally wherein the candidate gRNAs are not associated with the Cas nuclease and the Cas nuclease is not present in the plant cell prior to their introduction. While the step 302 comprises predicting a functional effect, according to some embodiments, the step 302 can comprise predicting one or more functional effects of an INDELS wherein the number of functional effects can range from 1 to N where N is any number greater than 1. According to some embodiments, the step 302 can comprise predicting functional effect(s) of one or more INDELS wherein the number of INDELS can range from 1 to N where N is any number greater than 1.

[0185] The method 300 shown in Figure 3 further includes the step 304 of selecting a preferred gRNA based on the predicted functional effect of the INDELS. According to some embodiments, the predicted functional effect(s) of the INDELS generated by candidate gRNAs and / or the editing efficiency of candidate gRNAs guide RNA can be used to select preferred gRNAs and / or to predict in-planta information. According to some embodiments, such in-planta information can comprise phenoty pic information (e.g., traits conferred on a plant by a given INDELS). Phenotype information includes predicted functional effects of one or more INDELS on plant yield (e.g, grain yield per unit area), biotic stress tolerance (e.g, tolerance to insects, nematodes, fungal pathogens, viral pathogens, and / or bacterial pathogens), and / or abiotic stress tolerance (e.g., tolerance to heat, cold, water surplus, and / or drought stress), herbicide tolerance, improved utilization of nutrients or w ater, modified lipid, carbohydrate, or protein composition, improved flavor or appearance, improved storage characteristics (e.g., resistance to bruising, browning, or softening), and / or altered morphology e.g., floral architecture or color, plant height, branching, root structure) in comparison to a control plant lacking the one or more INDELS.Agent Ref. No. P14786WOOO

[0186] According to some embodiments, at least one of the steps of the method 300 can be performed, at least partially, using the system 10 and / or any component(s) thereof.

[0187] Figure 4 shows a flow chart depicting a method 400 that involves predicting aspect(s) of gene editing and also involves selection of guide RNA based on said prediction(s). For example, according to some embodiments, the method 400 can be a method for selection of a preferred guide RNA for obtaining a plant or plant part comprising a desired DNA insertion, deletion, and / or substitution (INDELS) in a target plant gene encoding a protein via Cas nuclease mediated gene editing. The method 400 includes the step 402 of predicting a target gene editing frequency score for each of a plurality of distinct candidate guide RNAs (gRNAs) independently introduced into a plant cell with a Cas nuclease encoding DNA vector, optionally wherein the candidate gRNAs are not associated with the Cas nuclease and the Cas nuclease is not present in the plant cell prior to their introduction.

[0188] The method 400 shown in Figure 4 further includes the step 404 of predicting a functional effect of an INDELS obtained in a target plant gene via Cas nuclease mediated gene editing for each of the plurality of distinct candidate gRNAs.

[0189] The method 400 of Figure 4 further includes the step 406 of selecting a preferred gRNA based on the target gene editing frequency score and / or the predicted functional effect of the INDELS.

[0190] According to some embodiments, at least one of the steps of the method 400 can be performed, at least partially, using the system 10 and / or any component(s) thereof.

[0191] Figures 5A-G show a flow chart of a method 500 that involves predicting aspect(s) of gene editing and also involves selection of guide RNA based on said prediction(s). For example, according to some embodiments, the method 500 can be a method for selection of a preferred guide RNA for obtaining a plant or plant part comprising a desired DNA insertion, deletion, and / or substitution (INDELS) in a target plant gene encoding a protein via Cas nuclease mediated gene editing. The method 500 includes the step 502 of predicting a target gene editing frequency score for each of a plurality of distinct candidate guide RNAs (gRNAs) independently introduced into a plant cell with a Cas nuclease encoding DNA vector, wherein the candidate gRNAs are not associated with the Cas nuclease and the Cas nuclease is not present in the plant cell prior to their introduction. This step 502 of predicting a target gene editing frequency score for each of a plurality of distinct candidate guide RNAs can comprise steps 503, 505. 507, and / or 509 of Figure 5A.Agent Ref. No. P14786WOOO

[0192] The step 503 comprises independently introducing a plurality of distinct candidate guide RNAs into a plant cell with a Cas nuclease encoding DNA vector, wherein the candidate gRNAs are not associated with the Cas nuclease and the Cas nuclease is not present in the plant cell prior to their introduction. According to some embodiments, the candidate gRNAs and the Cas nuclease encoding DNA vector are introduced into at least one plant protoplast cell via PEG-mediated transfection and / or electroporation. The RNA molecules comprising the candidate gRNAs and the Cas nuclease encoding DNA vector can be introduced via a hybrid delivery assay. The hybrid delivery assay requires no protein purification or ribonucleoprotein (RNP) pre-complexing, but instead relies on the protein expression machinery' in the host cell to generate the Cas nuclease which then binds the gRNA to form the Cas nuclease-gRNA complex. Since the crRNA is provided as an RNA molecule within the cell, it is unnecessary to construct and introduce expression vectors encoding the gRNAs, which would be a bottleneck to guide RNA testing at scale. Without seeking to be limited by theory', it is believed that reliance on the host cell to generate the Cas nuclease from a DNA expression vector and the in vivo complexing of the Cas nuclease with introduced gRNAs to form RNPs (Cas nuclease complexed with crRNA of the gRNA) is similar to the pressures put on the CRISPR / Cas system in a whole plant gene editing system (e,g., a system for generating whole plants comprising gene edits). It is thus believed that the gene editing results obtained in plant cells subjected to hybrid delivery assays can be used for evaluation and selection of gRNAs which will be effective in whole plant gene editing systems (e,g. “plasmid transformation in planta’’ in FIG. 6).

[0193] Various treatments can be used for delivery of gene editing molecules to a plant cell. In the hybrid delivery' method, the gene editing molecules comprise a DNA molecule (e.g., plasmid) encoding a Cas nuclease and an RNA molecule comprising a guide RNA molecule which is recognized by the Cas nuclease. In in planta plasmid transformation methods, the gene editing molecules comprise one or more DNA molecules (e.g., plasmids) encoding both a Cas nuclease and an RNA molecule comprising a guide RNA molecule which is recognized by the Cas nuclease. In certain embodiments, one or more treatments is employed to deliver the gene editing molecules into a plant cell, e.g., through barriers such as a cell wall, a plasma membrane, a nuclear envelope, and / or other lipid bilayer. In certain embodiments, a composition comprising the gene editing molecules are delivered directly, for example by direct contact of the composition with a plant cell. Aforementioned compositions can beAgent Ref. No. P14786WOOO provided in the form of a liquid, a solution, a suspension, an emulsion, a reverse emulsion, a colloid, a dispersion, a gel. liposomes, micelles, an injectable material, an aerosol, a solid, a powder, a particulate, a nanoparticle, or a combination thereof can be applied directly to a plant, plant part, plant cell, or plant explant (e.g, through abrasion or puncture or otherwise disruption of the cell wall or cell membrane, by spraying or dipping or soaking or otherwise directly contacting, by microinjection). For example, a plant cell or plant protoplast is soaked in a liquid genome editing molecule-containing composition. In certain embodiments, the composition is delivered using negative or positive pressure, for example, using vacuum infiltration or application of hydrodynamic or fluid pressure. In certain embodiments, the composition is introduced into a plant cell or plant protoplast, e.g., by microinjection or bydisruption or deformation of the cell wall or cell membrane, for example by physical treatments such as by application of negative or positive pressure, shear forces, or treatment with a chemical or physical delivery agent such as surfactants, liposomes, or nanoparticles; see, e.g., delivery- of materials to cells employing microfluidic flow through a cell-deforming constriction as described in US Published Patent Application 2014 / 0287509, incorporated byreference in its entirety herein. Other techniques useful for delivering the composition to a eukaryotic cell, plant cell or plant protoplast include: ultrasound or sonication; vibration, friction, shear stress, vortexing, cavitation; centrifugation or application of mechanical force; mechanical cell wall or cell membrane deformation or breakage; enzy matic cell wall or cell membrane breakage or permeabilization; abrasion or mechanical scarification (e.g., abrasion with carborundum or other particulate abrasive or scarification with a file or sandpaper) or chemical scarification (e.g, treatment with an acid or caustic agent); and electroporation. In certain embodiments, the composition is provided by bacterially mediated (e.g., Agrobacterium sp., Rhizobium sp., Sinorhizobium sp., Mesorhizobium sp.. Bradyrhizobium sp.. Azobacter sp.. Phyllobacterium sp.) transfection of the plant cell or plant protoplast with a polynucleotide encoding the genome editing molecule(s); see, e.g., Broothaerts et al. (2005) Nature, 433:629 - 633). Any of these techniques or a combination thereof are alternatively- employed on a plant explant, plant part or tissue or intact plant (or seed) from which a plant cell is optionally subsequently obtained or isolated; in certain embodiments, the composition is delivered in a separate step after the plant cell has been isolated. Delivery of polynucleotides and gene editing molecules to plants and protoplasts in particular by PEG mediated transfection is disclosed in US Patent No. 11,634,722, which is incorporated hereinAgent Ref. No. P14786WOOO by reference in its entirety, and in Tonnies et al., 2023, DOI: 10.3791 / 64991. Examples of soy in planta transformation methods which can be used include those disclosed in US Patent Application No. 11,453,885 and WO 2024 / 015781. which are each incorporated herein by reference in their entireties.

[0194] Gene editing molecules used herein include an RNA directed DNA endonuclease (RdDe) and a guide RNA recognized by the RdDE as well as polynucleotides (e.g.. DNA) encoding the gene editing molecules. RNA directed DNA endonucleases (RdDe) systems comprising both an RdDE and a include a class 1 CRISPR type nuclease system, a class 2 type II Cas nuclease, a Cas9, a nCas9 nickase, a class 2 type V Cas nuclease, a Casl2a nuclease, a nCas 12a nickase, a Casl2d (CasY), a Casl2e (CasX), a Casl2b (C2cl), a Casl2c (C2c3), a Casl2i, a Casl2j, and a Casl4 nuclease systems. CRISPR technology for editing the genes of eukaryotes is disclosed in US Patent Application Publications 2016 / 0138008A1 and US2015 / 0344912A1, and in US Patents 8,697,359; 8,771,945; 8,945,839; 8,999,641; 8,993,233; 8,895,308; 8,865,406; 8,889,418; 8,871,445; 8,889,356; 8,932,814; 8,795,965; and 8,906,616. Cpfl endonuclease and corresponding guide RNAs and PAM sites are disclosed in US Patent Application Publication 2016 / 0208243 Al. Plant RNA promoters for expressing CRISPR guide RNA and plant codon-optimized CRISPR Cas9 endonuclease are disclosed in International Patent Application WO 2015 / 131101. Methods of using CRISPR technology for genome editing in plants are disclosed in US Patent Application Publications US 2015 / 0082478A1 and US 2015 / 0059010A1 and in International Patent Application PCT / US2015 / 038767 Al (published as WO 2016 / 007347 and claiming priority to US Provisional Patent Application 62 / 023,246). All of the patent publications referenced in this paragraph are incorporated herein by reference in their entirety.

[0195] Guide RNAs (sgRNAs or crRNAs and a tracrRNA) form an RNA-guided endonuclease / guide RNA complex which can specifically bind sequences in the gDNA target site that are adjacent to a protospacer adjacent motif (PAM) sequence. The ty pe of RNA- guided endonuclease typically informs the location of suitable PAM sites and design of crRNAs or sgRNAs. G-rich PAM sites, e.g., 5’-NGG are typically targeted for design of crRNAs or sgRNAs used with Type II Cas nucleases (e.g., Cas9 proteins). Examples of PAM sequences include 5’-NGG (Streptococcus pyogenes), 5’-NNAGAA (Streptococcus thermophilus CRISPR1), 5’-NGGNG (Streptococcus thermophilus CRISPR3), 5’-NNGRRT or 5’-NNGRR (Staphylococcus aureus Cas9, SaCas9), and 5’-NNNGATT (NeisseriaAgent Ref. No. P14786WOOO meningitidis). T-rich PAM sites (e g., 5'-TTN or 5’-TTTV, where "V" is A, C, or G) are typically targeted for design of crRNAs or sgRNAs used with Type V Cas nucleases (e.g, Casl2a proteins). In some instances. Casl2a can also recognize a 5?-CTA PAM motif. Other examples of potential Casl2a PAM sequences include TTN, CTN, TCN, CCN, TTTN, TCTN, TTCN, CTTN, ATTN, TCCN, TTGN, GTTN, CCCN, CCTN, TTAN, TCGN, CTCN, ACTN, GCTN, TCAN, GCCN, and CCGN (wherein N is defined as any nucleotide). Cpfl endonuclease and corresponding guide RNAs and PAM sites are disclosed in US Patent Application Publication 2016 / 0208243 Al, which is incorporated herein by reference for its disclosure of DNA encoding Cpfl endonucleases and guide RNAs and PAM sites. At least 1 or 17 nucleotides of gRNA spacer sequence are required by Cas9 for DNA cleavage to occur; for Cpfl at least 16 nucleotides of gRNA spacer sequence are needed to achieve detectable DNA cleavage and at least 18 nucleotides of gRNA sequence were reported necessary for efficient DNA cleavage in vitro; see Zetsche et al. (2015) Cell, 163:759 - 771. In practice, guide RNA spacer sequences are generally designed to have a length of 17 - 24 nucleotides (frequently 19. 20. or 21 nucleotides) and exact complementarity (i.e., perfect base-pairing) to the targeted gene or nucleic acid sequence; guide RNAs having less than 100% complementarity to the target sequence can be used (e , a gRNA with a length of 20 nucleotides and 1 - 4 mismatches to the target sequence) but can increase the potential for off-target effects. Certain aspects of the design of effective guide RNAs for use in plant genome editing is disclosed in US Patent Application Publication 2015 / 0082478 Al. the entire specification of which is incorporated herein by reference. Efficient gene editing has been achieved using a chimeric “single guide RNA” (“sgRNA”), an engineered (synthetic) single RNA molecule that mimics a naturally occurring crRNA-tracrRNA complex and contains both a tracrRNA (for binding the nuclease) and at least one crRNA (to guide the nuclease to the sequence targeted for editing); see, for example. Cong et al. (2013) Science, 339:819 - 823; Xing et al. (2014) BMC Plant Biol., 14:327 - 340.

[0196] The step 505 comprises filtering a set of alleles to remove allele(s) outside an expected editing window and / or to remove false positives. Additionally or alternatively, the step 505 can comprise analyzing plants wherein a percent of plants with at least one edited allele over 25% editing efficiency are desired.Agent Ref. No. P14786WOOO

[0197] The step 507 comprises determining: (i) an average editing efficiency in the plant cell for a subset of test gRNAs directed to a target gene, and (ii) an actual editing outcome in the plant or plant part for the subset of test gRNAs.

[0198] The step 509 comprises applying a linear model to correlate the average editing efficiency in the plant cell with the actual editing outcome in the plant or plant part for the subset of test gRNAs. According to at least some embodiments, the step 509 can further include the linear model being configured to apply least squares regression to the correlation of the average editing efficiency in the plant cell and the actual editing outcome in the plant or plant part for each test gRNA. The step 509 can further include the linear model being configured to derive an equation to determine a line of best fit of the correlation of the average editing efficiency in the plant cell and the actual editing outcome in the plant or plant part of each test gRNA based, at least in part, on the least squares regression. The step 509 can further include the linear model being configured to predict editing efficiency of a candidate gRNA using the derived equation.

[0199] As shown in Figure 5B, the method 500 can include the step 510 of predicting a functional effect of an INDELS obtained in a target plant gene via Cas nuclease mediated gene editing for each of the plurality of distinct candidate gRNAs. This step 510 can comprise the steps 511, 513, 515, 517, 519, 521, 523, 525, 526, 527, 529, 531, 533, and / or 535.

[0200] The step 511 comprises selecting a subset of gRNAs based on their target gene editing frequency score. In certain embodiments, target gene editing frequency scores obtained from a population of candidate gRNAs are ranked from the highest scores to the lowest scores and the subset of candidate gRNAs (e.g, top 2, 3, 5, or 10) are selected.

[0201] The step 513 comprises training a protein language model that utilizes machine learning and / or artificial intelligence (Al). According to some embodiments, the protein language model can be the ESM-lv model described herein and / or the PROVEAN model and / or algorithm described herein. The protein language model can be a transformer-based foundation model according to some embodiments, how ever any suitable type of model could be used. According to some embodiments, the protein language model can be pretrained in a self-supervised manner on a large and diverse set of natural proteins such as UniRef50, UniRef90, UniRefl 00, and / or any other suitable set of protein information. Additionally or alternatively, according to various embodiments, the model can be trained in other mannersAgent Ref. No. P14786WOOO such as via supervised learning, unsupervised learning, semi-supervised learning, reinforcement learning, and / or dimensionality reduction as well as other types.

[0202] According to some embodiments, Al can be used in one or more aspects of the present disclosure. Al is intelligence embodied by machines, such as computers and / or processors. While Al has many definitions, some have defined Al as utilizing machines and / or systems to mimic human cognitive ability such as decision-making and / or problem solving. Al has additionally been described as machines and / or systems that are capable of acting rationally such that they can discern their environment and efficiently and effectively take the necessary steps to maximize the opportunity to achieve a desired outcome. Goals of Al can include but are not limited to reasoning, problem-solving, knowledge representation, planning, learning, natural language processing, perception, motion and manipulation, social intelligence, and general intelligence. Al tools used to achieve these goals can include but are not limited to searching and optimization, logic, probabilistic methods, classification, statistical learning methods, artificial neural networks, machine learning, and deep learning.

[0203] In some embodiments, machine learning can be used in one or more aspects such as to train, improve, and / or drive the protein language model and / or algorithm. Machine learning is a subset of artificial intelligence. Machine learning aims to learn or train via training data in order to improve performance of a task or set of tasks. A machine learning algorithm and / or model can be developed such that it can be trained using training data to ultimately make predictions and / or decisions. Machine learning can include different approaches such as supervised learning, unsupervised learning, semi-supervised learning, reinforcement learning, and dimensionality reduction as well as other types. Supervised learning models are trained using training data that includes inputs and the desired output. This type of training data can be referred to as labeled data wherein the output provides a label for the input. The supervised learning model will be able to develop, through optimization or other techniques, a method and / or function that is used to predict the outcome of new inputs. Unsupervised learning models take in data that only includes inputs and engage in finding commonalities in the inputs such as grouping or clustering of aspects of the inputs. Thus, the training data for unsupervised learning does not include labeling and / or classification. Unsupervised learning models can make decisions for new data based on how alike or similar it is to existing data and / or to a desired goal. Examples of machine learning models include but are not limited to artificial neural networks, decision trees, support-vector machines, regression analysis,Agent Ref. No. P14786WOOOBayesian networks, and genetic algorithms. Examples of potential applications of machine learning include but are not limited to image segmentation and classification, ranking, recommendation systems, visual identity tracking, face verification, and speaker verification.

[0204] In some embodiments, deep learning can be used in one or more aspects. Deep learning is a subset of machine learning that utilizes a multi-layered approach. Examples of deep learning architectures include but are not limited to deep neural networks, deep belief networks, deep reinforcement learning, recurrent neural networks, and convolutional neural networks. Examples of fields wherein deep learning can be successfully applied include but are not limited to computer vision, speech recognition, natural language processing, machine translation, bioinformatics, medical image analy sis, and climate science. Deep learning models are commonly implemented as multi-layered artificial neural networks wherein each layer can be trained and / or can learn to transform particular aspects of input data into some sort of desired output.

[0205] The step 515 comprises integrating evolutionary information into the parameters of the protein language model. During training the model can be configured to leam a set of parameters that reflect integration of evolutionary information present in the diversity of the training set. According to some embodiments, such evolutionary information can be derived using evolutionary computation and / or evolutionary' algorithm(s).

[0206] The step 517 comprises inputting first and second protein sequence variants into the protein language model. According to some embodiments, the first and second variants can be chosen from variants predicted to be generated from the subset of gRNAs selected in the step 511.

[0207] According to various embodiments, the protein language model can function in a variety of different ways. For example, according to some embodiments, the protein language model can function according to steps 519, 521. and 523 of Figure 5C. The step 519 comprises utilizing the protein language model for: (i) masking a single amino acid position of the first sequence variant, (ii) predicting a probability distribution over all possible amino acids at the single amino acid position based on surrounding sequence context and the evolutionary information, (iii) determining a conditional probability of the single amino acid position based on the probability distribution, (iv) performing parts (i)-(iii) for each amino acid position of a length of the first sequence variant in a one-by-one manner, and (v)Agent Ref. No. P14786WOOO calculating a product of each conditional probability of each amino acid position of the length to produce an approximated joint probability of the first sequence variant.

[0208] The first protein sequence variant can be represented as S wherein the one masked amino acid is represented as Si. The surrounding sequence context can be represented as S-i and the evolutionary information integrated into the model parameters can be represented as E. The conditional probability of each masked amino acid position, as determined in part (iii) of the step 519, can be determined from the probability distribution represented as P(Si| S-i. E). The approximated joint probability of the first sequence variant, as determined in part (v) of the step 519, can be represented as P(S|E)=\S_bE). This approximated j oint probability can be interpreted as a pseudolikelihood of drawing the sequence from the dataset upon which the protein language model was trained.

[0209] The step 521 comprises utilizing the protein language model to perform parts (i)-(v) of the step 519 with the second sequence variant that was inputted into the protein language model. While step 519 describes manipulating the first sequence variant inputted into the model, step 521 comprises applying step 519 to the second sequence variant inputted into the model such that each part (i)-(v) of step 519 is performed on the second sequence variant. Thus, a conditional probability' of each amino acid position of the second sequence variant is determined, and an approximated joint probability of the second sequence variant is calculated via the step 521.

[0210] As an example, according to some embodiments, the step 521 can comprise utilizing the protein language model for: (i) masking a single amino acid position of the second sequence variant; (ii) predicting a probability distribution over all possible amino acids at the single amino acid position of the second sequence variant based on surrounding sequence context and the evolutionary information; (iii) determining a conditional probability of the single amino acid position of the second sequence variant based on the probability distribution; (iv) performing steps (i)-(iii) for each amino acid position of a length of the second sequence variant in a one-by-one manner; and (v) calculating a product of each conditional probability of each amino acid position of the length of the second sequence variant to produce an approximated joint probability of the second sequence variant.

[0211] According to some embodiments, the approximated joint probability of the second sequence variant can be calculated in the same and / or similar manner as the approximated joint probability of the first sequence variant is calculated in the step 519. For example,Agent Ref. No. P14786WOOO according to some embodiments, the second protein sequence variant can be represented as S’ wherein the one masked amino acid is represented as S’i. The surrounding sequence context can be represented as S’-i and the evolutionary information integrated into the model parameters can be represented as E. The conditional probability of each masked amino acid position of the second sequence variant, as determined in part (iii) of the step 521, can be determined from the probability distribution represented as P(S’i|S'-i, E). The approximated joint probability of the second sequence variant, as determined in part (v) of the step 521, can be represented as P(S’|E)= [Ji P(S't |SZ- 1, E). This approximated joint probability can be interpreted as a pseudolikelihood of drawing the sequence from the dataset upon which the protein language model was trained.

[0212] The step 523 comprises utilizing the protein language model to calculate a quotient of the approximated joint probability of the first sequence variant divided by the approximated joint probability of the second sequence variant. This quotient represents a ratio that can be used to evaluate the functional impact of each protein variant of the hybrid delivery assay with respect to a wild-type protein. This ratio can be referred to herein as the pseudolikelihood ratio. The pseudolikelihood ratio can be represented as PLR wherein PLR=P(S|E) / P(S jE) and wherein S refers to the first sequence variant and S’ refers to the second sequence variant. This pseudolikelihood ratio can be used to bin possible protein variants by phenotype class, which makes it possible to determine how likely it is for an inplanta edit to exhibit the desired phenotype regardless of which edited allele is produced. If the first sequence variant is more functional, then it is more likely to be favored by evolution and thereby represented in a large set of natural sequences, so that the pseudolikelihood ratio is greater than 1. Similarly, if the second sequence variant is more functional, then the pseudolikelihood ratio will be less than 1. Thus, this pseudolikelihood ratio can be used to predict the functional impact of protein variants produced by one or more different INDELS and / or the INDELS and determine the likelihood of an in-planta edit to exhibit the desired phenotype.

[0213] While the step 523 is shown to follow the step 521 in Figure 5C, the step 523 could follow the step 526. In other words, the step 523, which involves calculating a quotient of the approximated joint probabilities of the first and second sequence variants, could be applied whether steps 519 and 521 are applied (wherein a one-by-one approach is used to calculate the approximated joint probability of the first and second sequence variants) or steps 525 andAgent Ref. No. P14786WOOO526 are applied (wherein a group-by-group approach is used to calculate the joint probability of the first and second sequence variants).

[0214] The step 525 is an optional step. Step 525 comprises utilizing the protein language model to perform operations similar to parts (i)-(iv) of the step 519 by partitioning the first sequence variant into one or more groups of amino acid positions and proceeding in a group- by-group manner rather than the one-by-one manner. Operating in a one-by-one manner includes performing parts (i)-(iii) of the step 519 for each amino acid position along the length of each protein variant. This includes masking each single amino acid position, predicting a probability distribution over all possible amino acids at each single amino acid position based on surrounding sequence context and the evolutionary information, and determining a conditional probability of each single amino acid position based on the probability distribution. As noted, in the one-by-one approach, these parts of step 519 must occur at every amino acid position along the length of the protein variant. Thus, proceeding in the one-by-one manner can be computationally expensive in terms of computing power, necessary’ steps, and / or time for long sequence variants.

[0215] According to some embodiments, the step 525 can comprise utilizing the protein language model for: (i) partitioning the first sequence variant into one or more groups of amino acid positions; (ii) masking a group of amino acid positions of the one or more groups of amino acid positions of the first sequence variant; (iii) predicting a probability distribution associated with the group of amino acid positions based on surrounding sequence context and the evolutionary information; (iv) simultaneously determining a conditional probability of all amino acid positions of the group of amino acid positions based on the probability distribution; (v) performing steps (ii)-(iv) for each group of the one or more groups of amino acid positions of a length of the first sequence variant in a group-by-group manner; and (vi) calculating a product of each conditional probability’ of each group of the one or more groups of amino acid positions of the length to produce an approximated joint probability of the first sequence variant.

[0216] The step 526 is an optional step. The step 526 can comprise utilizing the protein language model to perform operations similar to parts (i)-(iv) of the step 521 for the second sequence variant by partitioning the second sequence variant into one or more groups of amino acid positions and proceeding in a group-by-group manner rather than the one-by-one manner. For example, according to some embodiments, the step 526 can comprise utilizingAgent Ref. No. P14786WOOO the protein language model for: (i) partitioning the second sequence variant into one or more groups of amino acid positions; (ii) masking a group of amino acid positions of the one or more groups of amino acid positions of the second sequence vanant; (lii) predicting a probability distribution associated with the group of amino acid positions of the second sequence variant based on surrounding sequence context and the evolutionary information; (iv) simultaneously determining a conditional probability of all amino acid positions of the group of amino acid positions of the second sequence variant based on the probability distribution associated with the second sequence variant; (v) performing steps (ii)-(iv) for each group of the one or more groups of amino acid positions of a length of the second sequence variant in a group-by-group manner; and (vi) calculating a product of each conditional probability of each group of the one or more groups of amino acid positions of the length of the second sequence variant to produce an approximated joint probability of the second sequence variant.

[0217] The steps 525 and 526 aim to increase efficiency by providing better scaling. Instead of computing conditional probabilities of each amino acid position by masking them one-by- one, each sequence variant can be partitioned into groups of size G amino acids, so that an entire group is masked and the conditional probability of all G amino acids in the group are predicted simultaneously. With this group-by-group approach (wherein L refers to the length of the first sequence variant, L’ refers to the length of the second sequence variant, and “ceil” refers to the least integer function which is also known as the ceiling function) determining conditional probabilities then takes only ceil(L / G)+ceil(L7G) forward passes of the protein language model, which amounts to approximately a G-fold speedup. In some instances some accuracy is lost as G increases.

[0218] If each sequence variant is partitioned into contiguous groups, valuable local contextual information from neighboring amino acids can be lost when running the protein language model. Therefore, according to some embodiments, each sequence variant is instead partitioned so that each group contains amino acids from random positions in the sequence variant.

[0219] The step 527 is an optional step. The step 527 can be performed if the lengths of the first and second protein variants inputted into the protein language model differ. The step 527 comprises utilizing the protein language model for calculating a quotient of a geometric mean of the approximated joint probability of the first sequence variant divided by a geometricAgent Ref. No. P14786WOOO mean of an approximated joint probability of the second sequence variant. This quotient can then represent the pseudolikelihood ratio that can be used to evaluate functional impact of protein variants and / or one or more INDELS, such as the INDELS. and determine the likelihood of an in-planta edit to exhibit the desired phenoty pe as described above with reference to the pseudolikelihood ratio of step 523. The geometric mean of the approximated joint probability of the first and second sequence variants can be used for variants of differing lengths such that the difference in lengths does not affect the pseudolikelihood ratio. For example, because the approximated joint probability of each variant is calculated my multiplying the conditional probability' of each amino acid position and / or each amino acid group and because each conditional probability is less than one. the variant with the larger length will be biased towards having a smaller approximated joint probability. Using the geometric mean of each variant’s approximated joint probability serves to reduce and / or eliminate this bias created by variants having differing lengths. The pseudolikelihood of the first sequence variant S having length L can be calculated using the equation P(S|E) = nf=1P( i |S-i , E). The pseudolikelihood of the second sequence variant S' having lengthL’ can be calculated using the equation P(S jE) = While the step 527 isshown to follow the step 526 in Figure 5D, the step 527 could follow the step 521. In other words, the step 527, which involves utilizing a geometric mean of the approximated j oint probabilities of the first and second sequence variants, could be applied whether steps 519 and 521 are applied (wherein a one-by-one approach is used to calculate the approximated joint probability of the first and second sequence variants) or steps 525 and 526 are applied (wherein a group-by -group approach is used to calculate the joint probability of the first and second sequence variants).

[0220] The steps 529, 531, and 533 comprise an additional and / or alternative way to arrive at the pseudolikelihood ratio that can be used to evaluate functional impact of protein variants and / or one or more INDELS, such as the INDELS, and determine the likelihood of an inplanta edit to exhibit the desired phenotype as described above with reference to the pseudolikelihood ratio of step 523. Steps 529, 531, and 533 can be performed in addition to and / or as an alternative to any of steps 519, 521, 523, 525, 526, and / or 527. Steps 529, 531, and 533 can be performed if the difference betw een the first and second sequence variants isAgent Ref. No. P14786WOOO a substitution rather than an insertion or deletion. Steps 529, 531, and 533 do not require any masking such that the steps 529, 531, and 533 can determine conditional probabilities and ultimately a pseudolikelihood ratio while only performing a single forward pass of each sequence variant. Thus, performing steps 529, 531, and 533 is less computationally expensive than performing steps 519, 521, and 523.

[0221] The step 529 comprises utilizing the protein language model for: (i) predicting a probability distribution over all possible amino acids at each amino acid position of the first sequence variant based on surrounding sequence context and the evolutionary information; (ii) determining a conditional probability of each amino acid position of the first sequence variant based on the probability distribution of each amino acid position; and (iii) calculating a product of each conditional probability of each amino acid position of the first sequence variant to produce an approximated joint probability of the first sequence variant.

[0222] The step 531 comprises utilizing the protein language model to perform step 529 with the second sequence variant inputted into the model rather than the first sequence variant. As an example, according to some embodiments, the step 531 can comprise utilizing the protein language model for: (i) predicting a probability distribution over all possible amino acids at each amino acid position of the second sequence variant based on surrounding sequence context and the evolutionary information; (ii) determining a conditional probability’ of each amino acid position of the second sequence variant based on the probability distribution of each amino acid position of the second sequence variant; and (iii) calculating a product of each conditional probability of each amino acid position of the second sequence variant to produce an approximated joint probability of the second sequence variant. According to some embodiments, the steps 529, 531, and 533 can be used if the second sequence variant and the first sequence variant have the same length.

[0223] The step 533 comprises utilizing the protein language model to calculate a quotient of the approximated joint probability of the first sequence variant divided by an approximated joint probability of the second sequence variant. This quotient can serve as the pseudolikelihood ratio that can be used to evaluate functional impact of protein variants and / or one or more INDELS, such as the INDELS, and determine the likelihood of an inplanta edit to exhibit the desired phenotype as described above with reference to the pseudolikelihood ratio of step 523. Again, according to some embodiments, the steps 529,Agent Ref. No. P14786WOOO531, and 533 can be used when the lengths of the first and second sequence variants are the same.

[0224] According to some embodiments, the steps 519. 521, 523. 525, 526, 527. 529, 531. and / or 533 can be performed using the ESM-lv protein language model. The ESM-lv protein language model is an open source model. See github.com / facebookresearch / esm?tab=MIT-l- ov-file#readme (date accessed: September 9, 2024) and www.biorxiv.org / content / 10. 1101 / 2021.07.09.450648v2 (date accessed: September 9, 2024), both of which are hereby incorporated by reference in their entireties. While steps 519, 521, 523, 525, 526, 527, 529, 531, and / or 533 can be performed using the ESM-lv protein language model, any suitable protein language model could be used. The ESM-lv model is described in further detail below.

[0225] The step 535 comprises utilizing the protein language model and / or any other suitable model and / or algorithm to compute an alignment-based dimensionless score of relative fitness regarding the first and second sequence variants. The step 535 can be performed additionally and / or alternatively to any of steps 519, 521, 523, 525, 526, 527. 529, 531, and / or 533. The protein language model and / or algorithm used to perform step 535 can be a different model and / or algorithm than that used to perform any of steps 519, 521, 523, 525, 526, 527, 529, 531, and / or 533. For example, the model and / or algorithm utilized to perform step 535 can be and / or comprise the PROVEAN model and / or algorithm. The PROVEAN model and / or algorithm is an open source model and / or algorithm. See www.jcvi.org / research / provean (date accessed: September 9, 2024) and www.jcvi.org / publications / predicting-functional-effect-amino-acid-substitutions-and- INDELS (date accessed: September 9, 2024), both of which are hereby incorporated by reference in their entireties. The PROVEAN algorithm / model can utilize and / or incorporate the BLAST alignment algorithm. See blast.ncbi.nlm.nih.gov / Blast.cgi (date accessed: September 9, 2024), which is hereby incorporated by reference in its entirety. However, according to various embodiments, any suitable model and / or algorithm can be used to perform step 535. The PROVEAN model and / or algorithm is described in further detail below.

[0226] The alignment-based dimensionless score of relative fitness produced by the model of step 535 is similar to the pseudolikelihood ratio produced by the protein language model of steps 519, 521, 523, 525. 526, 527, 529, 531, and 533 in that the relative fitness score can beAgent Ref. No. P14786WOOO used to estimate and / or predict the functional impact and / or effect(s) of one or more INDELS. such as the INDELS, and / or determine the likelihood of an in-planta edit to exhibit the desired phenotype.

[0227] The model of step 535 can be a machine learning model and can be trained, improved, and / or driven in the same and / or similar manner as the protein language model described above to be used for steps 519, 521, 523, 525, 526, 527. 529, 531, and / or 533. According to some embodiments, the model utilized for step 535 can utilize and / or comprise any machine learning techniques and / or training techniques described above regarding the method used for steps 519, 521, 523, 525, 526, 527, 529, 531, and / or 533.

[0228] The step 536, as shown in Figure 5F, comprises selecting a preferred gRNA based on the target gene editing frequency score and / or the predicted functional effect of the INDELS. As noted above, according to some embodiments, the predicted functional effect(s) of the INDELS and / or the preferred guide RNA can be used to select and / or predict in-planta information. According to some embodiments, such in-planta information can comprise phenoty pe information.

[0229] The step 536 of selecting a preferred gRNA can comprise the step 537. The step 537 comprises: (i) ranking a plurality of gRNAs based on the predicted target gene editing frequency score to obtain a subset of top-ranked gRNAs with the highest target gene editing frequency scores, (ii) identifying a desired allele among gene editing alleles produced in the plant cell by the top-ranked gRNA, and (iii) selecting a gRNA which is predicted to provide the desired allele. According to some embodiments, the number of top-ranked gRNAs is at least five. According to some embodiments, the number of top-ranked gRNAs is at least ten. According to various embodiments, the number of top-ranked gRNAs can range from 1 to N where N is any number greater than 1. Additionally or alternatively, the method 500 can comprise a step that ranks a plurality of gRNAs based on the predicted functional effect of the INDELS wherein a particular gRNA predicted to provide the desired functional effect is selected. Additionally or alternatively, according to some embodiments, the step 536 can comprise selecting a preferred gRNA based on a combination of the top-ranked gRNAs in terms of target gene editing score and the top-ranked gRNAs in terms of desired functional effect of the INDELS.

[0230] Additionally or alternatively, alleles can be characterized that will likely be generated in-planta based on the top alleles ranked in terms of editing efficiency. Additionally orAgent Ref. No. P14786WOOO alternatively, the hybrid delivery assay could be used to select guide RNAs based on a cassette of favorable alleles such that guide RNAs are selected that have more in-frame alleles or with more frameshift alleles depending on the goal of the gene edits. In certain embodiments, gRNAs which produce more frameshift alleles are selected when the goal is to produce amorphic alleles of a target gene. In certain embodiments, gRNAs which produce more in-frame alleles are selected when the goal is to produce hypomorphic alleles of a target gene. The hybrid delivery assay could also be used to select guide RNAs most likely to generate a particular allele of interest. For example, guide RNAs could be selected where one of the top alleles matched a desired allele based on prior research or guide RNAs could be selected because the top allele is predicted to have a desirable effect based on a model such as a protein language model.

[0231] Additionally or alternatively, steps 515, 517, 519, 521, 523, 525, 526, 527, 529, 531, 533, 535, 536, and 537 can be replaced with a step that comprises: (i) inputting a set of alleles produced by the subset of gRNAs into a protein language model; (ii) processing the set of alleles via the model; (iii) outputting a ratio from the model wherein the ratio can be used to evaluate the predicted functional effect of each allele of the set of alleles, and (iv) selecting one or more gRNAs from the subset of gRNAs based on the predicted functional effect of the alleles produced by the gRNA. Such a protein language model could be and / or comprise the ESM-lv model and / or the PROVEAN model and / or algorithm according to various embodiments.

[0232] The method 500 can further comprise the step 538. The step 538 comprises introducing into an experimental plant, experimental plant part, or experimental plant cell lacking the desired INDELS: (a) (i) the selected preferred gRNA or a polynucleotide encoding the selected preferred gRNA; and (ii) a Cas nuclease which binds the selected preferred gRNA or a polynucleotide encoding the Cas nuclease; or (b) a ribonucleoprotein complex comprising the Cas nuclease bound to the selected preferred gRNA.

[0233] According to some embodiments, the plant, plant part, experimental plant, and / or experimental plant part is and / or comprises a seed, leaf, stem, flower, meristem, or root. According to some embodiments, the plant cell and / or experimental plant cell comprises plant callus tissue, embryogenic callus tissue, or a protoplast.

[0234] The method 500 can further comprise the step 540. The step 540 comprises obtaining the plant or plant part comprising the desired INDELS from the experimental plant,Agent Ref. No. P14786WOOO experimental plant part, or experimental plant cell lacking the desired INDELS. Methods for obtaining regenerable plant structures and regenerating plants from plant cells or regenerable plant structures can be adapted from published procedures (Roest and Gihssen. Acta Bot. Neerl., 1989, 38(1), 1-23; Bhaskaran and Smith, Crop Sci. 30(6): 1328-1337; Ikeuchi et al.. Development, 2016, 143: 1442-1451). Methods for obtaining regenerable plant structures and regenerating plants from plant cells or regenerable plant structures can also be adapted from US Patent Application Publication Nos. 20170121722, 20220251587, and PCT application number WO 2024 / 015781, which are incorporated herein by reference in their entireties and specifically with respect to such disclosure.

[0235] According to some embodiments, the method 500 can further comprise the step 542. The step 542 comprises isolating, amplifying, and / or sequencing genomic DNA sequences containing the target plant gene from an experimental plant cell that had received the distinct candidate guide RNA and the Cas nuclease encoding DNA vector, wherein the isolated, amplified, and sequenced genomic DNA sequences can be used in the step of predicting the target gene editing frequency score and / or in the step of predicting the functional effect of the INDELS. Such isolation, amplification, and / or sequencing can be performed using any DNA extraction, DNA amplification, and / or DNA sequencing device(s) known in the art.

[0236] According to some embodiments, at least one of the steps of the method 500 can be performed, at least partially, using the system 10 and / or any component(s) thereof.

[0237] According to some embodiments, each of the methods 200, 300, 400. and 500 can be include steps that are interchangeable such that any step of any of methods 200, 300, 400, and 500 can be included as part of any of the other methods and / or can replace a step in any other the other methods. Additionally, while the method steps of the methods 200, 300, 400, and 500 of Figures 2-5G are shown in a particular order, the steps of each of the methods could be performed in any suitable order. Additionally, each of the methods 200. 300, 400. and 500 could be performed with more or fewer steps according to various embodiments.

[0238] Figure 6 shows a chart comparing aspects of different Cas nuclease / gRNA gene editing molecule delivery' approaches including: (1) Cas nuclease / gRNA RNP complex delivery (RNP) in protoplasts, (2) hybrid delivery in protoplasts (e.g.. delivery of a plasmid encoding the Cas nuclease and an RNA molecule comprising the gRNA), and (3) plasmid transformation in planta (e.g., delivery' of a plasmid encoding both the Cas nuclease and the RNA molecule comprising the gRNA to plant cells or plant parts). As can be seen in the chartAgent Ref. No. P14786WOOO of Figure 6, the hybrid delivery approach, which is comprised in many of the embodiments herein describing the hybrid deliver}’ assay, has higher sensitivity than the Cas nuclease / gRNA RNP complex assay while retaining advantages in seal abi 1 ity and throughput. Cas nuclease / gRNA RNP complex (RNP) testing in protoplasts is limited due to RNP being precomplexed with the crRNA of the gRNA in excess. RNP testing in protoplasts is further limited due to the fact that RNP delivered to cells in excess gives optimal editing performance but is not reflective of RNP levels in planta. Hybrid deliver}’ provides several advantages over RNP testing. For example, hybrid delivery allows for in vivo complexing of the introduced RNA molecule comprising the gRNA with the Cas nuclease produced by expression of the introduced DNA encoding the Cas nuclease w hich adds sensitivity to guide efficacy. Also, hybrid delivery allows for Cas nuclease protein expression in the host system to more closely approximate Cas nuclease protein expression which occurs in plant transformation procedures aimed at production of gene edited plants.

[0239] Figure 7 shows a graphical representation comparing gene editing outcomes (combined INDELS) when the three different approaches to Cas nuclease and guide RNA delivery described in Figure 6 are used with eight different guide RNAs. As can be seen in the graph of Figure 7, a correlation exists between hybrid delivery assay results and editing efficiency in explants.

[0240] Figure 8A shows a plot of the proportion of plants with high quality gene editing alleles obtained with in planta plasmid transformation with eight different gRNAs versus the Cambridge mean combined INDELS. As can be seen in the graph of Figure 8, a correlation exists between hybrid deliver}’ assay results and editing efficiency in explants. As used herein, the term “high quality’' with reference to gene editing alleles refers to an allele that is above 25% of the total alleles that are likely to be heritable (copy number is between 1 and 2).

[0241] Figure SB shows a plot of the most frequent alleles obtained with the CR1808 guide RNA obtained with in planta plasmid transformation-based editing where the alleles were either False or True in editing results obtained with that same guide RNA in the hybrid delivery assay. According to some embodiments, the edited alleles shown along the horizontal axis of the plot of Figure 8B refer to multiple different INDELS wherein a positive value refers to an INDELS downstream of a cut site (i.e., 3' to the cleavage site on the target strand complementary to the spacer RNA sequence of the gRNA), a negative valueAgent Ref. No. P14786WOOO refers to an INDELS upstream of a cut site (i.e., 5' to the cleavage site on the target strand complementary’ to the spacer RNA sequence of the gRNA), the number before the colon refers to the location of the INDELS in terms of number of base pairs from the cut site, the number after the colon refers to the size of the INDELS in terms of number of base pairs, and the letter refers to the type of INDELS. For example, -24:39D refers to an INDELS 24 base pairs upstream of the nuclease cut site, wherein the INDELS comprises 39 base pairs and the INDELS is a deletion. This same notation is used in Figures 14A-C and 17.

[0242] Figure 9A shows a plot of the percentage of high quality edited plants obtained with in planta plasmid transformation-based editing versus the mean FPF INDELS % for different gRNAs hybrid delivery' assay. As used herein, the term “FPF"’ refers to false positive filters. False positive filters are a means to ensure false positives are screened out. For example, use of false positive filters excludes single nucleotide polymorphisms (SNPs) that fall within an expected region and may be counted as an INDELS.

[0243] Figure 9B shows a plot of the percentage of high quality7edited plants obtained with in planta plasmid transformation-based editing versus the percentage of high quality edited explants obtained with in planta plasmid transformation-based editing for different gRNAs.

[0244] Figure 10 shows a comparison of editing efficiency results obtained with hybrid delivery7(Hybrid Mean FPF INDELS), with in planta plasmid transformation-based editing (Plant pct HQE), and with explants obtained from plasmid transformation-based editing of cultured plant cells (Explant pct HQE) with three different gRNAs (CR1820. CR1821. and CR1933). The CR1933gRNA would be selected for in planta plasmid transformation-based editing based on its higher editing efficiency in the hybrid delivery assay. The hybrid mean FPF bar represents the average combined editing frequency across all alleles in protoplasts with alleles outside the expected editing window removed. The plant pct HQE bar represents the percentage of plants with at least one allele with greater than 25% editing efficiency. The explant pct HQE bar represents the percentage of explants with at least one allele with greater than 25% editing efficiency. The data shown in Figure 10 is representative of the three gRNAs targeting a specific gene, which is referred to in Figure 10 as “GenelD 738."’

[0245] Figure 11 compares the percentage of instances where the correct gRNA is selected in a hybrid delivery assay versus the percentage of instances where the correct gRNA is selected by assaying explants obtained from plasmid transformation-based editing of cultured plant cells.Agent Ref. No. P14786WOOO

[0246] Figure 12 shows a table comparing gene editing results obtained using unoptimized gRNAs with results obtained using a gRNA selected by use of hybrid delivery assays and other methods disclosed herein. As seen in Figure 12, optimizing vectors with the highest efficiency guides for each target can improve multiplex editing efficiency. Further, using early sampling pipeline for selecting the top performing guides is resource intensive. Even further, hybrid deliver}' guide testing enables increased multiplexing efficiency faster. Additionally, hybrid delivery can test the best combinations of guides for multiplexing in a single vector. Hybrid delivery can be used to test the best combination of guide RNAs when multiplexing.

[0247] Figure 13 shows a graphical representation of population level editing efficiency- obtained in whole plant gene editing experiments (e.g., in planta plasmid transformationbased editing) versus results obtained in a hybrid delivery assay with plant protoplasts. As can be seen in Figure 13, population level editing efficiency is correlated between protoplasts and plants.

[0248] Figures 14A-C show graphical representations comparing alleles obtained with hybrid delivery based editing in protoplasts and alleles obtained with in planta plasmid transformation-based editing with the CR1808, CR1946, and CR1820 gRNAs.

[0249] Figure 15 shows the percentage of overlap of alleles obtained with hybrid delivery and obtained with in planta plasmid transformation-based editing with different gRNAs. As can be seen in Figure 15, a majority of high quality alleles are found in the top 5 alleles in terms of hybrid delivery data.

[0250] Figure 16 shows the percentage of overlap of the top 5 alleles obtained with hybrid delivery- and the top 10 alleles obtained with in planta plasmid transformation-based editing with different gRNAs.

[0251] Figure 17 shows a depiction of a bioinformatics pipeline related to qualitative prediction(s) of gene editing outcomes.

[0252] Figure 18 shows a depiction of predicting an edited allele effect on function of the encoded protein using a particular protein language model. Figure 18 shows the prediction of an edited allele effect using the ESM-lv protein language model as described above and below. The ESM-lv protein language model can be used in association with any method described herein such as any of the methods 200, 300, 400, and / or 500. According to some embodiments, a PROVEAN model and / or algorithm can be used rather than the ESM-lvAgent Ref. No. P14786WOOO model as described above and below. The PROVEAN model and / or algorithm can be used in association with any method described herein such as any of the methods 200, 300, 400, and / or 500. According to some embodiments, the model can be pretrained on natural protein sequences. According to some embodiments, the likelihood of the protein is given by the likelihood of its residues. According to some embodiments, the model is configured to output and / or produce a score, wherein said score indicates the likelihood of an edited sequence compared to a wild-type sequence. A higher score (z. e. , a higher likelihood) indicates that the edited allele is more functional.

[0253] Figure 19 shows a graphical representation of prediction(s) of gene editing outcomes based on the protein language model of Figure 18 using data on the effect of 314 single AA deletions in the PTEN tumor suppressor gene in humanized yeast cells. As noted in Figure 19, large variations in INDELS size can confound model predictions, but multiple INDELS of similar size do not tend to have the same issues of confounding the model.

[0254] Figures 20A-D show various graphical representations of predictions regarding the effect(s) of edited alleles using the protein language model of Figure 18 using data on the effect of 314 single AA deletions in the PTEN tumor suppressor gene in humanized yeast cells. Figure 20A shows essentially the same data as that of Figure 19, however. Figure 20A further includes data regarding length distribution and / or sequence length. As mentioned herein, large variations in INDELS size can be a confounding factor for the model, particularly for insertions as shown in Figure 20C. Additionally or alternatively, large variations in INDELS size can be interpreted as outliers, as shown in Figure 20D. The issue of large variations in INDELS size for insertions being a confounding factor for the model is mitigated by the fact that editing experiments are more likely to result in deletions than insertions.

[0255] Figure 21 shows a depiction of a pipeline for predicting editing effect(s) of one or more INDELS. The pipeline can include: nominating guides using a guide nomination tool, performing hybrid delivery guide testing, distributing edited alleles, qualitatively characterizing the alleles based on bioinformatics, performing an insertion, deletion, and / or substitution (INDELS), using one or more models to predict effect(s) of an INDELS. and receiving a quantitative score as output from the one or more models regarding the predicted effect(s) of the INDELS. According to various embodiments, the one or more models can comprise the ESM-lv model and / or the PROVEAN model as described herein. Prior to usingAgent Ref. No. P14786WOOO the pipeline shown in Figure 21, a user can first implement the following steps: (i) onboarding a dataset of one or more in-frame INDELS for model validation using allele frequency as an estimate of variant severity; (ii) implementing bioinformatics-based variant effect prediction algorithms (such as PROVEAN and / or any other such algorithm / model) to benchmark the models; (iii) implementing an ensemble of additional variant effect models (such as Tranception, PRIME, and / or any other such model); and (iv) scaling to a pipeline for integration into a tool used to nominate guide RNAs.

[0256] Figure 22 shows the schematic of a gene and various locations for editing a gene to modify gene expression.

[0257] Figure 23 show s a diagrammatic view of the architecture of the ESM-lv model according to at least some aspects of the present disclosure. As noted above, the ESM-lv model can be utilized by any of the system(s) and / or method(s) described herein including, but not limited to, the system 10, the method 200, the method 300, the method 400, and / or the method 500.

[0258] The ESM-lv model can be configured to accept a pair of protein sequence variants and compute a dimensionless score of relative fitness between the two variants. This score indicates the relative likelihood of sampling a protein sequence at random from a large distribution of naturally occurring proteins (z.e., Uniref50, UniRef90, UniRefl 00, and / or any other suitable set of protein information) before and after a mutation, and therefore how much more likely it is to resemble a functional protein. The relative fitness score can be used to estimate the impact of the mutation wherein variant 2 is a mutation of variant 1. The mutation in question may be an INDELS or a more complicated variant.

[0259] The ESM-lv model is an open source transformer-based protein language model (pLM). The ESM-lv model can be pretrained by self-supervision on a large and diverse set of natural proteins (z.e.. UniRef50, UniRef90, UniRefl 00. and / or any other suitable set of protein information). During training, the model can learn a set of parameters that reflect the integration of evolutionary' information present in the diversify of the training set.

[0260] According to some embodiments, the ESM-lv model can include zero or more model parameters wherein the number of model parameters can range from zero to N where N is any number greater than zero. According to various embodiments, the ESM-lv model can include three model parameters which include mask size, mask ty pe, and model name.Agent Ref. No. P14786WOOO

[0261] The mask size parameter can be an integer value according to some embodiments. However, the mask size parameter could be any suitable type of value. The mask size parameter can indicate the number of ammo acids to mask simultaneously in each scoring iteration of the model. Larger values of the mask size parameter can lead to lower accuracy but result in faster computations and be less computationally expensive. The mask size parameter can be set to zero to use a wild-type scoring strategy as described below. Setting the mask size parameter to zero causes faster computations but is only appropriate for substitutions rather than insertions or deletions. A user can specify the mask size based on the user’s preference as to how many amino acids to mask simultaneously. In this way, the protein language model can be modified based on user input. Alternatively, the mask size parameter could be learned by the model during training.

[0262] The mask type parameter can be a string according to some embodiments. However, the mask type parameter could be any suitable type of value. The mask type parameter can be used to indicate whether the model will mask a block of consecutive amino acid positions or a block of random amino acid positions in each scoring iteration of the model. For example, according to some embodiments, the string ‘"block” can be input as the mask type parameter to indicate that the model will mask one or more blocks of consecutive amino acid positions in each scoring iteration of the model. As a further example, according to some embodiments, the string “rand” can be input as the mask type parameter to indicate that the model will mask one or more samples of random amino acid positions in each scoring iteration of the model. A user can specify the mask type based on the user’s preference as to how to implement the masking strategy. In this way, the protein language model can be modified based on user input. Alternatively, the mask type parameter could be learned by the model during training. Masking random amino acid positions rather than consecutive amino acid positions can mitigate positional bias and can allow the model to see more local context when making predictions.

[0263] The model name parameter can be a string according to some embodiments. However, the model name parameter can be any suitable type of value. The model name parameter can be used to indicate which one of the ESM-lv models to use when performing the scoring. For example, the model name parameter can be the name of one of the ESM-lv models. ESM-lv has five separate models trained with different random initializations. A user can specify the model name to indicate which ESM-lv model will perform the scoringAgent Ref. No. P14786WOOO and / or predicting. In this way, the protein language model can be modified based on user input. Alternatively, the model name parameter could be learned by the model during training.

[0264] According to some embodiments, the ESM-lv model can include two or more model inputs wherein the number of model inputs can range from two to N where N is any number greater than two. According to various embodiments, the ESM-lv model can include three model inputs which include protein sequence variant 1, protein sequence variant 2. and protein ID.

[0265] The protein sequence variant 1 model input can be a string according to some embodiments. However, the protein sequence variant 1 input can be any suitable type of value. The protein sequence variant 1 input is and / or comprises the first protein sequence variant to be input into the model. According to some embodiments, the protein sequence variant 1 input can be a reference or wild-type sequence. A user can enter any desired protein sequence variant as the protein sequence variant 1 input. In this way, the protein language model can be modified based on user input.

[0266] The protein sequence variant 2 model input can be a string according to some embodiments. However, the protein sequence variant 2 input can be any suitable type of value. The protein sequence variant 2 input is and / or comprises the second protein sequence variant to be input into the model. A user can enter any desired protein sequence variant, that is a mutation of the protein sequence variant 1, as the protein sequence variant 2 input. In this way, the protein language model can be modified based on user input.

[0267] The protein ID model input can be a string according to some embodiments. However, the protein ID input can be any suitable type of value. The protein ID input is and / or comprises identification of the two protein sequence variants and / or the scoring thereof. A user can enter any desired name and / or identification information as the protein ID input. In this way, the protein language model can be modified based on user input.

[0268] According to some embodiments, the ESM-lv model can include one or more model outputs wherein the number of model outputs can range from one to N where N is any number greater than one. According to various embodiments, the ESM-lv model can include three model outputs which include absolute fitness score for protein sequence variant 1, absolute fitness score for protein sequence variant 2, and relative fitness score between protein sequence variants 1 and 2.Agent Ref. No. P14786WOOO

[0269] The absolute fitness score for protein sequence variant 1 output can be a float according to some embodiments. However, the absolute fitness score for protein sequence variant 1 output can be any suitable type of value.

[0270] The absolute fitness score for protein sequence variant 2 output can be a float according to some embodiments. However, the absolute fitness score for protein sequence variant 2 output can be any suitable type of value.

[0271] The relative fitness score between protein variant sequences 1 and 2 output can be a float according to some embodiments. However, the relative fitness score between protein variant sequences 1 and 2 output can be any suitable type of value. The relative fitness score between protein variant sequences 1 and 2 output is the dimensionless score that indicates relative likelihood of sampling a protein sequence at random from a large distribution of naturally occurring proteins before and after mutation, and therefore how much more likely it is to resemble a functional protein.

[0272] According to some embodiments, the ESM-lv model is configured to accept a protein sequence “S” as input. The model is then configured to mask the identity of one amino acid position ‘"Si”. The model is then configured to predict a probability distribution over all possible amino acids at the masked position conditioned on the surrounding sequence context “S-i” and the evolutionary information “E” integrated into the model parameters. The model can then determine the conditional probability of the original amino acid from the distribution P(Si|S-i. E).

[0273] The model is configured to include multiple approaches of calculating the relative fitness score between the two protein sequence variants that are input into the model. One approach can be referred to as the pseudolikelihood scoring strategy'.

[0274] When using the pseudolikelihood scoring strategy, for a protein sequence variant “S” having a length “L” that is input into the ESM-lv model, the model can mask the identity of each amino acid at each amino acid position one-by-one to produce L sequences. The model can then be used to predict the conditional probability of each masked amino acid of the protein sequence variant S. The model can then multiply each conditional probability' wherein the product thereof results in an approximated joint probability of the full sequence S conditioned on the evolutionary information E integrated into the model parameters. This approximated joint probability' can be represented as: P(S|E) = Hi P(Si|S-i, E). ThisAgent Ref. No. P14786WOOO approximated joint probability can be interpreted as a pseudolikelihood of drawing the first sequence variant from the dataset on which the model was trained.

[0275] The model can then repeat this procedure on the second protein sequence variant "S ’ having length “L”’ that is input into the model. For example, the model can mask the identity of each amino acid at each amino acid position one-by-one to produce L’ sequences. The model can then be used to predict the conditional probability of each masked amino acid of the second protein sequence variant S’. The model can then multiply each conditional probability wherein the product thereof results in an approximated joint probability’ of the full sequence S’ conditioned on the evolutionary information E integrated into the model parameters. This approximated joint probability can be represented as the following: P(S'|E) = fli P(S’i|S’-i, E). This approximated joint probability’ can be interpreted as a pseudolikelihood of drawing the second sequence variant from the dataset on which the model was trained.

[0276] Once the ESM-lv model has computed the approximated joint probability’ for each of the first protein sequence variant S and the second protein sequence variant S’, the model can compute the pseudolikelihood ratio 'PLR” wherein PLR = P(S|E) / P(S’|E). If the first protein sequence variant S is more functional, then S is more likely to be favored by evolution and thereby represented in a large set of natural sequences. If the first protein sequence variant S is more functional, then PLR is greater than 1. If the second protein sequence variant S’ is more functional, then S’ is more likely to be favored by evolution and thereby represented in a large set of natural sequences. If the second protein sequence variant S’ is more functional, then PLR is less than 1.

[0277] According to some embodiments, the pseudolikelihood scoring strategy involves 2*L forward passes of the method when performing the steps to compute the relative fitness score. According to some embodiments, the pseudolikelihood scoring strategy can be represented by the following equation:52 [k’g AS1' | Sty, £’) -■ log

[0278] Another approach in which the ESM-lv model is configured to calculate the relative fitness score between two protein sequence variants can be referred to as the wildtypeAgent Ref. No. P14786WOOO marginal probability approach. The wildty pe marginal probability' approach cannot be used if the difference and / or mutation between the two protein sequence variants input into the model includes insertion(s) and / or deletion(s). This is because positions of mutated residues of insertion(s) and / or deletion(s) are different from wildtype residues. The wildtype marginal probability7approach can be used when the mutation and / or difference between the two protein sequence variants input into the model comprises one or more substitution(s) rather than insertion(s) and / or deletion(s). This is because the positions of all the residues remain the same. The wildtype marginal probability’ approach is similar to the pseudolikelihood scoring strategy except that all scores are conditioned only on the full wildtype sequence without any masking.

[0279] Similar to the pseudolikelihood scoring strategy, the wildtype marginal probability approach can include determining the conditional probabilities of both the original and substituted amino acid from the distribution P(Si|S-i, E) for each amino acid position. Again, similar to the pseudolikelihood scoring strategy, the wildtype marginal probability approach can then include summing the conditional probabilities over all amino acid positions to determine an approximated joint probability of the first protein variant sequence S. as represented by the following: P(S|E) = fli P(Si| S-i, E). The same procedure can be conducted for the second protein variant sequence S’, wherein conditional probabilities of each amino acid position can be determined from the distribution P(S ’i|S’-i, E). The model can then sum each conditional probability over all amino acid positions to determine an approximated joint probability of the second protein sequence variant S’, as represented by the following: P(S’|E) = Hi P(S’i|S’-i, E).

[0280] The wildtype marginal probability approach can then include calculating a pseudolikelihood ratio PLR based on the two approximated joint probabilities based on the following: PLR = P(S|E) / P(S jE). Because the non-mutated residues and sequence that the model is conditioned on are the same in both numerator and the denominator of the pseudolikelihood ratio equation, the numerator and denominator cancel each other out. As a result, non-substitution positions in the pseudolikelihood sum need not be considered. Similar to the pseudolikelihood scoring strategy, for the wildtype marginal probability approach, if S is more functional, then S is more likely to be favored by evolution and thereby represented in a large set of natural sequences. Thus, if S is more functional, then PLR is greater than 1. Additionally, for wildty pe marginal probability, if S’ is more functional, then S’ is moreAgent Ref. No. P14786WOOO likely to be favored by evolution and thereby represented in a large set of natural sequences. Thus, if S’ is more functional, then PLR is less than 1.

[0281] The wildtype marginal probability approach can include setting the mask size model parameter to zero since no masking is involved. According to some embodiments, a user could set the mask size model parameter to zero and / or the model could automatically set the mask size model parameter to zero when using the wildtype marginal probability approach. According to some embodiments, the wildtype marginal probability approach includes only a single forw ard pass of the model rather than 2*L forward passes as required by the pseudolikelihood scoring strategy. Thus, the wildtype marginal probability approach is faster and less computationally expensive than the pseudolikelihood scoring strategy. According to some embodiments, the wildtype marginal probability approach can be represented by the following equation:

[0282] The ESM-lv model further includes a sequence length normalization feature. The sequence length normalization feature is not necessary for substitution variants because the sequence length between the first and second protein sequence variants would be the same. Rather, the sequence length normalization feature is appropriate to use w hen the first protein sequence variant S and the second protein sequence variant S’ differ by an in-frame INDELS such that they have different lengths L and L'. Sequence length normalization can be applied to protein sequence variants analyzed by the model. When computing the approximated joint probability of each of the first and second protein sequence variants S and S’ using the pseudolikelihood scoring strategy, a product of the conditional probabilities of each amino acid position of each protein sequence variant is computed. Thus, the longer protein sequence variant will have more terms in its product. Since each conditional probability is by definition less than or equal to one, the longer sequence is biased towards having a smaller approximated joint probability. The sequence length normalization feature can be used to compare the approximated joint probabilities of the first and second protein sequence variants S and S’ equitably. The sequence length nonnalization feature can apply a transformation that accounts for the multiplicative nature of independent probabilities and normalizes for theAgent Ref. No. P14786WOOO differing lengths and / or residues in each protein sequence variant. This transformation is performed by computing the geometric mean of the conditional probabilities rather than the product. Thus,

[0283] The two geometric mean equations above can be interpreted as each residue and / or amino acid position of S contributing 1 / L of the P(S|E) and each residue and / or amino acid position of S’ contributing 1 / L’ ofP(S’|E). According to some embodiments, the sequence length normalization feature applied to the pseudolikelihood scoring strategy can be represented by the following equation:

[0284] Figure 24 shows a diagrammatic depiction of a masking strategy capable of being used in conjunction with the ESM-lv model. Another feature of the ESM-lv model is that the model is capable of using different types of masking strategies. When performing the pseudolikelihood scoring strategy as described above, the model takes L + L’ forward passes of the protein sequence variants to compute a relative fitness score. This can be computationally expensive for long protein variant sequences. The model includes the abi li ty to perform masking in a group-by -group manner rather than a one-by-one manner, which provides for better efficiency and scaling. Instead of computing conditional probabilities of each amino acid position by masking them one-by-one, as described above regarding the pseudolikelihood scoring strategy, the amino acids can be partitioned into chunks and / or groups of size G. The size G of the chunk and / or group is determined based on the mask size model parameter (i.e., G = mask size model parameter). The model can then mask each chunk and / or group wherein the conditional probability of all G amino acids in the chunk and / or group is determined simultaneously.

[0285] It should be noted that if the protein sequence variant is partitioned into chunks and / or groups that include contiguous and / or consecutive amino acid positions, such as when the mask t pe model parameter is ’’block”, valuable local contextual information from neighboring amino acids will be lost when running the model. Thus, the model can partition the protein sequence variant so that each chunk and / or group comprises amino acid positionsAgent Ref. No. P14786WOOO from random positions in the protein sequence variant. The amino acid positions can be selected randomly without replacement. This random selection of amino acid positions in each chunk and / or group can occur when the mask type model parameter is “rand”. Masking random amino acid positions rather than consecutive amino acid positions mitigates potential bias and allows the model to see more local context when making predictions. Using group- by-group masking rather than one-by-one masking takes only ceil(LZG) + ceil(L7G) forward passes of the model, which results in approximately a G-fold speedup. It should be noted that “ceil” refers to the least integer function which is also known as the ceiling function. Thus, the group-by -group masking is more efficient and less computationally expensive than one- by-one masking. It should be noted that group-by-group masking can result in some accuracy being lost as G increases. In other words, larger values of mask size G leads to lower accuracy but results in faster and more efficient computation.

[0286] Figure 25 shows a workflow diagram incorporating use of the ESM-lv protein language model. The workflow^ diagram represents an example workflow that a user and / or a system could employ when using the ESM-lv model. Such a system could be any system described herein such as the system 10. While the workflow of Figure 25 represents one example of a workflow, any suitable workflow could be employed when using the ESM-lv model. Figure 25 shows that a user and / or a system can load dependencies; load the model; load input data; auto-batch (which can include a GPU profile); run the model in batches; compute, fill, and save relative fitness scores; and display and plot the output of the model. According to some embodiments, the workflow diagram incorporating use of the ESM-lv protein language model of Figure 25 could be used for individual protein sequence(s) rather than for batches of protein sequence(s). For example, according to some embodiments, autobatching is not necessary, and the model can be run on individual protein sequence(s) rather than batch(es).

[0287] As noted above, the PROVEAN model and / or algorithm can be utilized by any of the system(s) and / or method(s) described herein including, but not limited to, the system 10, the method 200, the method 300, the method 400, and / or the method 500.

[0288] As noted above, PROVEAN (Protein Variation Effect Analyzer) (see. e.g., w w.jcvi.org / research / provean) (date accessed: September 9, 2024) is an open source (see. e.g., www.j cvi . org / si tes / default / files / as sets / proj ects / pro vean / downl oads / LIC ENSE (date accessed: September 9, 2024), which is hereby incorporated by reference in its entirety)Agent Ref. No. P14786WOOO model, algorithm, and / or software tool for predicting biological impact of an amino acid INDELS in a protein sequence.

[0289] The PROVEAN model and / or algorithm can be configured to accept a pair of protein sequence variants and compute a dimensionless, alignment-based score of relative fitness between the two variants. This alignment-based score measures the relative change in sequence similarity of a query sequence to a protein sequence homolog before and after a mutation of the query sequence. The relative fitness score can be used to estimate the impact of the mutation on biological function wherein variant 2 is a mutation of variant 1. The mutation in question may be an INDELS or a more complicated variant.

[0290] According to some embodiments, the PROVEAN model and / or algorithm can include zero or more model parameters wherein the number of model parameters can range from zero to N where N is any number greater than zero. According to various embodiments, the PROVEAN model and / or algorithm can include three model parameters which include BLAST database, BLAST database director}7, and save interval.

[0291] The BLAST (see. e.g., blast.ncbi.nlm.nih.gov / Blast.cgi (date accessed: September 9, 2024)) database model parameter can be a string value according to some embodiments. However, the BLAST database model parameter could be any suitable type of value. The BLAST database model parameter can indicate the name of the protein BLAST database used to create multiple sequence alignments for each of the input protein sequence variants. The BLAST database could be a custom database and / or could be a standard protein nr database from the National Center for Biotechnology Information (NCBI) (see, e.g., www.nlm.nih.gov / ncbi / workshops / 2023-08_BLAST_evol / databases.html (date accessed: September 9, 2024), which is hereby incorporated by reference in its entirety). For example, according to some embodiments, the BLAST database can be a custom plant-specific protein database built using data from UniProtKB (see, e.g., ftp. uniprot. org / pub / databases / uniprot / current_release / knowledgebase / taxonomic_di visions / (date accessed: September 9, 2024), which is produced by UniProt® and which is hereby incorporated by reference in its entirety) and clustered to 100% minimum sequence similarity to remove duplicate and subfragment sequences. A user can specify the BLAST database based on the user’s preference. In this way, the model and / or algorithm can be modified based on user input. Alternatively, the BLAST database parameter could be learned by the model during training.Agent Ref. No. P14786WOOO

[0292] The BLAST database directory model parameter can be a string according to some embodiments. However, the BLAST database directory model parameter could be any suitable type of value. The BLAST database directory model parameter can refer to a directory and / or location where BLAST database(s) are stored. Thus, specifying the BLAST database directory model parameter can allow the system to correctly locate and access the desired BLAST database. A user can specify' the BLAST database directory' model parameter based on the user’s preference as to how to implement the masking strategy’. In this way. the model and / or algorithm can be modified based on user input. Alternatively, the BLAST database directory parameter could be learned by the model during training.

[0293] The save interval model parameter can be an integer value according to some embodiments. However, the save interval model parameter could be any suitable type of value. The save interval model parameter can be used to indicate the number of sequence pair batches to process before saving. According to some embodiments, a smaller save interval model parameter will run more slowly but will minimize lost work if workflow should fail. A user can specify the save interval model parameter based on the user's preference. In this way, the model and / or algorithm can be modified based on user input. Alternatively, the save interval parameter could be learned by the model during training.

[0294] According to some embodiments, the PROVEAN model and / or algorithm can include two or more model inputs wherein the number of model inputs can range from two to N where N is any number greater than two. According to various embodiments, the PROVEAN model and / or algorithm can include four model inputs which include protein ID for protein sequence variant 1, protein ID for protein sequence variant 2, protein sequence variant 1, and protein sequence variant 2.

[0295] The protein ID for protein sequence variant 1 model input can be a string according to some embodiments. However, the protein ID for protein sequence variant 1 input can be any suitable type of value. The protein ID for protein sequence variant 1 input is and / or comprises identification of protein sequence variant 1. A user can enter any desired name and / or identification information as the protein ID for protein sequence variant 1 input. In this way, the model and / or algorithm can be modified based on user input.

[0296] The protein ID for protein sequence variant 2 model input can be a string according to some embodiments. Hoyvever, the protein ID for protein sequence variant 2 input can be any suitable type of value. The protein ID for protein sequence variant 2 input is and / or comprisesAgent Ref. No. P14786WOOO identification of protein sequence variant 2. A user can enter any desired name and / or identification information as the protein ID for protein sequence variant 2 input. In this way, the model and / or algorithm can be modified based on user input.

[0297] The protein sequence for variant 1 model input can be a string according to some embodiments. However, the protein sequence for variant 1 input can be any suitable type of value. The protein sequence for variant 1 input is and / or comprises the first protein sequence variant to be input into the model. According to some embodiments, the protein sequence for variant 1 input can be a reference or wild-type sequence. A user can enter any desired protein sequence variant as the protein sequence for variant 1 input. In this way, the model and / or algorithm can be modified based on user input.

[0298] The protein sequence for variant 2 model input can be a string according to some embodiments. However, the protein sequence for variant 2 input can be any suitable type of value. The protein sequence for variant 2 input is and / or comprises the second protein sequence variant to be input into the model. According to some embodiments, the protein sequence for variant 2 input can be a reference or wild-type sequence. A user can enter any desired protein sequence variant, that is a mutation of the protein sequence for variant 1. as the protein sequence for variant 2 input. In this way, the model and / or algorithm can be modified based on user input.

[0299] According to some embodiments, the PROVEAN model and / or algorithm can include one or more model outputs wherein the number of model outputs can range from one to N where N is any number greater than one. According to various embodiments, the PROVEAN model and / or algorithm can include one model output which includes relative fitness score between protein sequence variants 1 and 2.

[0300] The relative fitness score between protein variant sequences 1 and 2 output can be a float according to some embodiments. However, the relative fitness score between protein variant sequences 1 and 2 output can be any suitable type of value. The relative fitness score between protein variant sequences 1 and 2 output is the dimensionless, alignment-based score that measures the relative change in sequence similarity of a query sequence to a protein sequence homolog before and after a mutation of the query sequence.

[0301] According to some embodiments, the PROVEAN model and / or algorithm can proceed by performing a series of operations. For example, according to some embodiments, the PROVEAN model and / or algorithm can be implemented and / or performed as described inAgent Ref. No. P14786WOOO journals. plos.org / plosone / article?id=10.1371 / joumal.pone.0046688 (date accessed: September 9. 2024), which is hereby incorporated by reference in its entirety. While the PROVEAN model could be implemented using many different approaches, the description of operations below is one example of an implementation of the PROVEAN model according to some embodiments. The first operation can comprise using the BLAST alignment algorithm (see, e.g., blast.ncbi.nlm.nih.gov / Blast.cgi (date accessed: September 9, 2024)) along with a specified BLAST database (such as blastdb), to gather homologous protein sequences that are similar to the query' sequence, wherein the query sequence refers to the first protein sequence variant (protein sequence for variant 1) input into the PROVEAN model and / or algorithm. The second protein sequence variant (protein sequence for variant 2) input into the PROVEAN model and / or algorithm can be a mutation of the query sequence. According to some embodiments, the e-value cutoff = 0.1. According to some embodiments, the e-value cutoff can be chosen for performance optimization. The next operation of the PROVEAN model and / or algorithm can comprise grouping the homologous sequence hits into clusters with at least 75% global sequence identity. While at least 75% global sequence identity’ is used for some embodiments, the clusters can be grouped with any’ suitable global sequence identity7. The next operation of the PROVEAN model and / or algorithm can comprise selecting the top N clusters that are most similar to the query' sequence to form the supporting sequence set. Any suitable number of clusters can be selected. For example, the PROVEAN model and / or algorithm can select the top 30 clusters according to some embodiments. The next operation of the PROVEAN model and / or algorithm can comprise, within each cluster, computing a delta alignment score between the query sequence and each supporting sequence and averaging the delta alignment scores. The delta alignment score can be calculated via the following equation:Regarding the delta alignment score equation above, ‘"Q” refers to the query sequence (e.g., protein sequence for variant 1) and “Q”’ refers to a mutated variant of the query' sequence Q (e.g., protein sequence for variant 2). As noted, the mutated variant and / or mutated sequence Q’ can refer to the second protein sequence variant inputted into the model and / or algorithm (protein sequence for variant 2). The term “A(Q, S)" represents the semi-global alignmentAgent Ref. No. P14786WOOO score, which is the Needleman-Wunsch global alignment score between the query7sequence Q and a support sequence, denoted as “S”, with no penalty on end gaps. Similarly, the term “A(Q’. S)” represents the semi-global alignment score, which is the Needleman-Wunsch global alignment score between the mutated variant Q’ of the query7sequence and a support sequence, denoted as “S”, with no penalty7on end gaps. The delta alignment score, denoted as “A(Q, Q', S)'’ is the difference between the semi-global alignment score of the mutated variant sequence Q' minus the semi-global alignment score of the query sequence Q. each with respect to a supporting sequence.

[0302] The next operation of the PROVEAN model and / or algorithm can comprise averaging the mean of the delta alignment scores across all clusters in the supporting set. The resulting average of the mean delta alignment score across all clusters in the supporting set is the relative fitness score between the first protein sequence variant (protein sequence for variant 1) input into the model and / or algorithm and the second protein sequence variant (protein sequence for variant 2) input into the model and / or algorithm. The relative fitness score can also be referred to as the PROVEAN score. The PROVEAN model and / or algorithm can further make a prediction based on the PROVEAN score. A predefined threshold can be set to evaluate the PROVEAN score. For example, a particular protein sequence variant could be considered deleterious if the PROVEAN score is less than or equal to the predefined threshold.

[0303] As noted above, a custom BLAST database can be constructed for use with the PROVEAN model and / or algorithm. A custom BLAST database can be constructed and / or created in any suitable manner. An example approach to constructing and / or creating a custom BLAST database can first comprise downloading raw data for plant taxonomic classification from any suitable source such as the current release of UniProtKB (see ftp.uniprot.org / pub / databases / uniprot / current_release / knowledgebase / taxonomic_divisions / (date accessed: September 9, 2024)). Then, the construction and / or creation of the custom BLAST database can further comprise converting the raw data to a suitable format. According to some embodiments, FASTA format is a suitable format. If the raw data was from the current release of UniProtKB, the raw data can be converted from SwissProt DAT format to FASTA format. According to some embodiments, this conversion can comprise using Biopython’s Bio.SwissProt library (see biopython.org / docs / latest / api / Bio.SwissProt.html (date accessed: September 9, 2024), whichAgent Ref. No. P14786WOOO is hereby incorporated by reference in its entirety ). The construction and / or creation of the custom BLAST database can further comprise clustering the raw data to 100% minimum sequence identity with 100% coverage, which serves to remove duplicates and subfragments. According to some embodiments, any suitable value of minimum sequence identity or coverage could be used. According to some embodiments, such clustering can be performed using the mmseqs easy-cluster command in MMseqs2 (see github.com / soedinglab / MMseqs2 (date accessed: September 9. 2024), which is hereby incorporated by reference in its entirety). The construction and / or creation of the custom BLAST database can further comprise filtering out all sequences with length less than or equal to a specified number of amino acids. According to some embodiments, this specified number of amino acids could be l l. According to some embodiments, such filtering could be performed using the SeqKit library (see bioinf.shenwei.me / seqkit / (date accessed: September 9, 2024), which is hereby incorporated by reference in its entirety). The construction and / or creation of the custom BLAST database can further comprise building the BLAST database. According to some embodiments, the BLAST database could be built using the makeblastdb command from the NCBI BLAST+ software tool (see blast.ncbi.nlm.nih.gov / doc / blast- help / downloadbl astdata.html (date accessed: September 9, 2024), which is hereby incorporated by reference in its entirety7). The construction and / or creation of the custom BLAST database can further comprise copying the constructed BLAST database to persistent storage in order to effectively store the BLAST database.

[0304] Figure 26 shows a workflow diagram incorporating use of the PROVEAN model and / or algorithm. The workflow diagram represents an example workflow that a user and / or a system could employ when using the PROVEAN model and / or algorithm. Such a system could be any system described herein such as the system 10. While the workflow of Figure 26 represents one example of a workflow, any suitable workflow could be employed when using the PROVEAN model and / or algorithm. Figure 26 shows that a user and / or a system can load dependencies; load input data; batch input data by reference sequence; run the PROVEAN model and / or algorithm in batches (which can include building a custom BLAST database); compute, fill, and save PROVEAN scores for protein sequence variants; and display and plot the output of the PROVEAN model and / or algorithm. According to some embodiments, the workflow diagram incorporating use of the PROVEAN model and / or algorithm of Figure 26 could be used for individual protein sequence(s) rather than forAgent Ref. No. P14786WOOO batches of protein sequence(s). For example, according to some embodiments, batching of input data is not necessary, and the model can be run on individual protein sequence(s) rather than batch(es).Embodiments

[0305] Various embodiments of the systems and methods described herein are set forth in the following set of numbered embodiments.

[0306] 1. A system for selection of a preferred guide RNA for obtaining a plant or plant part comprising a desired DNA insertion, deletion, and / or substitution (INDELS) in a target plant gene via Cas nuclease mediated gene editing, the system comprising: a processing unit; a non-transitory computer-readable medium that stores executable instructions that, when executed by the processing unit, perform operations, the operations comprising: (i) predicting a target gene editing frequency score for each of a plurality of distinct candidate guide RNAs (gRNAs) independently introduced into a plant cell with a Cas nuclease encoding DNA vector, wherein the candidate gRNAs are not associated with the Cas nuclease and the Cas nuclease is not present in the plant cell prior to their introduction; and (ii) selecting the preferred gRNA from the plurality of distinct candidate gRNAs for use in obtaining the plant or plant part comprising the desired INDELS in a target plant gene based on the target gene editing frequency score

[0307] 2. A system for selection of a preferred guide RNA for obtaining a plant or plant part comprising a desired DNA insertion, deletion, and / or substitution (INDELS) in a target plant gene encoding a protein via Cas nuclease mediated gene editing, the system comprising: a processing unit; a non-transitory computer-readable medium that stores executable instructions that, when executed by the processing unit, perform operations, the operations comprising: (i) predicting a functional effect of an INDELS obtained in a target plant gene via Cas nuclease mediated gene editing for each of a plurality of distinct candidate guide RNAs (gRNAs) independently introduced into a plant cell with a Cas nuclease encoding DNA vector, wherein the candidate gRNAs are not associated with the Cas nuclease and the Cas nuclease is not present in the plant cell prior to their introduction; and (ii) selecting the preferred gRNA based on the predicted functional effect of the INDELS.

[0308] 3. A system for selection of a preferred guide RNA for obtaining a plant or plant part comprising a desired DNA insertion, deletion, and / or substitution (INDELS) in a targetAgent Ref. No. P14786WOOO plant gene encoding a protein via Cas nuclease mediated gene editing, the system comprising: a processing unit; a non-transitory computer-readable medium that stores executable instructions that, when executed by the processing unit, perform operations, the operations comprising: (i) predicting a target gene editing frequency score for each of a plurality of distinct candidate guide RNAs (gRNAs) independently introduced into a plant cell with a Cas nuclease encoding DNA vector, wherein the candidate gRNAs are not associated with the Cas nuclease and the Cas nuclease is not present in the plant cell prior to their introduction; (ii) predicting a functional effect of an INDELS obtained in a target plant gene via Cas nuclease mediated gene editing for each of the plurality of distinct candidate gRNAs; and (iii) selecting the preferred gRNA based on the target gene editing frequency score and the predicted functional effect of the INDELS.

[0309] 4. The system of embodiments 1 or 3, wherein the candidate gRNAs and the Cas nuclease encoding DNA vector are introduced into at least one plant protoplast cell via PEG- mediated transfection and / or electroporation.

[0310] 5. The system of embodiments 1 or 3, wherein the predicting a target gene editing frequency score step further comprises filtering a set of alleles to remove allele(s) outside an expected editing window and / or to remove false positive(s).

[0311] 6. The system of embodiments 1 or 3, wherein the predicting a target gene editing frequency score step further comprises determining: (i) an average editing efficiency in the plant cell for a subset of test gRNAs directed to the target gene; and (ii) an actual editing outcome in the plant or plant part for the subset of test gRNAs.

[0312] 7. The system of embodiment 6, wherein the predicting a target gene editing frequency score step further comprises applying a linear model to correlate the average editing efficiency in the plant cell with the actual editing outcome in the plant or plant part for the subset of test gRNAs.

[0313] 8. The system of embodiment 7, wherein the linear model is configured to: (i) apply least squares regression to the correlation of the average editing efficiency in the plant cell and the actual editing outcome in the plant or plant part for each test gRNA; (ii) derive an equation to determine a line of best fit of the correlation of the average editing efficiency in the plant cell and the actual editing outcome in the plant or plant part of each test gRNA based, at least in part, on the least squares regression; and (iii) predict editing efficiency of a candidate gRNA using the derived equation.Agent Ref. No. P14786WOOO

[0314] 9. The system of any one of embodiments 3-8, wherein the selecting step comprises: (i) ranking a plurality of gRNAs based on the predicted target gene editing frequency score to obtain a subset of top-ranked gRNAs with the highest target gene editing frequency scores; (ii) identifying a desired allele among gene editing alleles produced in the plant cell by the top-ranked gRNA; (iii) selecting a gRNA which is predicted to provide the desired allele.

[0315] 10. The system of embodiment 9, wherein the number of top-ranked gRNAs is at least five.

[0316] 11. The system of embodiments 9 or 10, wherein the number of top-ranked gRNAs is at least ten.

[0317] 12. The system of any one of embodiments 3-11, wherein the predicting of the functional effect of the INDELS step further comprises selecting a subset of gRNAs based on their target gene editing frequency score.

[0318] 13. The system of embodiment 12, wherein the predicting of the functional effect of the INDELS step further comprises: inputting a set of alleles produced by the subset of gRNAs into a protein language model; processing the set of alleles via the model; outputting a ratio from the model wherein the ratio can be used to evaluate the predicted functional effect of each allele of the set of alleles, and selecting one or more gRNAs from the subset of gRNAs based on the predicted functional effect of the alleles produced by the gRNA.

[0319] 14. The system of any one of embodiments 2-13, wherein the predicting of the functional effect of the INDELS step comprises using a model that utilizes machine learning and is trained in a self-supervised manner.

[0320] 15. The system of embodiment 14, wherein evolutionary information is integrated into parameters of the model.

[0321] 16. The system of embodiments 14 or 15, wherein the predicting of the functional effect of the INDELS step further comprises inputting first and second protein sequence variants into the model.

[0322] 17. The system of embodiment 16, wherein the predicting of the functional effect of the INDELS step further comprises: (i) masking a single amino acid position of the first sequence variant; (ii) predicting a probability distribution over all possible amino acids at the single amino acid position based on surrounding sequence context and the evolutionary information; (iii) determining a conditional probability' of the single amino acid positionAgent Ref. No. P14786WOOO based on the probability distribution; (iv) performing steps (i)-(iii) for each amino acid position of a length of the first sequence variant in a one-by-one manner; and (v) calculating a product of each conditional probability of each amino acid position of the length to produce an approximated joint probability of the first sequence variant.

[0323] 18. The system of embodiments 16 or 17, wherein the predicting of the functional effect of the INDELS step further comprises: (i) masking a single amino acid position of the second sequence variant: (ii) predicting a probability distribution over all possible amino acids at the single amino acid position of the second sequence variant based on surrounding sequence context and the evolutionary information; (iii) determining a conditional probability of the single amino acid position of the second sequence variant based on the probability distribution; (iv) performing steps (i)-(iii) for each amino acid position of a length of the second sequence variant in a one-by-one manner; and (v) calculating a product of each conditional probability of each amino acid position of the length of the second sequence variant to produce an approximated j oint probability of the second sequence variant.

[0324] 19. The system of embodiment 18, wherein the predicting of the functional effect of the INDELS step further comprises calculating a quotient of the approximated joint probability of the first sequence variant divided by the approximated joint probability of the second sequence variant.

[0325] 20. The system of embodiment 18, wherein the second sequence variant differs from the first sequence variant by an in-frame INDEL such that the length of the second sequence variant differs from the length of the first sequence variant, and wherein the predicting of the functional effect of the INDELS step further comprises calculating a quotient of a geometric mean of the approximated joint probability of the first sequence variant divided by a geometric mean of the approximated joint probability of the second sequence variant.

[0326] 21. The system of embodiment 16, wherein the predicting of the functional effect of the INDELS step further comprises: (i) partitioning the first sequence variant into one or more groups of amino acid positions; (ii) masking a group of amino acid positions of the one or more groups of amino acid positions of the first sequence variant; (iii) predicting a probability distribution associated with the group of amino acid positions based on surrounding sequence context and the evolutionary information; (iv) simultaneously determining a conditional probability' of all amino acid positions of the group of amino acidAgent Ref. No. P14786WOOO positions based on the probability7distribution; (v) performing steps (ii)-(iv) for each group of the one or more groups of amino acid positions of a length of the first sequence variant in a group-by-group manner; and (vi) calculating a product of each conditional probability of each group of the one or more groups of amino acid positions of the length to produce an approximated joint probability7of the first sequence variant.

[0327] 22. The system of embodiments 16 or 21, wherein the predicting of the functional effect of the INDELS step further comprises: (i) partitioning the second sequence variant into one or more groups of amino acid positions; (ii) masking a group of amino acid positions of the one or more groups of amino acid positions of the second sequence variant; (iii) predicting a probability7distribution associated with the group of amino acid positions of the second sequence variant based on surrounding sequence context and the evolutionary7information; (iv) simultaneously determining a conditional probability of all amino acid positions of the group of amino acid positions of the second sequence variant based on the probability distribution associated with the second sequence variant; (v) performing steps (ii)- (iv) for each group of the one or more groups of amino acid positions of a length of the second sequence variant in a group-by-group manner; and (vi) calculating a product of each conditional probability' of each group of the one or more groups of amino acid positions of the length of the second sequence variant to produce an approximated joint probability7of the second sequence variant.

[0328] 23. The system of embodiment 22. wherein the predicting of the functional effect of the INDELS step further comprises calculating a quotient of the approximated joint probability' of the first sequence variant divided by the approximated joint probability of the second sequence variant.

[0329] 24. The system of claim 22, wherein the second sequence variant differs from the first sequence variant by an in-frame 1NDEL such that the length of the second sequence variant differs from the length of the first sequence variant, and wherein the predicting of the functional effect of the INDELS step further comprises calculating a quotient of a geometric mean of the approximated joint probability7of the first sequence variant divided by a geometric mean of the approximated joint probability of the second sequence variant.

[0330] 25. The system of any one of embodiment 22-24, wherein each of the one or more groups of amino acid positions of the first sequence variant and each of the one or moreAgent Ref. No. P14786WOOO groups of amino acid positions of the second sequence variant comprise random amino acid positions.

[0331] 26. The system of any one of embodiments 16-25. wherein the predicting of the functional effect of the INDELS step further comprises: (i) predicting a probability distribution over all possible amino acids at each amino acid position of the first sequence variant based on surrounding sequence context and the evolutionary' information; (ii) determining a conditional probability of each amino acid position of the first sequence variant based on the probability distribution of each amino acid position; and (iii) calculating a product of each conditional probability of each amino acid position of the first sequence variant to produce an approximated j oint probability of the first sequence variant.

[0332] 27. The system of any one of embodiments 16-26. wherein the predicting of the functional effect of the INDELS step further comprises: (i) predicting a probability distribution over all possible amino acids at each amino acid position of the second sequence variant based on surrounding sequence context and the evolutionary information; (ii) determining a conditional probability of each amino acid position of the second sequence variant based on the probability distribution of each amino acid position of the second sequence variant; and (iii) calculating a product of each conditional probability of each amino acid position of the second sequence variant to produce an approximated joint probability of the second sequence variant; wherein the second sequence variant and the first sequence variant have the same length.

[0333] 28. The system of embodiment 27, wherein the predicting of the functional effect of the INDELS step further comprises calculating a quotient of the approximated joint probability of the first sequence variant divided by the approximated joint probability of the second sequence variant.

[0334] 29. The system of any one of embodiments 2-28, wherein the predicting of the functional effect of the INDELS step further comprises using a model that utilizes machine learning and is trained in a self-supervised manner.

[0335] 30. The system of embodiment 29, wherein the predicting of the functional effect of the INDELS step further comprises inputting first and second protein sequence variants into the model.Agent Ref. No. P14786WOOO

[0336] 31. The system of embodiment 30, wherein the predicting of the functional effect of the INDELS step further comprises using the model to compute an alignment-based dimensionless score of relative fitness regarding the first and second sequence variants.

[0337] 32. The system of any one of embodiments 2-31, wherein the predicted functional effect of the INDELS and / or the preferred gRNA can be used to select and / or predict inplanta information.

[0338] 33. The system of embodiment 32. wherein the in-planta information comprises phenotype information.

[0339] 34. The system of any one of embodiments 1-33, further comprising an experimental plant, experimental plant part, or experimental plant cell lacking the desired INDELS wherein: (a) (i) the selected preferred gRNA or a polynucleotide encoding the selected preferred gRNA; and (ii) a Cas nuclease which binds the selected preferred gRNA or a polynucleotide encoding the Cas nuclease are introduced into the experimental plant, experimental plant part, or experimental plant cell lacking the desired INDELS; or (b) a ribonucleoprotein complex comprising the Cas nuclease bound to the selected preferred gRNA is introduced into the experimental plant, experimental plant part, or experimental plant cell lacking the desired INDELS.

[0340] 35. The system of embodiment 34, wherein the plant or plant part comprising the desired INDELS is obtained from the experimental plant, experimental plant part, or experimental plant cell lacking the desired INDELS.

[0341] 36. The system of embodiments 34 or 35, wherein: (i) the plant, plant part, experimental plant, and / or experimental plant part is and / or comprises a seed, leaf, stem, flower, meristem, or root; or (ii) the plant cell and / or experimental plant cell comprises plant callus tissue, embryogenic callus tissue, or a protoplast.

[0342] 37. The system of any one of embodiments 2-36, further comprising one or more DNA extraction, DNA amplification, and / or DNA sequencing device(s) configured to isolate, amplify, and / or sequence genomic DNA sequences containing the target plant gene from an experimental plant cell that had received the distinct candidate guide RNA and the Cas nuclease encoding DNA vector, wherein the isolated, amplified, and / or sequenced genomic DNA sequences can be used in the step of predicting the target gene editing frequency score and / or in the step of predicting the functional effect of the INDELS.Agent Ref. No. P14786WOOO

[0343] 38. A method for selection of a preferred guide RNA for obtaining a plant or plant part comprising a desired DNA insertion, deletion, and / or substitution (INDELS) in a target plant gene via Cas nuclease mediated gene editing, the method comprising: (i) predicting a target gene editing frequency score for each of a plurality of distinct candidate guide RNAs (gRNAs) independently introduced into a plant cell with a Cas nuclease encoding DNA vector, wherein the candidate gRNAs are not associated with the Cas nuclease and the Cas nuclease is not present in the plant cell prior to their introduction; and (ii) selecting the preferred gRNA from the plurality of distinct candidate gRNAs for use in obtaining the plant or plant part comprising the desired INDELS in a target plant gene based on the target gene editing frequency score.

[0344] 39. A method for selection of a preferred guide RNA for obtaining a plant or plant part comprising a desired DNA insertion, deletion, and / or substitution (INDELS) in a target plant gene encoding a protein via Cas nuclease mediated gene editing, the method comprising: (i) predicting a functional effect of an INDELS obtained in a target plant gene via Cas nuclease mediated gene editing for each of a plurality of distinct candidate guide RNAs (gRNAs) independently introduced into a plant cell with a Cas nuclease encoding DNA vector, wherein the candidate gRNAs are not associated with the Cas nuclease and the Cas nuclease is not present in the plant cell prior to their introduction; and (ii) selecting the preferred gRNA based on the predicted functional effect of the INDELS.

[0345] 40. A method for selection of a preferred guide RNA for obtaining a plant or plant part comprising a desired DNA insertion, deletion, and / or substitution (INDELS) in a target plant gene encoding a protein via Cas nuclease mediated gene editing, the method comprising: (i) predicting a target gene editing frequency score for each of a plurality of distinct candidate guide RNAs (gRNAs) independently introduced into a plant cell with a Cas nuclease encoding DNA vector, wherein the candidate gRNAs are not associated with the Cas nuclease and the Cas nuclease is not present in the plant cell prior to their introduction; (ii) predicting a functional effect of an INDELS obtained in a target plant gene via Cas nuclease mediated gene editing for each of the plurality of distinct candidate gRNAs; and (iii) selecting the preferred gRNA based on the target gene editing frequency score and the predicted functional effect of the INDELS.Agent Ref. No. P14786WOOO

[0346] 41. The method of embodiments 38 or 40, wherein the candidate gRNAs and the Cas nuclease encoding DNA vector are introduced into at least one plant protoplast cell via PEG-mediated transfection and / or electroporation.

[0347] 42. The method of embodiments 38 or 40, wherein the predicting a target gene editing frequency score step further comprises filtering a set of alleles to remove allele(s) outside an expected editing window and / or to remove false positive(s).

[0348] 43. The method of embodiments 38 or 40, wherein the predicting a target gene editing frequency score step further comprises determining: (i) an average editing efficiency in the plant cell for a subset of test gRNAs directed to the target gene; and (ii) an actual editing outcome in the plant or plant part for the subset of test gRNAs.

[0349] 44. The method of embodiment 43, wherein the predicting a target gene editing frequency score step further comprises applying a linear model to correlate the average editing efficiency in the plant cell with the actual editing outcome in the plant or plant part for the subset of test gRNAs.

[0350] 45. The method of embodiment 44, wherein the linear model is configured to: (i) apply least squares regression to the correlation of the average editing efficiency in the plant cell and the actual editing outcome in the plant or plant part for each test gRNA; (ii) derive an equation to determine a line of best fit of the correlation of the average editing efficiency in the plant cell and the actual editing outcome in the plant or plant part of each test gRNA based, at least in part, on the least squares regression; and (iii) predict editing efficiency of a candidate gRNA using the derived equation.

[0351] 46. The method of any one of embodiments 40-45, wherein the selecting step comprises: (i) ranking a plurality of gRNAs based on the predicted target gene editing frequency score to obtain a subset of top-ranked gRNAs with the highest target gene editing frequency scores; (ii) identifying a desired allele among gene editing alleles produced in the plant cell by the top-ranked gRNA; and (iii) selecting a gRNA which is predicted to provide the desired allele.

[0352] 47. The method of embodiment 46, wherein the number of top-ranked gRNAs is at least five.

[0353] 48. The method of embodiments 46 or 47, wherein the number of top-ranked gRNAs is at least ten.Agent Ref. No. P14786WOOO

[0354] 49. The method of any one of embodiments 40-48, wherein the predicting of the functional effect of the INDELS step further comprises selecting a subset of gRNAs based on their target gene editing frequency score.

[0355] 50. The method of embodiment 49, wherein the predicting of the functional effect of the INDELS step further comprises: inputting a set of alleles produced by the subset of gRNAs into a protein language model; processing the set of alleles via the model; outputting a ratio from the model wherein the ratio can be used to evaluate the predicted functional effect of each allele of the set of alleles, and selecting one or more gRNAs from the subset of gRNAs based on the predicted functional effect of the alleles produced by the gRNA.

[0356] 51. The method of any one of embodiments 39-50, wherein the predicting of the functional effect of the INDELS step comprises using a model that utilizes machine learning and is trained in a self-supervised manner.

[0357] 52. The method of embodiment 51, wherein evolutionary information is integrated into parameters of the model.

[0358] 53. The method of embodiments 51 or 52, wherein the predicting of the functional effect of the INDELS step further comprises inputting first and second protein sequence variants into the model.

[0359] 54. The method of embodiment 53, wherein the predicting of the functional effect of the INDELS step further comprises: (i) masking a single amino acid position of the first sequence variant; (ii) predicting a probability distribution over all possible amino acids at the single amino acid position based on surrounding sequence context and the evolutionary information; (iii) determining a conditional probability of the single amino acid position based on the probability distribution; (iv) performing steps (i)-(iii) for each amino acid position of a length of the first sequence variant in a one-by-one manner; and (v) calculating a product of each conditional probability of each ammo acid position of the length to produce an approximated joint probability of the first sequence variant.

[0360] 55. The method of embodiments 53 or 54, wherein the predicting of the functional effect of the INDELS step further comprises: (i) masking a single amino acid position of the second sequence variant: (ii) predicting a probability distribution over all possible amino acids at the single amino acid position of the second sequence variant based on surrounding sequence context and the evolutionary information; (iii) determining a conditional probability of the single amino acid position of the second sequence variant based on the probabilityAgent Ref. No. P14786WOOO distribution; (iv) performing steps (i)-(iii) for each amino acid position of a length of the second sequence variant in a one-by-one manner; and (v) calculating a product of each conditional probability of each amino acid position of the length of the second sequence variant to produce an approximated j oint probability of the second sequence variant.

[0361] 56. The method of embodiment 55, wherein the predicting of the functional effect of the INDELS step further comprises calculating a quotient of the approximated joint probability of the first sequence variant divided by the approximated joint probability of the second sequence variant.

[0362] 57. The method of embodiment 55, wherein the second sequence variant differs from the first sequence variant by an in-frame INDEL such that the length of the second sequence variant differs from the length of the first sequence variant, and wherein the predicting of the functional effect of the INDELS step further comprises calculating a quotient of a geometric mean of the approximated joint probability of the first sequence variant divided by a geometric mean of the approximated joint probability of the second sequence variant.

[0363] 58. The method of embodiment 53, wherein the predicting of the functional effect of the INDELS step further comprises: (i) partitioning the first sequence variant into one or more groups of amino acid positions; (ii) masking a group of amino acid positions of the one or more groups of amino acid positions of the first sequence variant; (iii) predicting a probability distribution associated with the group of amino acid positions based on surrounding sequence context and the evolutionary information; (iv) simultaneously determining a conditional probability of all amino acid positions of the group of amino acid positions based on the probability7distribution; (v) performing steps (ii)-(iv) for each group of the one or more groups of amino acid positions of a length of the first sequence variant in a group-by-group manner; and (vi) calculating a product of each conditional probability of each group of the one or more groups of amino acid positions of the length to produce an approximated joint probability7of the first sequence variant.

[0364] 59. The method of embodiments 53 or 58, wherein the predicting of the functional effect of the INDELS step further comprises: (i) partitioning the second sequence variant into one or more groups of amino acid positions; (ii) masking a group of amino acid positions of the one or more groups of amino acid positions of the second sequence variant; (iii) predicting a probability7distribution associated with the group of amino acid positions of theAgent Ref. No. P14786WOOO second sequence variant based on surrounding sequence context and the evolutionary information; (iv) simultaneously determining a conditional probability of all amino acid positions of the group of amino acid positions of the second sequence variant based on the probability distribution associated with the second sequence variant; (v) performing steps (ii)- (iv) for each group of the one or more groups of amino acid positions of a length of the second sequence variant in a group-by-group manner; and (vi) calculating a product of each conditional probability of each group of the one or more groups of amino acid positions of the length of the second sequence variant to produce an approximated joint probability of the second sequence variant.

[0365] 60. The method of embodiment 59, wherein the predicting of the functional effect of the INDELS step further comprises calculating a quotient of the approximated joint probability of the first sequence variant divided by the approximated joint probability of the second sequence variant.

[0366] 61. The method of embodiment 59, wherein the second sequence variant differs from the first sequence variant by an in-frame INDEL such that the length of the second sequence variant differs from the length of the first sequence variant, and wherein the predicting of the functional effect of the INDELS step further comprises calculating a quotient of a geometric mean of the approximated joint probability of the first sequence variant divided by a geometric mean of the approximated joint probability of the second sequence variant.

[0367] 62. The method of any one of embodiments 59-61, wherein each of the one or more groups of amino acid positions of the first sequence variant and each of the one or more groups of amino acid positions of the second sequence variant comprise random amino acid positions.

[0368] 63. The method of any one of embodiments 53-62, wherein the predicting of the functional effect of the INDELS step further comprises: (i) predicting a probability distribution over all possible amino acids at each amino acid position of the first sequence variant based on surrounding sequence context and the evolutionary information; (ii) determining a conditional probability of each amino acid position of the first sequence variant based on the probability distribution of each amino acid position; and (iii) calculating a product of each conditional probability of each amino acid position of the first sequence variant to produce an approximated j oint probability of the first sequence variant.Agent Ref. No. P14786WOOO

[0369] 64. The method of any one of embodiments 53-63, wherein the predicting of the functional effect of the INDELS step further comprises: (i) predicting a probability distribution over all possible amino acids at each ammo acid position of the second sequence variant based on surrounding sequence context and the evolutionary information; (ii) determining a conditional probability of each amino acid position of the second sequence variant based on the probability distribution of each amino acid position of the second sequence variant; and (iii) calculating a product of each conditional probability of each amino acid position of the second sequence variant to produce an approximated joint probability of the second sequence variant; wherein the second sequence variant and the first sequence variant have the same length.

[0370] 65. The method of embodiment 64, wherein the predicting of the functional effect of the INDELS step further comprises calculating a quotient of the approximated joint probability of the first sequence variant divided by the approximated joint probability of the second sequence variant.

[0371] 66. The method of any of embodiments 39-65. wherein the predicting of the functional effect of the INDELS step further comprises using a model that utilizes machine learning and is trained in a self-supervised manner.

[0372] 67. The method of embodiment 66, wherein the predicting of the functional effect of the INDELS step further comprises inputting first and second protein sequence variants into the model.

[0373] 68. The method of embodiment 67, wherein the predicting of the functional effect of the INDELS step further comprises using the model to compute an alignment-based dimensionless score of relative fitness regarding the first and second sequence variants.

[0374] 69. The method of any one of embodiments 39-68, wherein the predicted functional effect of the INDELS and / or the preferred guide RNA can be used to select and / or predict in-planta information.

[0375] 70. The method of embodiment 69, wherein the in-planta information comprises phenoty pe information.

[0376] 71. The method of any one of embodiments 38-70, further comprising introducing into an experimental plant, experimental plant part, or experimental plant cell lacking the desired INDELS: (a) (i) the selected preferred gRNA or a polynucleotide encoding the selected preferred gRNA; and (ii) a Cas nuclease which binds the selected preferred gRNA orAgent Ref. No. P14786WOOO a polynucleotide encoding the Cas nuclease; or (b) a ribonucleoprotein complex comprising the Cas nuclease bound to the selected preferred gRNA.

[0377] 72. The method of embodiment 71, further comprising obtaining the plant or plant part comprising the desired INDELS from the experimental plant, experimental plant part, or experimental plant cell lacking the desired INDELS.

[0378] 73. The method of embodiments 71 or 72, wherein: (i) the plant, plant part, experimental plant, and / or experimental plant part is and / or comprises a seed, leaf. stem, flower, meristem, or root; or (ii) the plant cell and / or experimental plant cell comprises plant callus tissue, embryogenic callus tissue, or a protoplast.

[0379] 74. The method of any one of embodiments 39-73, further comprising isolating, amplifying, and / or sequencing genomic DNA sequences containing the target plant gene from an experimental plant cell that had received the distinct candidate guide RNA and the Cas nuclease encoding DNA vector, wherein the isolated, amplified, and sequenced genomic DNA sequences can be used in the step of predicting the target gene editing frequency score and / or in the step of predicting the functional effect of the INDELS.EXAMPLESExample 1. Guide nomination tool

[0380] To improve guide designs for gene editing, a tool for nominating guides was developed.

[0381] The tool nominates guides based on user input. Required input information for the tool includes crop, cultivar, version, geneids, nuclease, and target modification. Crops are not limited to but include Zea Mays (com), Glycine max (soybean), Solanum lycopersicum (tomato). Oryza sativa (rice), and Triticum aestivum (wheat). Cultivar refers to which specific germplasm / cultivar or line, such as com B104 or soy NINF1170. Version specifies which genome version of the specified crop and cultivar to use. Geneid refers to the gene of interest to edit using the gene id. Geneid is specific to the annotation, not the human readable name. The nuclease includes various editing nucleases which include Cas9Casl2a, and other Type V Cas nucleases. Target modification refers to the desired gene edits, for example: promoter insertions or deletions, coding sequence knockouts, and premature stops (see FIG. 22).

[0382] Potential guides RNA spacer sequences are pre-filtered by in silico analysis to meet recommendations for obtaining high quality edits while avoiding off target effects. Off-targetAgent Ref. No. P14786WOOO sites are in some cases minimized by selecting spacers with at least 3 mismatches recommended, but in other cases a single mismatch to a target sequence is sufficient and can be used. Ambiguous bases “N” are also avoided in most cases to minimize potential off target effects. Spacer sequence content can aim for GC content in the 30-70% range and limits repetitive sequence and polymer runs (e.g., AAAA, TTTT, GGGG, CCCC, ‘ATATAT’, ‘ ACACAC’, ‘AGAGAG’, and 2-mer repeats). Poly-T is especially problematic for RNA PolIII termination, so PAM sites that are TTTT are also avoided.

[0383] Guides can be further selected from the filtered guides to meet additional recommendations which include position in the target gene for the intended purpose. Edits aimed at producing knockout alleles can be obtained in certain cases by selecting guide RNA spacers which target the first third of the CDS. Guides for Promoter Fine Tuning can target sequences spread evenly over 2+kb upstream of the Transcription start site. Guides for transcriptional enhancer insertions can be obtained in certain cases by selecting guide RNA spacers which target sequences located 100-500bp upstream of the Transcription Start Site.

[0384] Consideration can also be given for future cloning needs such as presence of restriction enzyme cut-sites. For example, cut-sites that are used in golden gate cloning of vectors, like Aarl (CACCTGC), SapI (GCTCTTC), or Esp3I (CGTCTC) can be avoided.Secondary structure can be minimized in the spacer, while the direct repeat (DR) should have the necessary structure for Cas protein recognition. The Vienna RNAfold online tool can be used to verify secondary structure.

[0385] The tool also suggested polymerase chain reaction primers assay the edit of the recommended guides. These primers are validated on wild type DNA from the crop and cultivar prior to testing the guides in protoplast so results can be assayed.Example 2. Assay to improve guide design

[0386] Testing of gRNAs in protoplasts by delivery of Cas nuclease / gRNA RNP complexes has several advantages over plasmid transformation in planta for guide testing because it is high-throughput, has a faster turn-around time, and has great scalability. However, RNP testing in protoplasts also has several limitations. The Cas nuclease is precomplexed with crRNA in excess. The Cas nuclease / gRNA RNP complex is delivered to cells in excess which gives optimal editing performance but is not reflective of Cas nuclease / gRNA RNP levels achieved with in planta plasmid transformation systems where the Cas nuclease and gRNA are delivered to the plant cells via DNA expression vectors. Hybrid delivery inAgent Ref. No. P14786WOOO protoplast allows in vivo complexing of the Cas nuclease produced by from a DNA expression vector and gRNA provided as RNA. Such hybrid delivery assays in plant protoplasts are believed to more closely resemble the conditions present in the in planta plasmid transformation systems than Cas nuclease / gRNA RNP delivery assays. (See FIG. 6)

[0387] To determine if a hybrid deliver}' is an effective form for guide selection eight guides were tested in RNP assay, hybrid delivery and explants.

[0388] Protoplast isolation

[0389] To perform the RNP and hybrid deliver}’ assays, protoplasts were isolated from tissue from healthy 10-11 day old soybean plants grown in a growth chamber. Stem tissue was collected above the cotyledons with a razor, then any true leaves and browning tissue was removed. The stems were sliced into 0.5-lmm coins and placed in a petri dish. lOmL of enzyme solution containing mannitol, cellulase RS, macroenzyme RIO, and 2-(N- morpholino)ethanesulfonic (MES) was added per 1g plant tissue. Lidded petri dishes were placed in a vacuum desiccator to help with tissue infiltration for 10 minutes in the dark. Petri dishes were wrapped in parafilm and placed in a shaking incubator for 15-17 hours at 25 RMP and 26°C. The protoplast / enzyme solution was filtered through a 40um nylon mesh filter into a conical tube. Residual protoplasts w ere rinsed from the petri dish with a buffer containing Sodium Chloride, Calcium Chloride, Potassium Chloride, and MES and placed in a second conical tube which was shaken vigorously for 10 seconds to release any protoplast still bound to the tissue. The contents of the second conical tube were filtered through the 40um filter and added to the original conical tube. The tube w as spun down and the supernatant was removed. The pellet was washed and then resuspended in the buffer twice. Cells were rested on ice or in 4°C for 30 minutes to 4 hours in the dark then spun down and resuspended in the buffer.

[0390] Soy protoplast hybrid delivery assay

[0391] Cells w ere spun dow i and supernatant was removed, the cells were resuspended at 2xlO5-lxlO6cells / mL in buffer containing mannitol, magnesium chloride, and 2-(N- morpholino)ethanesulfonic (MES). To transfect samples 1.2nmol gRNA and 20ug plasmid DNA was added to 200uL of the cell suspension and mixed by gentle tapping. 40% PEG C (0.25M Mannitol, 0.2M CaCh, 40% PEG 4000) was gently added at a 1: 1 volume and mixed by gentle tapping. The reaction was incubated at room temperature for 15 minutes before W5 solution containing sodium chloride, potassium chloride, and MES was added to stop theAgent Ref. No. P14786WOOOPEG reaction. The samples were spun down and resuspended in W5. Cells were plated and incubated for 48 hours at 28°C in the dark. Cells were harvested by pelleting at 200g for 5 min. Supernatant was removed and the cell pellet was frozen on dry ice.

[0392] Soy protoplast RNP assay

[0393] Cells were spun down and supernatant was removed, the cells were resuspended at 2X105-1X106cells / mL in buffer containing 0.6M mannitol, 15mM magnesium chloride, and 4mM 2-(N-morpholino)ethanesulfonic (MES), pH5.7. To transfect samples, 20uL of an RNP mixture (e.g., gRNA complexed with the Cas nuclease with Salmon Sperm DNA carrier) was added to 200uL of the cell suspension and mixed by gentle tapping. 40% PEG C (0.25M Mannitol, 0.2M CaCE. 40% PEG 4000) was gently added at a 1 : 1 volume and mixed bygentle tapping. The reaction was incubated at room temperature for 15 minutes before at least 880 pL W5 solution (154mM NaCl, 125mM CaC12, 5mM KC1, 2mM MES pH5.7) was added to stop the PEG reaction. The samples were spun down and resuspended in W5. Cells were plated and incubated for 48 hours at 28°C in the dark. Cells were harvested by pelleting at 200g for 5 min. Supernatant was removed and the cell pellet was frozen on dry ice.

[0394] Soy in planta plasmid transformation assay

[0395] Surface sterilized soybean seed previously stored at about 4°C were imbibed either in water for 16-18 hours or imbibed for 6 to 7 hours on solid imbibition media. Explants were prepared by splitting the imbibed seeds on the longitudinal axis with the hypocotyl facing downward. The radicle attached to each cotyledon is trimmed to between 0.75 and 1mm in length and any primary leaves are removed with a scalpel. The explant is stored in water.

[0396] Explants were subjected to Agrobacterium infection and co-cultivation with Agrobacterium containing the gRNA and Cas nuclease expression vectors along with controls which lack the vectors. Once the A. tumefaciens is resuspended in the infection media and brought to an OD of 0.6 (at 600nM). 50ml of the tumefaciens solution was pipetted directly onto sterilized dry seeds in 100x25ml petri dishes. The petri dishes were then wrapped twice in parafdm and placed inside a black Flambeau box. The Flambeau box was placed on a plate shaker overnight (about 20 hours). Plates are then removed from the Flambeau box and the A. tumefaciens solution was pipetted out and devitalized. Seeds are then cut and placed on 2 disks of filter paper covered with 2ml co-cultivation (CC) media in 100x25ml petri dishes as per the standard procedure and placed in light at a 24-hourAgent Ref. No. P14786WOOO light / dark photoperiod, wherein about 6 to 10 hours of the photoperiod are dark (e.g., 16 hour / 8 hour day / night photoperiod) for 5 days.

[0397] The explants were transferred from the co-cultivation media to a shoot induction media (SIM). After 5 days of co-cultivation, the half-seed explants were transferred to SIM containing glyphosate (15 mg / L). For Shoot induction, 25x100mm plates can be used with 50 mL media per plate. Coty ledons were submerged into the media at a 45° to 90° angle with the hypocotyl end down. Unwrapped petri dishes were placed in the transparent Flambeau boxes and incubated at 27°C, a 24-hour light / dark photoperiod, wherein about 6 to 10 hours of the photoperiod are dark (e.g., 16 hour / 8 hour day / night photoperiod) for a total of 2 weeks.

[0398] Explants with shoots were transferred from SIM to Shoot Elongation Media (SEM) containing Glyphosate (15 mg / L) and then cultivated in SEM 1 for about 2 weeks at 27°C. 16 / 8 (light / dark). Explants were then sampled after 2 weeks on SEMI for gene edits.

[0399] Imbibition MediaMedia Name > Imbibition MediaIngredients(stock concentration)],MS Modified Basal 2.22 g / LMedium withGamborg vitaminsSucrose 20g / LMES 0.59 g / LNoble Agar 5 g / L pH 5.7

[0400] Infection, CC, and SIM MediaName — Infection CC media SIM- mediaIngredients(stock concentration)Agent Ref. No. P14786WOOO 5.4 5.4 5.7 LAVE 0.25 mg / L 0.25 mg / L 1.67 mg / L 1.67mg / L 1.11 mg / LAcetosy ringone 40.00 mg / L 40.00 mg / L*(20mg / mL,DMSO)L-Cysteine 400.00 mg / L(50mg / mL)DTT 154.20 mg / L(50mg / mL)Glyphosate 15.00 mg / L(lOmg / mL)Timentin 50.00 mg / L(lOOmg / mL)Vancomycin 50.00 mg / L(50mg / mL)Cefotaxime (100 mg / L) 200.00 mg / L

[0401] SEMPRE-AUTOCLAVE MS Modified Basal 4.44 g / L Medium with Gamborg Vitamins Sucrose 40.00 g / L MES 0.59 g / LNoble Agar 7.00 g / L pH 5.7POST-AUTOCLAVEL- Asparagine (50 50.00 mg / L mg / mL) L-Pyroglutamic acid 100 mg / L (lOOmg / mL) IAA (Img / mL) 0.10 mg / L GA3 (Img / mL) 0.50 mg / L trans-Zeatin riboside 1.00 mg / L (Img / mL) IBA (Img / mL) Glyphosate (lOmg / mL) 15.00 mg / L Timentin* 100.00 mg / L (lOOmg / mL) Vancomycin 50.00 mg / L (50mg / mL)Agent Ref. No. P14786WOOOCefotaxime 200.00 mg / L(lOOmg / mL)

[0402] Library Preparation

[0403] Genomic DNA (gDNA) was isolated from the frozen cell pellets as described in the Maxwell RSC Plant DNA Kit technical manual revised 10 / 21 located on the world wide web at “www.promega.com / resources / protocols / technical-manuals / 101 / maxwell-rsc-plant-dna- kit-protocol / '’ (date accessed: September 9, 2024) (which is hereby incorporated by reference in its entirety) accessed on June 6. 2024. Short amplicons were generated from the gDNA by polymerase chain reaction using Phusion Flash High-Fidelity PCR Master Mix and custom primers. Libraries were created by adding adapters to the amplicons by polymerase chain reaction using KAPA HiFi HotStart Ready Mix. These libraries were pooled, and the resulting product was cleaned using KAPA SPRI Pure Beads as described in KAPA Pure Beads Technical Data sheet located on the world wide web at “elabdoc- prod. roche.com / eLD / api / downloads / 88253649-67 Ob-ee 11 - 1 c91 - 005056a772fd?countryIsoCode=pi” (date accessed: September 9, 2024) (which is hereby incorporated by reference in its entirety) accessed on July 24, 2024. The library' was quantified using a Qubit 3.0 Fluorometer and checked for quality using a D5000 ScreenTape Assay on a TapeStation System D5000.

[0404] Correlation between hybrid delivery assay results and editing efficiency in explants

[0405] Eight guides were tested in hybrid delivery’ and in in planta plasmid transformation assays (e.g, explants). Results show correlation between hybrid delivery assay results and editing efficiency in explants (see FIG. 7). The hybrid delivery assay results correlate with the proportion of high quality edits (see FIG. 8A) and the specific edits found (see FIG. 8B) in explants. Results show the hybrid delivery assay can be used to also predict in-planta editing efficiency (see FIG. 9A & B). Furthermore, hybrid delivery assay can be used to predict the best guides for further use in explants and plants (see FIGs. 10 & 11).

[0406] Optimizing vectors with the highest efficiency guides for each target can improve multiplex editing efficiency. This is important because using an early sampling pipeline for selecting the top performing guides is resource intense. Hybrid delivery guide testing enablesAgent Ref. No. P14786WOOO increased multiplexing efficiency faster and can test for the best combinations of guides for multiplexing in a single vector see FIG. 12).

[0407] Data shows that population level editing efficiency is also correlated between protoplasts and plants (see FIG. 13). Data also shows that there is overlap between the top 5 alleles in hybrid delivery' assay results and the top 10 most frequently occurring alleles in plants (see FIGs. 14A-C). Data also shows that the majority' of high quality' alleles are found in the top 5 hybrid delivery results (see FIG. 15).Example 3. Coding Variant Effect Prediction with ESM-lv

[0408] Given a pair of protein sequences, a dimensionless score of relative fitness between the two variants was computed using ESM-lv, an open-source transformer based protein language model (pLM) (see FIG. 18). This score indicated the relative likelihood of sampling a protein sequence at random from a large distribution of naturally occurring proteins (z.e., Uniref90) before and after a mutation, and therefore how much more likely' it is to resemble a functional protein. The relative fitness score can be used to estimate the impact of the mutation which takes variant 1 variant 2. The mutation in question may be a SNP, an INDELS, or a more complicated variant.

[0409] Model parameters included mask_size, mask_type, and model_name. Mask_size is the number of amino acids to mask simultaneously in each scoring iteration. Larger mask_size values lead to lower accuracy but are faster. For SNP variants mask size can be set to 0 to use the wild-type scoring strategy’ which is much faster.

[0410] The model outputs a score ref, score, and score delta. The score ref is the absolute fitness score for protein sequence variant 1. The score is the absolute fitness score for protein sequence variant 2. Score_delta is the relative fitness score between protein sequence variants 1 and 2.

[0411] The relative fitness score varies based on the type of mutations. The default relative fitness score is pseudolikelihood with 2 times length forward passes. For a sequence of length L, the identity’ of the amino acid at each position can be masked one-by-one to produce L sequences, and the pLM can be used on each to predict the conditional probability of each masked amino acid. Taking the product of the independent conditional probabilities of each amino acid results in an approximation of the joint probability of the full sequence conditioned only on the evolutionary’ information E integrated into the model parameters P(S|E)=IIi P(Si|S-i, E). This approximated joint probability can be interpreted as aAgent Ref. No. P14786WOOO pseudolikelihood of drawing the sequence from the dataset the pLM was trained on. By repeating this procedure on another sequence S’, the pseudolikelihood ratioPLR = P(S|E) / P(S’|E) can be computed. If S is more functional, then it is more likely to be favored by evolution and thereby represented in a large set of natural sequences, so that PLR > 1. Similarly, if S’ is more functional, then PLR < 1.

[0412] If the mutation is a substitution, the positions of all the residues remain the same and the wildtype marginal strategy can be used. This is very similar to the pseudolikelihood strategy except that all scores are conditioned only on the full wildtype sequence without any masking, and therefore it only requires one forw ard pass through the model to compute.

[0413] The conditional probabilities of both the original and substituted amino acid can be determined from the distribution P(Si|S, E). This can be done for each position, and a pseudolikelihood can be computed as before by summing over the positionsP(S|E)=IIi P(Si|S, E). Then, a pseudolikelihood ratio can be computed as beforePLR = P(S|E) / P(S’|E). Because the non-mutated residues and the sequence we condition on are the same in both numerator and denominator, they cancel out, and as a result, we need not consider non-substitution positions in the pseudolikelihood sum. As before, If S is more functional, then it is more likely to be favored by evolution and thereby represented in a large set of natural sequences, so that PLR > 1. Similarly, if S’ is more functional, the PLR < 1.

[0414] When protein sequences S and S’ differ by an in-frame INDELS, they may have two different lengths L and L’ but will otherwise have mostly the same amino acid content. However, the pseudolikelihood of each sequence is computed by taking a product of the conditional probabilities of each of its amino acids, and so the longer sequence will have more terms in this product. Since each probability is by definition smaller than or equal to one, the longer sequence will be biased towards having a smaller pseudolikelihood.

[0415] In order to compare the pseudolikelihoods of the two sequences equitably, then, we must apply some transformation that accounts for the multiplicative nature of independentAgent Ref. No. P14786WOOO probabilities and normalizes for the different number of events (i.e., residues in each sequence). This is done by computing the geometric mean of conditional probabilities rather than the product P(S|E) = Lx / lli P(Si|S-i, E) and P(S’|E) = L'^IIi P(S'i|S'-i, E). Effectively, this can be interpreted as each residue of S contributing 1 / L of the P(S|E) and each residue of S' contributing 1 / L' of P(S'|E).

[0416] Validating ESM-lv predictions

[0417] ESM-vl predicted results were compared with cell proliferation assay results measuring the effect of 314 single amino deletions in the PTEN tumor suppressor gene in humanized yeast cells (see FIGs. 19 and / or 20A). Results showed correlation between predicted and observed effect of PTEN single amino acid deletions, p = 0.576, p value = 4e" 29

[0418] ESM-vl predicted results were compared with cell proliferation assay results measuring the effect of single and double amino acid deletions in the P53_HUMAN gene (obtained from ProteinGym Indel dataset found at huggingface.co / datasets / ICML2022 / ProteinGym / blob / refs%2Fconvert%2Fparquet / ProteinGy m_reference_file_indels.csv (date accessed: September 9, 2024), which is hereby incorporated by reference in its entirety) in humanized yeast cells (see FIG. 20B). Results showed correlation between predicted and observed effect of single and double amino acid deletions, p = 0. 199, p value = 2e’04.

[0419] ESM-vl predicted results were compared with viral adeno-associated virus (AAV) capsid production measuring the effect of insertions ranging from 1 to 14 amino acids in the CAPSD_AAV2S gene (obtained from ProteinGym Indel dataset found at huggingface. co / datasets / ICML2022 / ProteinGym / blob / refs%2F convert%2F parquet / ProteinGy m_reference_file_indels.csv (date accessed: September 9, 2024)) (see FIG. 20C). Results showed variations in INDELS size can be a confounding factor for the model, particularly for large insertions, p = 0.512, p value = 2e’16. This problem is mitigated by the fact that editing experiments are more likely to result in deletions.

[0420] ESM-vl predicted results were compared with enzyme activity measuring the effect of one or more INDELS ranging from approximately -40 to 80 amino acids in theAgent Ref. No. P14786WOOOA0A1 J4YT16 9PROT (obtained from ProteinGym Indel dataset found at huggingface. co / datasets / ICML2022 / ProteinGy m / blob / refs%2F convert%2Fparquet / ProteinGy m_reference_file_indels.csv (date accessed: September 9. 2024)) (see FIG. 20D). Results showed variations in INDELS size can be a confounding factor for the model, particularly for a large INDELS and can be interpreted as outliers, p = 0.202, p value = 4e'02.Example 4. Coding Variant Effect Prediction with PROVEAN

[0421] Given a pair of protein sequences, a dimensionless score of relative fitness between the two variants was computed using PROVEAN (Protein Variation Effect Analyzer) which is an open-source algorithm and software tool This dimensionless alignment-based score measures the relative change in sequence similarity of a query sequence to a protein sequence homolog before and after a mutation of the query sequence. The relative fitness score was used to estimate the impact of the mutation which takes variant 1 variant 2 on biological function. The mutation in question may be a SNP, an INDELS, or a more complicated variant.

[0422] Model parameters included: blastdb, blastdb dir, and save interval. Blastdb is the name of the protein BLAST database used to create multiple sequence alignments for each of the protein variants. The default database is a custom plant-specific protein database built using data from UniProtKB and clustered to 100% minimum sequence similarity to remove duplicate and subfragment sequences. Blastdb_dir is the directory where BLAST databases are stored. Save interval is the number of sequence pair batches to process before saving.

[0423] Input data values included: id ref, id alt, sequence ref, and sequence alt. Id ref is the protein ID for variant 1 (i.e., reference or wild-type sequence). Id alt is the protein ID for variant 2. Sequence_ref and sequence_alt are the protein sequences for variant 1 and variant 2, respectively.

[0424] The output (pro vean_s core) is presented as the relative fitness between protein variants 1 and 2.

[0425] The PROVEAN algorithm then proceeds based on the following steps. Using the BLAST alignment algorithm along with a specified BLAST database (i. e. , blastdb), the algorithm gathers homologous protein sequences that are similar to the query sequence. The algorithm groups the homologous sequence hits into clusters with at least 75% global sequence identity. The algorithm selects the top 30 clusters that are most similar to the query' sequence to form the supporting sequence set. Within each cluster the delta alignment scoreAgent Ref. No. P14786WOOO was computed between the query sequence and each supporting sequence, and then averaged. Given a query sequence Q, a mutated variant is given by Q'. The semi-global alignment score A(Q, S) is the Needleman-Wunsch global alignment score between the query sequence and a support sequence S with no penalty on end gaps. The delta alignment score is the difference between the semi-global alignment scores of the mutated and query' sequences, each with respect to a supporting sequence. The algorithm then averages the mean delta alignment scores across each all clusters in the supporting sequence set. The result is the provean score.

[0426] Therefore, as understood from the present disclosure, the system(s) and / or method(s) disclosed herein can be used to efficiently and effectively predict a target gene editing frequency score for a plurality of candidate guide RNAs, predict functional effect(s) of one or more INDELS, and / or select at least one preferred guide RNA based on the predicted target gene editing frequency scores and / or the predicted functional effect(s) of the one or more INDELS. The system(s) and / or method(s) described herein further provide for the use of a hybrid delivery assay' to perform gene editing, which improves the efficiency and efficacy of such editing. Thus, the present disclosure allows a researcher to conduct genetic editing experiments and to perform genetic editing of plants in a more efficient and effective manner than what is known in the prior art.

[0427] From the foregoing, it can be seen that the present disclosure accomplishes at least all of the stated objectives.

[0428] All cited patents and patent publications referred to in this application are incorporated herein by reference in their entirety'. All of the materials and methods disclosed and claimed herein can be made and used without undue experimentation as instructed by the above disclosure and illustrated by the examples. Although the materials and methods of this disclosure have been described in terms of embodiments and illustrative examples, it will be apparent to those of skill in the art that substitutions and variations can be applied to the materials and methods described herein without departing from the concept, spirit, and scope of the disclosure. All such similar substitutes and modifications apparent to those skilled in the art are deemed to be within the spirit, scope, and concept of the disclosure as encompassed by the embodiments of the disclosures recited herein and the specification and appended claims.

Claims

Agent Ref. No. P14786WOOOCLAIMSWhat is claimed is:

1. A system for selection of a preferred guide RNA for obtaining a plant or plant part comprising a desired DNA insertion, deletion, and / or substitution (INDELS) in a target plant gene via Cas nuclease mediated gene editing, the system comprising: a processing unit; a non-transitory computer-readable medium that stores executable instructions that, when executed by the processing unit, perform operations, the operations comprising:(i) predicting a target gene editing frequency score for each of a plurality of distinct candidate guide RNAs (gRNAs) independently introduced into a plant cell with a Cas nuclease encoding DNA vector, wherein the candidate gRNAs are not associated with the Cas nuclease and the Cas nuclease is not present in the plant cell prior to their introduction; and(ii) selecting the preferred gRNA from the plurality of distinct candidate gRNAs for use in obtaining the plant or plant part comprising the desired INDELS in a target plant gene based on the target gene editing frequency score.

2. A system for selection of a preferred guide RNA for obtaining a plant or plant part comprising a desired DNA insertion, deletion, and / or substitution (INDELS) in a target plant gene encoding a protein via Cas nuclease mediated gene editing, the system comprising: a processing unit; a non-transitory computer-readable medium that stores executable instructions that, when executed by the processing unit, perform operations, the operations comprising:(i) predicting a functional effect of an INDELS obtained in a target plant gene via Cas nuclease mediated gene editing for each of a plurality of distinct candidate guide RNAs (gRNAs) independently introduced into a plant cell with a Cas nuclease encoding DNA vector, wherein the candidate gRNAs are not associated with the Cas nuclease and the Cas nuclease is not present in the plant cell prior to their introduction; and(ii) selecting the preferred gRNA based on the predicted functional effect of theINDELS.Agent Ref. No. P14786WOOO3. A system for selection of a preferred guide RNA for obtaining a plant or plant part comprising a desired DNA insertion, deletion, and / or substitution (INDELS) in a target plant gene encoding a protein via Cas nuclease mediated gene editing, the system comprising: a processing unit; a non -transitory computer-readable medium that stores executable instructions that, when executed by the processing unit, perform operations, the operations comprising:(i) predicting a target gene editing frequency score for each of a plurality of distinct candidate guide RNAs (gRNAs) independently introduced into a plant cell with a Cas nuclease encoding DNA vector, wherein the candidate gRNAs are not associated with the Cas nuclease and the Cas nuclease is not present in the plant cell prior to their introduction;(ii) predicting a functional effect of an INDELS obtained in a target plant gene viaCas nuclease mediated gene editing for each of the plurality of distinct candidate gRNAs; and(iii) selecting the preferred gRNA based on the target gene editing frequency score and the predicted functional effect of the INDELS.

4. The system of claim 1 or 3, wherein the candidate gRNAs and the Cas nuclease encoding DNA vector are introduced into at least one plant protoplast cell via PEG-mediated transfection and / or electroporation.

5. The system of claim 1 or 3, wherein the predicting a target gene editing frequency score step further comprises filtering a set of alleles to remove allele(s) outside an expected editing window and / or to remove false positive(s).

6. The system of claim 1 or 3, wherein the predicting a target gene editing frequency score step further comprises determining:(i) an average editing efficiency in the plant cell for a subset of test gRNAs directed to the target gene; and(ii) an actual editing outcome in the plant or plant part for the subset of test gRNAs.Agent Ref. No. P14786WOOO7. The system of claim 6, wherein the predicting a target gene editing frequency score step further comprises applying a linear model to correlate the average editing efficiency in the plant cell with the actual editing outcome in the plant or plant part for the subset of test gRNAs.

8. The system of claim 7, wherein the linear model is configured to:(i) apply least squares regression to the correlation of the average editing efficiency in the plant cell and the actual editing outcome in the plant or plant part for each test gRNA;(ii) derive an equation to determine a line of best fit of the correlation of the average editing efficiency in the plant cell and the actual editing outcome in the plant or plant part of each test gRNA based, at least in part, on the least squares regression; and(iii) predict editing efficiency of a candidate gRNA using the derived equation.

9. The system of claim 3. wherein the selecting step comprises:(i) ranking a plurality of gRNAs based on the predicted target gene editing frequency score to obtain a subset of top-ranked gRNAs with the highest target gene editing frequency scores;(ii) identifying a desired allele among gene editing alleles produced in the plant cell by the top-ranked gRNA;(iii) selecting a gRNA which is predicted to provide the desired allele.

10. The system of claim 9, wherein the number of top-ranked gRNAs is at least five.1 1. The system of claim 9, wherein the number of top-ranked gRNAs is at least ten.

12. The system of claim 3, wherein the predicting of the functional effect of the INDELS step further comprises selecting a subset of gRNAs based on their target gene editing frequency score.Agent Ref. No. P14786WOOO13. The system of claim 12, wherein the predicting of the functional effect of the INDELS step further comprises: inputting a set of alleles produced by the subset of gRNAs into a protein language model; processing the set of alleles via the model; outputting a ratio from the model wherein the ratio can be used to evaluate the predicted functional effect of each allele of the set of alleles, and selecting one or more gRNAs from the subset of gRNAs based on the predicted functional effect of the alleles produced by the gRNA.

14. The system of claim 2 or 3, wherein the predicting of the functional effect of the INDELS step comprises using a model that utilizes machine learning and is trained in a selfsupervised manner.

15. The system of claim 14, wherein evolutionary information is integrated into parameters of the model.

16. The system of claim 1 , wherein the predicting of the functional effect of the INDELS step further comprises inputting first and second protein sequence variants into the model.

17. The system of claim 16, wherein the predicting of the functional effect of the INDELS step further comprises:(i) masking a single amino acid position of the first sequence variant;(ii) predicting a probability distribution over all possible amino acids at the single amino acid position based on surrounding sequence context and the evolutionary information;(iii) determining a conditional probability of the single amino acid position based on the probability distribution;(iv) performing steps (i)-(iii) for each amino acid position of a length of the first sequence variant in a one-by-one manner; andAgent Ref. No. P14786WOOO(v) calculating a product of each conditional probability of each amino acid position of the length to produce an approximated joint probability of the first sequence variant.

18. The system of claim 17, wherein the predicting of the functional effect of the INDELS step further comprises:(i) masking a single amino acid position of the second sequence variant;(ii) predicting a probability distribution over all possible amino acids at the single amino acid position of the second sequence variant based on surrounding sequence context and the evolutionary information;(iii) determining a conditional probability of the single amino acid position of the second sequence variant based on the probability distribution;(iv) performing steps (i)-(iii) for each amino acid position of a length of the second sequence variant in a one-by-one manner; and(v) calculating a product of each conditional probability of each amino acid position of the length of the second sequence variant to produce an approximated joint probability of the second sequence variant.

19. The system of claim 18, wherein the predicting of the functional effect of the INDELS step further comprises calculating a quotient of the approximated joint probability of the first sequence variant divided by the approximated joint probability of the second sequence variant.

20. The system of claim 18, wherein the second sequence variant differs from the first sequence variant by an in-frame INDELS such that the length of the second sequence vanant differs from the length of the first sequence variant, and wherein the predicting of the functional effect of the INDELS step further comprises calculating a quotient of a geometric mean of the approximated joint probability of the first sequence variant divided by a geometric mean of the approximated joint probability of the second sequence variant.

21. The system of claim 16, wherein the predicting of the functional effect of the INDELS step further comprises:Agent Ref. No. P14786WOOO(i) partitioning the first sequence variant into one or more groups of amino acid positions;(ii) masking a group of amino acid positions of the one or more groups of amino acid positions of the first sequence variant;(iii) predicting a probability7distribution associated wi th the group of amino acid positions based on surrounding sequence context and the evolutionary information;(iv) simultaneously determining a conditional probability of all amino acid positions of the group of amino acid positions based on the probability distribution;(v) performing steps (ii)-(iv) for each group of the one or more groups of amino acid positions of a length of the first sequence variant in a group-by-group manner; and(vi) calculating a product of each conditional probability of each group of the one or more groups of amino acid positions of the length to produce an approximated joint probability of the first sequence variant.

22. The system of claim 21, wherein the predicting of the functional effect of the INDELS step further comprises:(i) partitioning the second sequence variant into one or more groups of amino acid positions;(ii) masking a group of amino acid positions of the one or more groups of amino acid positions of the second sequence variant;(iii) predicting a probability distribution associated with the group of amino acid positions of the second sequence variant based on surrounding sequence context and the evolutionary information;(iv) simultaneously determining a conditional probability of all amino acid positions of the group of amino acid positions of the second sequence variant based on the probability distribution associated with the second sequence variant;(v) performing steps (ii)-(iv) for each group of the one or more groups of amino acid positions of a length of the second sequence variant in a group-by-group manner; andAgent Ref. No. P14786WOOO(vi) calculating a product of each conditional probability of each group of the one or more groups of amino acid positions of the length of the second sequence variant to produce an approximated joint probability of the second sequence variant.

23. The system of claim 22, wherein the predicting of the functional effect of the INDELS step further comprises calculating a quotient of the approximated joint probability of the first sequence variant divided by the approximated joint probability of the second sequence variant.

24. The system of claim 22, wherein the second sequence variant differs from the first sequence variant by an in-frame INDELS such that the length of the second sequence variant differs from the length of the first sequence variant, and wherein the predicting of the functional effect of the INDELS step further comprises calculating a quotient of a geometric mean of the approximated joint probability’ of the first sequence variant divided by a geometric mean of the approximated joint probability of the second sequence variant.

25. The system of claim 22, wherein each of the one or more groups of amino acid positions of the first sequence variant and each of the one or more groups of amino acid positions of the second sequence variant comprise random amino acid positions.

26. The system of claim 16, wherein the predicting of the functional effect of the INDELS step further comprises:(i) predicting a probability distribution over all possible amino acids at each amino acid position of the first sequence variant based on surrounding sequence context and the evolutionary information;(ii) determining a conditional probability of each amino acid position of the first sequence variant based on the probability' distribution of each amino acid position; and(iii) calculating a product of each conditional probability of each amino acid position of the first sequence variant to produce an approximated joint probability of the first sequence variant.Agent Ref. No. P14786WOOO27. The system of claim 26, wherein the predicting of the functional effect of the INDELS step further comprises:(i) predicting a probability distribution over all possible amino acids at each amino acid position of the second sequence variant based on surrounding sequence context and the evolutionary information;(ii) determining a conditional probability of each amino acid position of the second sequence variant based on the probability distribution of each amino acid position of the second sequence variant; and(iii) calculating a product of each conditional probability' of each amino acid position of the second sequence variant to produce an approximated joint probability’ of the second sequence variant; wherein the second sequence variant and the first sequence variant have the same length.

28. The system of claim 27, wherein the predicting of the functional effect of the INDELS step further comprises calculating a quotient of the approximated joint probability of the first sequence variant divided by the approximated joint probability of the second sequence variant.

29. The system of claim 2 or 3, wherein the predicting of the functional effect of the INDELS step further comprises using a model that utilizes machine learning and is trained in a self-supervised manner.

30. The system of claim 29, wherein the predicting of the functional effect of the INDELS step further comprises inputting first and second protein sequence variants into the model.

31. The system of claim 30, wherein the predicting of the functional effect of the INDELS step further comprises using the model to compute an alignment-based dimensionless score of relative fitness regarding the first and second sequence variants.Agent Ref. No. P14786WOOO32. The system of claim 2 or 3, wherein the predicted functional effect of the INDELS and / or the preferred gRNA can be used to select and / or predict in-planta information.

33. The system of claim 32, wherein the in-planta information comprises phenotype information.

34. The system of any one of claims 1, 2, or 3, further comprising an experimental plant, experimental plant part, or experimental plant cell lacking the desired INDELS wherein:(a)(i) the selected preferred gRNA or a polynucleotide encoding the selected preferred gRNA; and(ii) a Cas nuclease which binds the selected preferred gRNA or a polynucleotide encoding the Cas nuclease are introduced into the experimental plant, experimental plant part, or experimental plant cell lacking the desired INDELS; or(b) a ribonucleoprotein complex comprising the Cas nuclease bound to the selected preferred gRNA is introduced into the experimental plant, experimental plant part, or experimental plant cell lacking the desired INDELS.

35. The system of claim 34, wherein the plant or plant part comprising the desired INDELS is obtained from the experimental plant, experimental plant part, or experimental plant cell lacking the desired INDELS.

36. The system of claim 34, wherein:(i) the plant, plant part, experimental plant, and / or experimental plant part is and / or comprises a seed, leaf, stem, flower, meristem, or root or(ii) the plant cell and / or experimental plant cell comprises plant callus tissue, embryogenic callus tissue, or a protoplast.

37. The system of claim 2 or 3, further comprising one or more DNA extraction, DNA amplification, and / or DNA sequencing device(s) configured to isolate, amplify, and / or sequence genomic DNA sequences containing the target plant gene from an experimentalAgent Ref. No. P14786WOOO plant cell that had received the distinct candidate guide RNA and the Cas nuclease encoding DNA vector, wherein the isolated, amplified, and / or sequenced genomic DNA sequences can be used in the step of predicting the target gene editing frequency score and / or in the step of predicting the functional effect of the INDELS.

38. A method for selection of a preferred guide RNA for obtaining a plant or plant part comprising a desired DNA insertion, deletion, and / or substitution (INDELS) in a target plant gene via Cas nuclease mediated gene editing, the method comprising:(i) predicting a target gene editing frequency score for each of a plurality of distinct candidate guide RNAs (gRNAs) independently introduced into a plant cell with a Cas nuclease encoding DNA vector, wherein the candidate gRNAs are not associated with the Cas nuclease and the Cas nuclease is not present in the plant cell prior to their introduction; and(ii) selecting the preferred gRNA from the plurality of distinct candidate gRNAs for use in obtaining the plant or plant part comprising the desired INDELS in a target plant gene based on the target gene editing frequency score.

39. A method for selection of a preferred guide RNA for obtaining a plant or plant part comprising a desired DNA insertion, deletion, and / or substitution (INDELS) in a target plant gene encoding a protein via Cas nuclease mediated gene editing, the method comprising:(i) predicting a functional effect of an INDELS obtained in a target plant gene via Cas nuclease mediated gene editing for each of a plurality of distinct candidate guide RNAs (gRNAs) independently introduced into a plant cell with a Cas nuclease encoding DNA vector, wherein the candidate gRNAs are not associated with the Cas nuclease and the Cas nuclease is not present in the plant cell prior to their introduction; and(ii) selecting the preferred gRNA based on the predicted functional effect of theINDELS.

40. A method for selection of a preferred guide RNA for obtaining a plant or plant part comprising a desired DNA insertion, deletion, and / or substitution (INDELS) in a target plant gene encoding a protein via Cas nuclease mediated gene editing, the method comprising:Agent Ref. No. P14786WOOO(i) predicting a target gene editing frequency score for each of a plurality of distinct candidate guide RNAs (gRNAs) independently introduced into a plant cell with a Cas nuclease encoding DNA vector, wherein the candidate gRNAs are not associated with the Cas nuclease and the Cas nuclease is not present in the plant cell prior to their introduction;(ii) predicting a functional effect of an INDELS obtained in a target plant gene viaCas nuclease mediated gene editing for each of the plurality of distinct candidate gRNAs; and(iii) selecting the preferred gRNA based on the target gene editing frequency score and the predicted functional effect of the INDELS.

41. The method of claim 38 or 40, wherein the candidate gRNAs and the Cas nuclease encoding DNA vector are introduced into at least one plant protoplast cell via PEG-mediated transfection and / or electroporation.

42. The method of claim 38 or 40. wherein the predicting a target gene editing frequency score step further comprises filtering a set of alleles to remove allele(s) outside an expected editing window and / or to remove false positive(s).

43. The method of claim 38 or 40, wherein the predicting a target gene editing frequency score step further comprises determining:(i) an average editing efficiency in the plant cell for a subset of test gRNAs directed to the target gene; and(ii) an actual editing outcome in the plant or plant part for the subset of test gRNAs.

44. The method of claim 43, wherein the predicting a target gene editing frequency score step further comprises applying a linear model to correlate the average editing efficiency in the plant cell with the actual editing outcome in the plant or plant part for the subset of test gRNAs.Agent Ref. No. P14786WOOO45. The method of claim 44, wherein the linear model is configured to:(i) apply least squares regression to the correlation of the average editing efficiency in the plant cell and the actual editing outcome in the plant or plant part for each test gRNA;(ii) derive an equation to determine a line of best fit of the correlation of the average editing efficiency in the plant cell and the actual editing outcome in the plant or plant part of each test gRNA based, at least in part, on the least squares regression; and(iii) predict editing efficiency of a candidate gRNA using the derived equation.

46. The method of claim 40, wherein the selecting step comprises:(i) ranking a plurality of gRNAs based on the predicted target gene editing frequency score to obtain a subset of top-ranked gRNAs with the highest target gene editing frequency scores;(ii) identifying a desired allele among gene editing alleles produced in the plant cell by the top-ranked gRNA; and(iii) selecting a gRNA which is predicted to provide the desired allele.

47. The method of claim 46, wherein the number of top-ranked gRNAs is at least five.

48. The method of claim 46, wherein the number of top-ranked gRNAs is at least ten.

49. The method of claim 40, wherein the predicting of the functional effect of theINDELS step further comprises selecting a subset of gRNAs based on their target gene editing frequency score.

50. The method of claim 49, wherein the predicting of the functional effect of the INDELS step further comprises: inputting a set of alleles produced by the subset of gRNAs into a protein language model; processing the set of alleles via the model; outputting a ratio from the model wherein the ratio can be used to evaluate the predicted functional effect of each allele of the set of alleles, andAgent Ref. No. P14786WOOO selecting one or more gRNAs from the subset of gRNAs based on the predicted functional effect of the alleles produced by the gRNA.

51. The method of claim 39 or 40, wherein the predicting of the functional effect of the INDELS step comprises using a model that utilizes machine learning and is trained in a selfsupervised manner.

52. The method of claim 51, wherein evolutionary' information is integrated into parameters of the model.

53. The system of claim 52, wherein the predicting of the functional effect of the INDELS step further comprises inputting first and second protein sequence variants into the model.

54. The method of claim 53, wherein the predicting of the functional effect of the INDELS step further comprises:(i) masking a single amino acid position of the first sequence variant;(ii) predicting a probability' distribution over all possible amino acids at the single amino acid position based on surrounding sequence context and the evolutionary information;(iii) determining a conditional probability of the single amino acid position based on the probability distribution;(iv) performing steps (i)-(iii) for each amino acid position of a length of the first sequence variant in a one-by-one manner; and(v) calculating a product of each conditional probability of each amino acid position of the length to produce an approximated joint probability^ of the first sequence variant.

55. The method of claim 54. wherein the predicting of the functional effect of the INDELS step further comprises:(i) masking a single amino acid position of the second sequence variant;Agent Ref. No. P14786WOOO(ii) predicting a probability distribution over all possible amino acids at the single amino acid position of the second sequence variant based on surrounding sequence context and the evolutionary information;(iii) determining a conditional probability of the single amino acid position of the second sequence variant based on the probability distribution;(iv) performing steps (i)-(iii) for each amino acid position of a length of the second sequence variant in a one-by-one manner; and(v) calculating a product of each conditional probability of each amino acid position of the length of the second sequence variant to produce an approximated joint probability of the second sequence variant.

56. The method of claim 55, wherein the predicting of the functional effect of the INDELS step further comprises calculating a quotient of the approximated joint probability of the first sequence variant divided by the approximated joint probability of the second sequence variant.

57. The method of claim 55, wherein the second sequence variant differs from the first sequence variant by an in-frame INDEL such that the length of the second sequence variant differs from the length of the first sequence variant, and wherein the predicting of the functional effect of the INDELS step further comprises calculating a quotient of a geometric mean of the approximated joint probability of the first sequence variant divided by a geometric mean of the approximated j oint probability of the second sequence variant.

58. The method of claim 53, wherein the predicting of the functional effect of the INDELS step further comprises:(i) partitioning the first sequence variant into one or more groups of amino acid positions;(ii) masking a group of amino acid positions of the one or more groups of amino acid positions of the first sequence variant;(iii) predicting a probability distribution associated with the group of amino acid positions based on surrounding sequence context and the evolutionary information;Agent Ref. No. P14786WOOO(iv) simultaneously determining a conditional probability of all amino acid positions of the group of amino acid positions based on the probability distribution;(v) performing steps (ii)-(iv) for each group of the one or more groups of amino acid positions of a length of the first sequence variant in a group-by-group manner; and(vi) calculating a product of each conditional probability of each group of the one or more groups of amino acid positions of the length to produce an approximated joint probability of the first sequence variant.

59. The system of claim 58, wherein the predicting of the functional effect of the INDELS step further comprises:(i) partitioning the second sequence variant into one or more groups of amino acid positions;(ii) masking a group of amino acid positions of the one or more groups of amino acid positions of the second sequence variant;(iii) predicting a probability distribution associated with the group of amino acid positions of the second sequence variant based on surrounding sequence context and the evolutionary information;(iv) simultaneously determining a conditional probability of all amino acid positions of the group of amino acid positions of the second sequence variant based on the probability distribution associated with the second sequence variant;(v) performing steps (ii)-(iv) for each group of the one or more groups of amino acid positions of a length of the second sequence variant in a group-by-group manner; and(vi) calculating a product of each conditional probability of each group of the one or more groups of amino acid positions of the length of the second sequence variant to produce an approximated joint probability of the second sequence variant.

60. The system of claim 59, wherein the predicting of the functional effect of the INDELS step further comprises calculating a quotient of the approximated joint probabilityAgent Ref. No. P14786WOOO of the first sequence variant divided by the approximated joint probability' of the second sequence variant.

61. The system of claim 59, wherein the second sequence variant differs from the first sequence variant by an in-frame INDEL such that the length of the second sequence variant differs from the length of the first sequence variant, and wherein the predicting of the functional effect of the INDELS step further comprises calculating a quotient of a geometric mean of the approximated joint probability of the first sequence variant divided by a geometric mean of the approximated j oint probability of the second sequence variant.

62. The method of claim 59, wherein each of the one or more groups of amino acid positions of the first sequence variant and each of the one or more groups of amino acid positions of the second sequence variant comprise random amino acid positions.

63. The method of claim 53, wherein the predicting of the functional effect of the INDELS step further comprises:(i) predicting a probability distribution over all possible amino acids at each amino acid position of the first sequence variant based on surrounding sequence context and the evolutionary information;(ii) determining a conditional probability of each amino acid position of the first sequence variant based on the probability distribution of each amino acid position; and(iii) calculating a product of each conditional probability of each amino acid position of the first sequence variant to produce an approximated joint probability of the first sequence variant.

64. The method of claim 63, wherein the predicting of the functional effect of the INDELS step further comprises:(i) predicting a probability distribution over all possible amino acids at each amino acid position of the second sequence variant based on surrounding sequence context and the evolutionary information;Agent Ref. No. P14786WOOO(ii) determining a conditional probability of each amino acid position of the second sequence variant based on the probability distribution of each amino acid position of the second sequence variant; and(iii) calculating a product of each conditional probability of each amino acid position of the second sequence variant to produce an approximated joint probability' of the second sequence variant; wherein the second sequence variant and the first sequence variant have the same length.

65. The method of claim 64, wherein the predicting of the functional effect of the INDELS step further comprises calculating a quotient of the approximated joint probability' of the first sequence variant divided by the approximated joint probability of the second sequence variant.

66. The method of claim 39 or 40, wherein the predicting of the functional effect of the INDELS step further comprises using a model that utilizes machine learning and is trained in a self-supervised manner.

67. The method of claim 66, wherein the predicting of the functional effect of the INDELS step further comprises inputting first and second protein sequence variants into the model.

68. The method of claim 67, wherein the predicting of the functional effect of the INDELS step further comprises using the model to compute an alignment-based dimensionless score of relative fitness regarding the first and second sequence variants.

69. The method of claim 39 or 40, w herein the predicted functional effect of the INDELS and / or the preferred guide RNA can be used to select and / or predict in-planta information.

70. The method of claim 69. wherein the in-planta information comprises phenotype information.Agent Ref. No. P14786WOOO71. The method of any one of claims 38, 39, or 40, further comprising introducing into an experimental plant, experimental plant part, or experimental plant cell lacking the desired INDELS:(a)(i) the selected preferred gRNA or a polynucleotide encoding the selected preferred gRNA; and(ii) a Cas nuclease which binds the selected preferred gRNA or a polynucleotide encoding the Cas nuclease; or(b) a ribonucleoprotein complex comprising the Cas nuclease bound to the selected preferred gRNA.

72. The method of claim 71 , further comprising obtaining the plant or plant part comprising the desired INDELS from the experimental plant, experimental plant part, or experimental plant cell lacking the desired INDELS.

73. The method of claim 71. wherein:(i) the plant, plant part, experimental plant, and / or experimental plant part is and / or comprises a seed, leaf, stem, flower, meristem, or root; or(ii) the plant cell and / or experimental plant cell comprises plant callus tissue, embryogenic callus tissue, or a protoplast.

74. The method of claim 39 or 40, further comprising isolating, amplifying, and / or sequencing genomic DNA sequences containing the target plant gene from an experimental plant cell that had received the distinct candidate guide RNA and the Cas nuclease encoding DNA vector, wherein the isolated, amplified, and sequenced genomic DNA sequences can be used in the step of predicting the target gene editing frequency score and / or in the step of predicting the functional effect of the INDELS.

75. System(s), method(s), apparatus(es). or kit(s) as herein described.

76. System(s), method(s), apparatus(es), or kit(s) incorporating novel and unobvious aspect(s) of the present disclosure as they are as substantially show n or described.Agent Ref. No. P14786WOOO77. Alternative system(s), method(s), apparatus(es), or kit(s) incorporating other novel and unobvious aspect(s) of the present disclosure as they are as substantially shown or described.

78. The system(s), method(s), apparatus(es), or kit(s) according to any one of the preceding claims further comprising improvement(s), structure(s), and / or step(s) known to one of ordinary skill in the art.

79. The system(s), method(s), apparatus(es), or kit(s) according to any one of the preceding claims wherein an element of the same possesses a functional capability as substantially shown or described and / or a functional capability known to one of ordinary skill in the art.

80. The system(s), method(s), apparatus(es), or kit(s) according to any one of the preceding claims within its applicable or intended environment.