Biomolecular fitness inference using machine learning for drug discovery by directed evolution

The molecular discovery system uses deep learning to infer biomolecular fitness from sequencing data, addressing challenges in target-binding candidate screening by accurately predicting on-target and off-target binding and enhancing the discovery of high-fit, low-frequency biomolecules.

JP2025536942APending Publication Date: 2025-11-12GENENTECH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025522550
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-20
Filing Date
2023-10-19
Publication Date
2025-11-12

AI Technical Summary

Technical Problem

Current multiplexed target-binding candidate screening systems face challenges in simultaneously selecting many nucleotide-containing peptide libraries for binding to desired targets due to issues like sample-to-sample variation and data complexity, making it difficult to identify high-fit, low-frequency biomolecules.

Method used

A molecular discovery system utilizing deep learning techniques infers biomolecular fitness based on sequencing time series data from directed evolution experiments, accounting for genetic drift and mutation processes, and disentangling on-target and off-target binding to rank biomolecules effectively.

Benefits of technology

The system accurately predicts on-target and off-target binding fitness, discovers novel and distinct genotypes, and improves the quantity and diversity of identified molecules by identifying low-frequency, high-fit molecules that conventional systems miss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025536942000001_ABST
    Figure 2025536942000001_ABST
Patent Text Reader

Abstract

In one embodiment, a method includes accessing a biomolecular representation of a first biomolecule and processing the biomolecular representation with a machine learning model trained using sequencing time series data. The sequencing time series data is obtained from directed evolution of a population of biomolecules over multiple enrichment rounds, where the population of biomolecules in each enrichment round is a set of unique biomolecules relative to each other enrichment round. The sequencing time series data in each enrichment round includes a biomolecule frequency of each biomolecule in the population of biomolecules in the respective enrichment round. The training includes learning an inferred fitness score for the population of biomolecules for each enrichment round by predicting the biomolecule frequency of the population of biomolecules in each enrichment round based on the biomolecule frequency of the population of biomolecules in a previous enrichment round. The method further includes outputting the inferred fitness score for the first biomolecule.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Priority This application claims the benefit under 35 U.S.C. § 119(e) of U.S. Provisional Patent Application No. 63 / 417,961, filed October 20, 2022, which is incorporated herein by reference.

[0002] Technical Field The present disclosure relates generally to systems and methods for inferring biomolecular fitness, and more particularly to machine learning techniques for biomolecule selection. [Background technology]

[0003] background Directed evolution through iterative mutation and human-designed selection is a powerful approach for drug discovery, such as large molecule drug discovery. Mutation is an important part of directed evolution. Directed evolution approaches for drug discovery use genetic strategies (e.g., DNA-encoded, RNA-encoded, or phage-based) to create very large yet specific libraries of molecules whose amplification is driven by the target of interest. In other words, directed evolution approaches can discover drug-like biomolecules, such as macrocycles, with novel activities of interest.

[0004] Current multiplexed target-binding candidate screening analysis systems have difficulty simultaneously selecting many nucleotide-containing peptide libraries for binding to desired targets due to problems such as sample-to-sample variation and data complexity. Therefore, there is a need for improved multiplexed target-binding candidate screening analysis systems and methods to facilitate the simultaneous selection of candidate binding agents for desired binding targets, such as proteins. Summary of the Invention

[0005] Overview of Certain Embodiments This paper provides a system and method for biomolecular fitness inference.The problem is how to select individual biomolecules with improved binding ability from directed evolution experiments.One solution is to infer the fitness of biomolecules based on sequencing time series data obtained from directed evolution experiments, and then use machine learning techniques to select biomolecules based on the inferred fitness.

[0006] In certain embodiments, one molecular discovery system described herein utilizes deep learning techniques for biomolecular fitness selection. By way of example and not limitation, the molecular discovery system may include a deep neural network that estimates molecular fitness, which represents at least the biological activity of a molecule, and uses fitness to rank biomolecules (e.g., other types of peptide therapeutics, such as macrocyclic peptides, small molecules, bicyclic peptides, or any molecules that can be associated with or tagged to DNA-encoded libraries or other similar technologies) for selection. The molecular discovery system may establish a fitness inference problem that takes into account on-target and off-target time-series DNA sequencing data. The molecular discovery system may utilize maximum likelihood solutions of nonlinear dynamical systems triggered by fitness-based competition. In contrast to previous studies that focused on two-round enrichment prediction, the disclosed approach can in principle learn from multiple time-series rounds. By ranking molecules based on fitness, the molecular discovery system may identify low-frequency, high-fit molecules that would otherwise be missed by conventional systems. The molecular discovery system can accurately predict on-target and off-target binding fitness, discover novel and distinct genotypes of biomolecules, and improve the quantity and diversity of identified molecules. Experiments show that jointly training a sequence-to-fitness deep learning model (e.g., Transformer) to infer fitness improves performance over a baseline model without deep learning and a two-round enrichment baseline.

[0007] In certain embodiments, the molecular discovery system may access a biomolecular representation of a first biomolecule. The molecular discovery system may then process the biomolecular representation of the first biomolecule using a machine learning model. The machine learning model was trained using sequencing time series data related to the biomolecular frequency of a particular biomolecule. The sequencing time series data was obtained from the directed evolution of a population of biomolecules over multiple enrichment rounds. The population of biomolecules in each enrichment round may be a set of biomolecules unique to each other enrichment round. As used herein, "unique" indicates that the population of biomolecules may be mutated in each enrichment round, such that the population of biomolecules in each enrichment round has at least some unique biomolecules relative to the population of biomolecules in each other enrichment round. The sequencing time series data for each enrichment round may include the biomolecular frequency of each biomolecule in the population of biomolecules in each enrichment round. The training may include learning an inferred fitness score for the population of biomolecules for each enrichment round by predicting biomolecule frequencies for the population of biomolecules in each enrichment round given biomolecule frequencies for the population of biomolecules in one or more prior enrichment rounds. The molecular discovery system may further output an inferred fitness score for the first biomolecule based on processing the biomolecular representation of the first biomolecule by the machine learning model.

[0008] Biomolecule fitness selection presents certain technical challenges. One technical challenge may include effectively establishing a fitness inference task for biomolecules that undergo mutation in enrichment rounds. A solution proposed by the embodiments disclosed herein to address this challenge may be to develop a model of evolutionary dynamics that optimizes the inferred fitness by taking into account biomolecule frequency, since such a model accounts for the genetic drift and mutation processes of biomolecules in enrichment rounds. Another technical challenge may include disentangling sequencing time-series data by both on-target binding strength and off-target binding. A solution proposed by the embodiments disclosed herein to address this challenge may be to consider off-target fitness, infer total fitness from standard directed evolution data, and then infer on-target fitness, since this approach may enable inferences regarding the relative contributions of on-target and off-target fitness. Another technical challenge may include the measured enrichment being unreliable for lower counts due to high assay noise. To address this challenge, the solution presented by the embodiments disclosed herein may be to use Dirichlet-Multinomial loss to optimize the fitness inference task, since Dirichlet-Multinomial loss may account for the increased difficulty of predicting biomolecule frequencies in the current round from previous rounds when the total read count is low.

[0009] Certain embodiments disclosed herein may provide one or more technical advantages. A technical advantage of an embodiment may include identifying the fitness of biomolecules not present in the original enrichment round because the molecular discovery system utilizes a deep learning model that can infer the fitness of unseen biomolecules. Another technical advantage of an embodiment may include more effective selection of discovered hits because the molecular discovery system infers the fitness of biomolecules that exhibit biological activity and then determines discovered hits based on such biological activity. Another technical advantage of an embodiment may include the ability to filter out low-specificity binders from the selection results because the molecular discovery system can determine both on-target and off-target binding. Another technical advantage of an embodiment may include the ability to discover novel and distinct genotypes because the molecular discovery system can identify genotypes with high fitness but low frequency (and vice versa) that may be overlooked through standard approaches for biomolecule discovery. Certain embodiments disclosed herein may not provide any, some, or all of the above technical advantages. One or more other technical advantages may be readily apparent to one skilled in the art upon consideration of the drawings, description, and claims of the present disclosure.

[0010] In certain embodiments, the techniques described herein include accessing, by one or more computing systems, a biomolecular representation of a first biomolecule; and processing the biomolecular representation of the first biomolecule with a machine learning model, wherein the machine learning model is trained using sequencing time series data related to the biomolecular frequency of the particular biomolecule, the sequencing time series data being obtained from directed evolution of a population of biomolecules over multiple enrichment rounds, the population of biomolecules in each enrichment round being a unique set of biomolecules relative to each other enrichment round, and the sequencing time series data related to each enrichment round being a unique set of biomolecules relative to each other enrichment round. The method includes processing a biomolecular representation of a first biomolecule, wherein the lineage data includes a biomolecule frequency for each biomolecule in the population of biomolecules in each enrichment round, and training includes learning an inferred fitness score for the population of biomolecules for each enrichment round by predicting a biomolecule frequency for the population of biomolecules in each enrichment round given the biomolecule frequencies of the population of biomolecules in one or more prior enrichment rounds; and outputting, with a machine learning model, an inferred fitness score for the first biomolecule based on the processing of the biomolecular representation of the first biomolecule.

[0011] In certain embodiments, the technology described herein relates to a method further comprising determining whether a biological activity associated with the first biomolecule meets predetermined criteria for selection based on the inferred fitness score for the first biomolecule.

[0012] In certain embodiments, the technology described herein relates to methods wherein the multiple enrichment rounds comprise at least three enrichment rounds, and at least one of the enrichment rounds is a control round in which the population of biomolecules is analyzed without the presence of the target protein.

[0013] In certain embodiments, the technology described herein relates to methods, wherein the inferred fitness score for a first biomolecule is indicative of the biological activity of the first biomolecule against a target protein.

[0014] In certain embodiments, the technology described herein relates to a method in which learning an inferential fitness score in training a machine learning model includes optimizing a Dirichlet multinomial loss function, wherein the Dirichlet multinomial loss function utilizes an overdispersed multinomial distribution to account for the increasing difficulty associated with predicting biomolecule frequencies of a population of biomolecules in each enrichment round given biomolecule frequencies of the population of biomolecules in prior enrichment rounds.

[0015] In certain embodiments, the technology described herein relates to a method, wherein training the machine learning model further comprises calculating a Dirichlet loss negative log-likelihood between the predicted biomolecule frequencies and the actual biomolecule frequencies as a negative log-likelihood.

[0016] In certain embodiments, the technology described herein relates to a method, wherein the inferred fitness score for a first biomolecule comprises an on-target fitness score associated with binding of the first biomolecule to a target protein.

[0017] In certain embodiments, the technology described herein relates to a method, wherein the inferred fitness score for a first biomolecule includes an off-target fitness score associated with binding of the first biomolecule to a test device instead of a target protein.

[0018] In certain embodiments, the technology described herein relates to a method in which the inferred fitness score for a first biomolecule includes an on-target fitness score associated with the first biomolecule binding to a target protein and an off-target fitness score associated with the first biomolecule binding to a test device instead of the target protein, the method further including determining the binding specificity of the first biomolecule based on a ratio of the on-target fitness score to the off-target fitness score.

[0019] In certain embodiments, the technology described herein relates to methods in which the machine learning model includes one or more neural networks.

[0020] In certain embodiments, the technology described herein relates to a method in which one or more neural networks include a first neural network trained to predict on-target fitness scores associated with a biomolecule and a second neural network trained to predict off-target fitness scores associated with the biomolecule.

[0021] In certain embodiments, the technology described herein relates to a method further comprising generating a biomolecular representation of a first biomolecule, wherein the first biomolecule is a polypeptide corresponding to a first genotype, the generating comprising determining a plurality of amino acids of the first biomolecule, applying a function for each amino acid of the plurality of amino acids to determine a feature representation for each amino acid, and generating a genotype representation corresponding to the first genotype based on the plurality of feature representations associated with the plurality of amino acids.

[0022] In certain embodiments, the technology described herein relates to methods in which the first biomolecule is within a population of biomolecules in multiple rounds of enrichment.

[0023] In certain embodiments, the technology described herein relates to methods in which the first biomolecule is not within the population of biomolecules in multiple rounds of enrichment.

[0024] In certain embodiments, the technology described herein relates to methods wherein the sequencing time series data comprises DNA sequencing time series data, and wherein the biomolecule frequencies of particular biomolecules indicate genotype frequencies.

[0025] In certain embodiments, the technology described herein relates to a method further including: processing a plurality of biomolecular representations associated with a plurality of respective second biomolecules through a machine learning model to determine a plurality of inferred fitness scores for each of the plurality of second biomolecules; and selecting one or more second biomolecules that meet predetermined criteria for selection based on the inferred fitness scores for the plurality of second biomolecules, wherein one or more of the selected second biomolecules each is associated with a low relative biomolecule frequency in a final round of the plurality of enrichment rounds.

[0026] In certain embodiments, the techniques described herein relate to methods that further include generating a genotype space based on biomolecule frequencies and inferred fitness scores for a plurality of second biomolecules, and selecting one or more second biomolecules by identifying one or more second biomolecules from one or more regions within the genotype space, wherein each of the one or more regions is associated with a particular biomolecule frequency range and a particular biomolecule fitness range.

[0027] In certain embodiments, the technology described herein relates to a method, wherein the training further comprises pre-training an off-target model, wherein the pre-training comprises identifying one or more off-target enrichment rounds from a plurality of enrichment rounds, and pre-training the off-target model based on sequencing time-series data in the one or more off-target enrichment rounds.

[0028] In certain embodiments, the technology described herein relates to a method, wherein the training further includes accessing sequencing time series data from one or more on-target enrichment rounds from a plurality of enrichment rounds, and generating an on-target model based on the accessed sequencing time series data from the one or more on-target enrichment rounds and an off-target model.

[0029] In certain embodiments, the technology described herein relates to methods in which the first biomolecule is a macrocycle.

[0030] In certain embodiments, the technology described herein relates to methods in which a population of biomolecules is amplified by polymerase chain reaction (PCR) in each of multiple enrichment rounds.

[0031] In certain embodiments, the technology described herein relates to a method further including: processing a plurality of biomolecular representations associated with a plurality of respective second biomolecules through a machine learning model to determine a plurality of inferred fitness scores for each of the plurality of second biomolecules; and selecting one or more distinct biomolecules from the plurality of second biomolecules based on the inferred fitness scores for the plurality of second biomolecules, wherein the one or more distinct biomolecules satisfy predetermined criteria for selection, and wherein one or more of the distinct biomolecules are each associated with a low relative biomolecule frequency in a final round of the plurality of enrichment rounds.

[0032] In certain embodiments, the technology described herein includes one or more non-transitory computer-readable storage media embodying software that, when executed, accesses a biomolecular representation of a first biomolecule and processes the biomolecular representation of the first biomolecule with a machine learning model, the machine learning model being trained using sequencing time series data related to the biomolecular frequency of the particular biomolecule, the sequencing time series data being obtained from directed evolution of a population of biomolecules over multiple enrichment rounds, the population of biomolecules in each enrichment round being a unique set of biomolecules relative to each other enrichment round, and each enrichment round the sequencing time series data in the enrichment round includes a biomolecule frequency for each biomolecule in the population of biomolecules in each enrichment round, the training includes learning an inferred fitness score for the population of biomolecules for each enrichment round by predicting a biomolecule frequency for the population of biomolecules in each enrichment round taking into account the biomolecule frequencies of the population of biomolecules in one or more prior enrichment rounds, and the machine learning model is operable to output an inferred fitness score for the first biomolecule based on processing a biomolecular representation of the first biomolecule.

[0033] In certain embodiments, the techniques described herein include one or more processors and a non-transitory memory coupled to the processor and including instructions executable by the processor, wherein the processor, when executing the instructions, accesses a biomolecular representation of a first biomolecule and processes the biomolecular representation of the first biomolecule with a machine learning model, the machine learning model being trained using sequencing time series data related to biomolecular frequencies of the particular biomolecule, the sequencing time series data being obtained from directed evolution of a population of biomolecules over multiple enrichment rounds, the population of biomolecules in each enrichment round being characterized by a unique characteristic of the biomolecules relative to each other enrichment round. wherein the sequencing time series data at each enrichment round includes a biomolecule frequency for each biomolecule in the population of biomolecules at the respective enrichment round, and training includes learning an inferred fitness score for the population of biomolecules for each enrichment round by predicting the biomolecule frequency for the population of biomolecules at the respective enrichment round taking into account the biomolecule frequencies of the population of biomolecules in one or more prior enrichment rounds, and the machine learning model is operable to output the inferred fitness score for the first biomolecule based on processing the biomolecular representation of the first biomolecule.

[0034] In certain embodiments, the techniques described herein include accessing, by one or more computing systems, a plurality of biomolecular representations of a plurality of respective biomolecules; and processing the plurality of biomolecular representations of the plurality of respective biomolecules with a machine learning model, wherein the machine learning model is trained using sequencing time series data related to biomolecular frequencies of particular biomolecules, the sequencing time series data being obtained from directed evolution of a population of biomolecules over multiple enrichment rounds, the population of biomolecules in each enrichment round being a unique set of biomolecules relative to each other enrichment round, the sequencing time series data in each enrichment round including the biomolecular frequency of each biomolecule in the population of biomolecules in each respective enrichment round, and the training is performed using sequencing time series data related to biomolecular frequencies of particular biomolecules over one or more prior enrichment rounds. and selecting one or more biomolecules that meet a predetermined criterion for selection based on the inferred fitness scores for the plurality of biomolecules, wherein one or more of the selected biomolecules are each associated with a low relative biomolecular frequency in a final round of the plurality of enrichment rounds.

[0035] In certain embodiments, the techniques described herein relate to methods further comprising generating a genotype space based on biomolecule frequencies and inferred fitness scores for the plurality of biomolecules, wherein selecting one or more biomolecules that meet predetermined criteria for selection comprises identifying one or more biomolecules from one or more regions within the genotype space, each of the one or more regions being associated with a particular biomolecule frequency range and a particular biomolecule fitness range.

[0036] The embodiments disclosed herein are merely examples, and the scope of the disclosure is not limited thereto. Particular embodiments may include all, some, or none of the components, elements, features, functions, operations, or steps of the embodiments disclosed herein. Embodiments of the present invention are disclosed in the accompanying claims, which are directed, inter alia, to methods, storage media, systems, and computer program products. Any feature recited in one claim category, e.g., a method, may also be claimed in another claim category, e.g., a system. Dependencies or references in the accompanying claims are selected for formality reasons only. However, just as any combination of a claim and its features is disclosed and may be claimed without regard to the dependencies selected in the accompanying claims, any subject matter resulting from an intentional reference to any preceding claim (e.g., multiple dependencies) may likewise be claimed. Subject matter that may be claimed includes not only combinations of features set forth in the accompanying claims, but also any other combinations of features in the claims, and each feature recited in a claim may be combined with any other feature or combination of features in the claims. Furthermore, any of the embodiments and features described or shown in this specification may be claimed in a separate claim and / or in any combination with any of the embodiments or features described or shown in this specification or with any of the features of the appended claims. [Brief explanation of the drawings]

[0037] [Figure 1] 1 depicts a system diagram illustrating an example of a molecular discovery system, according to some exemplary embodiments.

[0038] [Figure 2A] 1 illustrates an exemplary deep learning fitness model with an off-target oracle.

[0039] [Figure 2B] 1 illustrates another exemplary deep learning fitness model.

[0040] [Figure 3] An exemplary distribution of enrichment in each pair of rounds is shown.

[0041] [Figure 4] 1 shows exemplary on-target versus off-target fitness for predictions on 100,000 holdout genotypes.

[0042] [Figure 5A] 1 shows an exemplary genotype space generated based on final round frequencies and final round frequency winners.

[0043] [Figure 5B] 1 shows an exemplary genotype space generated based on fitness and fitness winners.

[0044] [Figure 5C] 1 shows an exemplary final round frequency versus fitness.

[0045] [Figure 6] 1 illustrates an exemplary method for biomolecular fitness inference.

[0046] [Figure 7] 1 illustrates an exemplary computer system. DETAILED DESCRIPTION OF THE INVENTION

[0047] Description of exemplary embodiments Introduction Macrocycles are a promising class of drug candidates intermediate between small and large molecules. One protocol uses DNA-encoded libraries (DELs) to discover macrocycles with very large library sizes up to 10 trillion. This protocol allows for the coupling of any DNA codon with any natural or unnatural amino acid (aa) via a codon table constructed by the scientist. This known 1:1 mapping allows for next-generation sequencing (NGS) readout of the amino acids in macrocycle peptides. To discover hits, the library undergoes multiple rounds of selection, involving repeated steps of incubation with the target protein, washing away non-binders, amplification, and retranslation of the DNA sequence into macrocycle peptides. Amplification can have an intentionally high error rate to introduce mutations during intermediate selection steps. This protocol can be intensive, and its complexity can complicate probability models of observations. The molecules undergoing selection in the aforementioned protocol can be peptides, p-linkers, and DNA.

[0048] Furthermore, it may be of particular interest to identify shorter macrocycle hits (<9 aa) because they are more likely to be cell-permeable but bind to the target protein much weaker (specifically, with weaker binding affinity) than longer macrocycles (10–14 aa). Short macrocycles tend to have less diversity. While conventional methods may have reasonable potency, drug-likeness and cell-permeability may need to be subsequently optimized, which can be challenging. As a result, it may be advantageous to first find hits that are cell-permeable and drug-like, and then optimize potency. Conventional methods can only perform selection on short macrocycles, but selection can be overly harsh because hits do not have meaningful enrichment. Instead, noise may dominate, resulting in a poor signal-to-noise ratio. One particular type of "noise" may be nonspecific binding of the peptide-DNA linker to the target protein. For example, linker enrichment may be about 0.1%, which is acceptable for large macrocycles with peak enrichment >1% (10:1 signal-to-noise ratio). However, short macrocycles may have peak enrichment of about 0.1% (1:1 SNR), meaning that the amplification capacity of each short macrocycle may be "obscured" by the linker—weak binders still amplify because their linkers are bound to the target, and therefore strong binders are not distinguished. Amplification capacity indicates the ability to survive the next round of the selection process, which corresponds to the challenge of binding to the target protein and surviving physical washing, which can kinetically remove weak binders.

[0049] However, as discussed herein, laboratory directed evolution can be augmented with machine learning techniques to improve the activity and diversity of discovered hits, such as specific macrocycle genotypes of interest, typically with expanded binding sites, allowing for increased binding affinity and selectivity. Discovered hits may exhibit improved performance, such as improved ability to bind to antigens, such as viral antigens, tumor antigens, etc. In many directed evolution experiments, the primary outcome may include DNA sequencing data measuring the frequency of competing biomolecules (e.g., macrocycles) in the evolving biomolecule population. However, biomolecule frequency may not be an accurate indicator of biological activity. To discover hits, one may wish to sort biomolecules by their biological activity. This raises an inference problem: inferring the biological activity of biomolecules given their frequency.

[0050] Molecular discovery systems can be used to perform directed evolution. In a DNA-encoded library (DEL) setting, the molecular discovery system can translate DNA sequences into peptides or small proteins using a known codon table. In one embodiment, the alphabet size of the DNA sequences can be 4 and the alphabet size of the small proteins can be 20. The peptides are physically connected to the DNA sequences by linkers.

[0051] At each time step, round, or generation of directed evolution, the molecular discovery system may perform activity-based selection. The multiple enrichment rounds may include at least three enrichment rounds, at least one of which may be a control round that analyzes the population of biomolecules without the presence of the target protein. 13 -10 14A set of more than 10 peptide-linker-DNA compounds is challenged to bind to a target protein immobilized on a solid support. The molecular discovery system 100 may perform a washing step to remove peptide-linker-DNA compounds with weak binding strength and extract survivors. In a single round of selection, candidates may be challenged to bind to the target, and then a washing step may be performed that may remove some binders (the washing may be intended to remove primarily very weak binders or non-binders).

[0052] At each time step, round, or generation of directed evolution, there can be mutations and amplifications. The DNA sequence of the binding compound, and therefore the peptide identity, can be determined by DNA sequencing. The peptide linker-DNA compound can be amplified by polymerase chain reaction (PCR), a multi-step process that doubles the DNA molecule at each step by replication. In other words, a population of biomolecules can be amplified by polymerase chain reaction (PCR) at each of multiple enrichment rounds. PCR can be performed at a rate of 1 x 10 per nucleotide per step. -5 Mutations can be introduced.

[0053] After directed evolution is applied to a population of biomolecules, the molecular discovery system described herein trains a deep learning model for biomolecular fitness selection. By way of example and not limitation, the molecular discovery system may train a deep neural network that estimates molecular fitness, which represents at least the biological activity of a molecule, and uses fitness to rank biomolecules (e.g., macrocycles) for selection. The molecular discovery system may establish a fitness inference problem that considers on-target and off-target time-series DNA sequencing data. The molecular discovery system may utilize maximum likelihood solutions of nonlinear dynamical systems induced by fitness-based competition. In contrast to previous studies that focused on two-round enrichment prediction, the disclosed approach can in principle learn from multiple time-series rounds. By ranking molecules based on fitness, the molecular discovery system can identify low-frequency, high-fit molecules that would otherwise be missed by conventional systems. The molecular discovery system may accurately predict on-target and off-target binding fitness, discover novel and distinct genotypes of biomolecules, and improve the quantity and diversity of identified molecules. Experiments show that inferring fitness while jointly training a sequence-to-fitness deep learning model (e.g., transformer) improves performance over a baseline model without deep learning and a two-round enrichment baseline.

[0054] In certain embodiments, the molecular discovery system may access a biomolecular representation of a first biomolecule. The molecular discovery system may then process the biomolecular representation of the first biomolecule using a machine learning model. The machine learning model was trained using sequencing time series data related to the biomolecular frequency of a particular biomolecule. The sequencing time series data was obtained from the directed evolution of a population of biomolecules over multiple enrichment rounds. The population of biomolecules in each enrichment round may be a set of biomolecules unique to each other enrichment round. As used herein, "unique" indicates that the population of biomolecules may be mutated in each enrichment round, such that the population of biomolecules in each enrichment round has at least some unique biomolecules relative to the population of biomolecules in each other enrichment round. The sequencing time series data for each enrichment round may include the biomolecular frequency of each biomolecule in the population of biomolecules in each enrichment round. The training may include learning an inferred fitness score for the population of biomolecules for each enrichment round by predicting biomolecule frequencies for the population of biomolecules in each enrichment round given biomolecule frequencies for the population of biomolecules in one or more prior enrichment rounds. The molecular discovery system may further output an inferred fitness score for the first biomolecule based on processing the biomolecular representation of the first biomolecule by the machine learning model.

[0055] Automated system for target binding candidate analysis FIG. 1 shows a system diagram illustrating an example of a molecular discovery system 100 according to some demonstrative embodiments. Referring to FIG. 1 , the molecular discovery system 100 may include a molecular discovery engine 110, an analysis engine 120, a client device 130, a data store 145, a machine learning model 150, and an automation facility 160. As shown in FIG. 1 , the molecular discovery engine 110, the analysis engine 120, the data store 145, the client device 130, and the automation facility 160 may be communicatively coupled via a network 140. The network 140 may be a wired and / or wireless network, including, for example, a local area network (LAN), a virtual local area network (VLAN), a wide area network (WAN), a public land mobile network (PLMN), the Internet, etc. The client device 130 may be a processor-based device, including, for example, a workstation, a desktop computer, a laptop computer, a smartphone, a tablet computer, a wearable device, etc. The data store 145 may be a database, including, for example, a relational database, a graph database, an in-memory database, a non-SQL (NoSQL) database, etc. In some exemplary embodiments, the molecular discovery engine 110 and / or the analysis engine 120 may be configured to support deep learning-based biomolecule fitness selection. In certain aspects, the biomolecule may be a macrocycle.

[0056] In some exemplary embodiments, data store 145 may store data such as a DNA-encoded library including multiple DNA sequences that may be stored as peptides or small proteins. Molecular discovery engine 110 and / or analysis engine 120 may execute analysis workflows that include applying various computational analyses based at least in part on the data in data store 145. Analysis engine 120 may train one or more machine learning models, such as machine learning model 150, based at least in part on the data in data store 145, for downstream analysis tasks, etc. For example, results of a workflow may be used as training data for training a neural network (or another type of machine learning model 150).

[0057] Referring again to FIG. 1 , molecular discovery engine 110 and / or analysis engine 120 may execute an analysis workflow based on one or more user inputs received from client device 130. For example, as shown in FIG. 1 , analysis engine 120 may generate user interface 135 for display on client device 130. One or more user inputs that may be received via user interface 135 may specify one or more subsets of data contained in data store 145. One or more visual representations of at least some of the results of the analyses performed by molecular discovery engine 110 and / or analysis engine 120 may be displayed as part of user interface 135. User interface 135 may be interactive such that the type of visual representation and the content presented therein may be updated in response to one or more user inputs.

[0058] In certain embodiments, the molecular discovery system 100 may utilize an automated analysis workflow to rapidly discover potent biomolecules. The automated analysis workflow may be executed by an automation facility 160. The automated analysis workflow may enable one round of selection for target binding and simultaneous visual observation and comparison of multiple libraries for inter-round comparison. The automated analysis workflow may include receiving quantitative information for multiple libraries of DNA-containing compositions. Receiving the quantitative information may include collecting quantitative data for input molecules, positive molecules, and negative molecules. The quantitative data may be generated after a series of automated library generation, target binding selection, and DNA measurement by a quantitative method. The automated analysis workflow may then transfer the quantification data and store the data in a database. Data transfer may include data scraping, in which a computer program extracts data from human-readable output coming from another program, such as Excel. The database stores each automatically exported dataset and all associated metadata (e.g., date, time, plate barcode, etc.). The automated analysis workflow may assign the dataset to a corresponding selection round. According to some embodiments, each exported dataset may be assigned to a corresponding round selected by the user. The automated analysis workflow can further visualize the data on a graphical display surface. Once each dataset is assigned to a corresponding round, all heat maps and charts are generated according to pre-set criteria or filter selections.

[0059] Further information regarding automated target binding analysis can be found in U.S. Patent Application No. 17 / 502022, filed October 14, 2021, particularly paragraphs 0065-0085, which is incorporated by reference in its entirety, among other discussions in that patent application.

[0060] Models of evolutionary dynamics A challenge exists in how to select distinct biomolecules with improved binding abilities from directed evolution experiments. One solution is to infer biomolecular fitness by parameterizing the logarithmic relative fitness value for each genotype based on sequencing time-series data obtained from the directed evolution experiment, and then select biomolecules based on the inferred fitness. Another solution is to infer biomolecular fitness based on sequencing time-series data obtained from the directed evolution experiment, and then utilize machine learning techniques to select biomolecules based on the inferred fitness. In certain embodiments, the molecular discovery system 100 may use a model of evolutionary dynamics to handle the task of fitness inference. By way of example and not limitation, the molecular discovery system 100 may include a deep neural network that estimates biomolecular fitness, which represents at least the biological activity of the biomolecule, and uses the fitness to rank biomolecules (e.g., macrocycles) for selection.

[0061] Let G be the number of unique genotypes in the population. At time or round t, let TIFF2025536942000002.tif5170, and total count as TIFF2025536942000003.tif6170. This disclosure uses repeated subscripts for indexing. TIFF2025536942000004.tif4170 is the count of the i-th genotype at time t. In population genetics, absolute fitness TIFF2025536942000005.tif5170 can be defined as follows: TIFF2025536942000006.tif7170In this disclosure, W describes the change in the number of genotypes due to differences in binding strength.

[0062] The molecular discovery system 100 infers absolute fitness from DNA sequencing time series data. The inferred fitness score for a first biomolecule may indicate the biological activity of the first biomolecule against a target protein. The inferred fitness score for a first biomolecule may include an on-target fitness score associated with binding of the first biomolecule to the target protein. t is shown as the count of sequencing reads, and TIFF2025536942000007.tif7170 is shown as the total read depth. The sequencing time series data may include DNA sequencing time series data, and the biomolecule frequency of a particular biomolecule may indicate the genotype frequency. DNA sequencing is t ~polynomial(n t / N t ,M t ), so that the genotype frequency, p t =n t / N t Only n can be estimated. t Without , we cannot use equation (1) to infer W. Equation (1) can be rewritten as follows (based on the relative fitness proposition): TIFF2025536942000008.tif9170

[0063] In one embodiment, the proof of the proposal regarding relative fitness is as follows.

[0064] Using TIFF2025536942000009.tif9170, we get: TIFF2025536942000010.tif104170

[0065] Equation (2) is p t ,p t+1 Given , we show that W can be identified only up to an unknown constant of proportionality. Specifically, if for any positive c, equation (2) holds for some W, then equation (2) can also hold for the absolute fitness value cW (based on the proposition about identifiability up to proportionality).

[0066] In one embodiment, the proposition regarding identifiability up to proportionality is as follows:

[0067] Assume TIFF2025536942000011.tif9170. Next, for any positive scalar c, TIFF2025536942000012.tif10170

[0068] The proof of the proposition about identifiability up to proportionality is as follows. TIFF2025536942000013.tif40170

[0069] Due to this inidentifiability, the inference problem can be formulated as follows: for t=0,1,...,T, m t Given, we infer the relative fitness w = cW, where c > 0 is unknown.

[0070] Based on equation (2), the molecular discovery system 100 can simulate a nonlinear dynamical system in the future in time, given an initial p0 and relative fitness W. This simple model may ignore genetic drift and lack the mutation process. Ignoring genetic drift reduces the likelihood of a 10 10 The population size may be reasonable, and the low mutation rate of PCR may allow for ignoring mutation effects for fitness inference purposes. Rewriting equation (2) using W in matrix notation, the present disclosure provides a priori round frequency p t Given , we define the fitness model prediction at round t+1 as follows: TIFF2025536942000014.tif7170

[0071] where: TIFF2025536942000015.tif4170 shows point multiplication. Based on Equation (3), the molecule discovery system 100 calculates the noise distribution according to different noise models, t. TIFF2025536942000016.tif5170 is exactly Optimizing W to reflect the frequency of biomolecules can be used to estimate the fitness of biomolecules. Developing a model of evolutionary dynamics that optimizes inferred fitness by taking into account the frequency of biomolecules can be an effective solution to address the technical challenge of effectively establishing the fitness inference task for biomolecules that undergo mutation in enrichment rounds, since such a model takes into account the genetic drift and mutation processes of biomolecules in enrichment rounds. Fitness can have the following properties: Fitness values ​​cannot be negative. Furthermore, under strict mathematical assumptions, W (i.e., the multiplication of fitness and the proportion of genotypes) can only increase over time under the model. The proportion of genotypes can increase if their fitness is above average and decrease if their fitness is below average. Increasing selection stringency can have the effect of increasing the variance or range of fitness. Increasing washing stringency or selection intensity can help better distinguish between weak and strong binders. However, if you wash too much, you risk washing everything away (killing everything).

[0072] In practice, the time series data may be affected not only by on-target binding strength, but also by off-target binding, where the molecule binds to the instrument instead of the target. In this case, the inferred fitness score for the first biomolecule may include an off-target fitness score associated with the first biomolecule binding to the test instrument instead of the target protein. The molecular discovery system 100 uses Equation (3) to derive the off-target data: the off-target fitness score from the directed evolution point where the target is not present and only the instrument is present. TIFF2025536942000018.tif7170. Furthermore, the molecular discovery system 100 can infer from standard directed evolution data: TIFF2025536942000019.tif7170, where both the target and the instrument are present. total =W on +W off Assuming that W onInferring total fitness from standard directed evolution data while accounting for off-target fitness and then inferring on-target fitness may be an effective solution to address the technical challenge of handling the impact of both on-target binding strength and off-target binding on sequencing time-series data, as this approach may enable inferences regarding the relative contributions of on-target and off-target fitness.

[0073] The challenge may involve matching the unknown proportionality constants in these two fitness inference problems, and in general, the inferred W off scales, in an unknown way, with W total This may be shown as follows: TIFF2025536942000020.tif12170

[0074] In one embodiment, the present disclosure discloses the following proposition for equation (4) regarding the upper limit of relative off-target activity: Assume the file is TIFF2025536942000021.tif12170. As a result, TIFF2025536942000022.tif9170.

[0075] The proof of the above proposition is as follows.

[0076] Using the property that absolute fitness is non-negative, TIFF2025536942000023.tif30170

[0077] This relationship holds for each entry in the vector.

[0078] As can be seen, the present disclosure demonstrates an upper bound on c1 / c2, which allows inferences regarding the relative contributions of on-target and off-target fitness. TIFF2025536942000024.tif8170

[0079] Alternatively, the molecular discovery system 100 may find W by fitting to the data. on W off A scale for σ may be learned. The present disclosure uses this approach in experiments. While the present disclosure describes a model of evolutionary dynamics that includes specific covariates, the present disclosure contemplates that the model can be readily extended to include additional covariates, if present.

[0080] Many conventional approaches learn enrichment scores that may be equivalent to relative fitness at two time points, but not at more than two time points (i.e., enrichment suggestions are not equivalent to relative fitness at time points ≠ p ≠ p ). This proposition may be stated as follows: If some fitness W holds such that p ≠ p ≠ p , then Assume TIFF2025536942000025.tif9170. And enrichment p2 / p1 ≠ p1 / p0. The proof is as follows.

[0081] Rearranging the equation for relative fitness, each enrichment has a numerator of W and a denominator of Equal to the fraction of TIFF2025536942000026.tif7170. The numerators are the same for p2 / p1 and p1 / p0, but the denominators are not. Therefore, the enrichment is not equal.

[0082] Enrichment is equivalent to the relative fitness at the two time points. Specifically, an algebraic rewrite of equation (2) shows that enrichment is proportional to absolute fitness at the two time points. Relative fitness is also proportional to absolute fitness.

[0083] If there are more than two time points where each pair of time points matches the same W, computational enrichment may result in different enrichment values ​​for each pair of time points. Thus, enrichment cannot equate to relative fitness.

[0084] In contrast to previous studies, the embodiments disclosed herein present a principled method for datasets with more than two time points. For hit prioritization, biologists may rank compounds by frequency at the final time point. However, the motto is not "survival of the most frequent" but "survival of the fittest." In other words, ranking by fitness may identify low-frequency, high-fitness "rising stars" that would otherwise be missed by ranking by frequency alone.

[0085] methodology Once a model of evolutionary dynamics is established, the molecular discovery system 100 can perform fitness inference on biomolecules for biomolecule selection. Fitness inference can be split into two components. One component may include a differentiable parameterization of W to map from genotype to fitness. Another component may include the predicted frequency from Equation (3). TIFF2025536942000027.tif5170 is the observed frequency p t+1 It may also include a differentiable loss function that is used to optimize how well the ensemble matches the ensemble. Given these two components, fitness may be inferred using first-order optimization.

[0086] In particular embodiments, fitness inference may be based on different loss functions. One example of a loss function may be the multinomial negative log-likelihood loss. A simple approach to learning fitness may be to take the predicted frequencies from equation (3) and treat them as event probabilities in a multinomial distribution. Then, the log-likelihood of a round of read counts given fitness may be: TIFF2025536942000028.tif5170 where, TIFF2025536942000029.tif7170 uses the event probability p and the event count C This is the multinomial likelihood of TIFF2025536942000030.tif4170.

[0087] Another exemplary loss function may be a Dirichlet multinomial negative log-likelihood loss. In certain embodiments, learning an inferential fitness score in training a machine learning model may include optimizing a Dirichlet multinomial loss function. The Dirichlet multinomial loss function utilizes an overdispersed multinomial distribution to account for the increasing difficulty associated with predicting biomolecule frequencies of a population of biomolecules in each enrichment round, given the biomolecule frequencies of the population of biomolecules in previous enrichment rounds. Training the machine learning model may further include calculating a Dirichlet loss negative log-likelihood between the predicted biomolecule frequencies and the actual biomolecule frequencies as a negative log-likelihood.

[0088] By construction, the relative fitness, i.e., equation (2), may take into account frequency rather than read count at round t. However, in practice, inherent high assay noise may reduce the reliability of measured enrichment at lower counts. The aforementioned multinomial distribution may not take into account the prior round count, preventing modeling of this aspect. Therefore, the present disclosure proposes that p t From p t+1 To account for the increased difficulty (i.e., lower confidence) of predicting β, we introduce the Dirichlet-multinomial (DM), an overdispersed multinomial distribution. Optimizing fitness inference tasks using the Dirichlet-multinomial loss can be an effective solution to address the technical challenge of measured enrichment being less reliable for lower counts due to high assay noise, because the Dirichlet-multinomial loss can account for the increased difficulty of predicting biomolecule frequencies in the current round from previous rounds when the total read count is low.

[0089] The DM distribution is the posterior predictive distribution of the Dirichlet prior on a multinomial model for count data. If we update the Dirichlet prior with cardinality α using the observed counts for each category k, the posterior is a Dirichlet distributed with cardinality α+k. The total cardinality is From TIFF2025536942000031.tif5170 Based on this observation, the present disclosure proposes a method for determining the DM likelihood. Considering TIFF2025536942000033.tif5170, in this case, TIFF2025536942000034.tif5170 is calculated as in equation (3), TIFF2025536942000035.tif5170 is the total read count. Note that the expected frequency of this distribution (like a multinomial distribution) is TIFF2025536942000036.tif5170, but the total density is Since the image is TIFF2025536942000037.tif5170, we consider the total count. For TIFF2025536942000038.tif4170, we obtain the following log-likelihood: TIFF2025536942000039.tif5170 where, TIFF2025536942000040.tif7170 is a count with density parameter α and total count C The Dirichlet-Multinomial Likelihood of TIFF2025536942000041.tif4170. The DM loss can be computed differentiably using Pyro [7] as the negative DM log-likelihood.

[0090] The molecular discovery system 100 may utilize various approaches to parameterizing fitness. One approach may include defining one log-relative fitness value per genotype, referred to in this disclosure as per-genotype fitness. Another approach may include using a neural network to map from genotype to log-relative fitness, referred to in this disclosure as a deep learning fitness model.

[0091] In e, one logw parameter may be trained for each genotype. The molecular discovery system 100 may randomly initialize the log-fitness values ​​using a standard normal distribution. Genotype-specific fitness may have the advantage of training in seconds on a GPU and having relatively few hyperparameters to tune. It may have the disadvantage of not being able to make predictions for genotypes it was not trained on and not taking advantage of the fact that similar genotypes are likely to have similar fitnesses. By way of example and not limitation, the L-BFGS algorithm has been found to be effective for training genotype-specific fitnesses. The molecular discovery system 100 may select a learning rate for each loss function using a grid search.

[0092] In a deep learning fitness model, a neural network model may parameterize the genotype-to-logw mapping. In other words, the machine learning model may comprise one or more neural networks. The model may learn common genotype motifs described in a latent space, and these genotype motifs are associated with high / low fitness values. Thus, the model may exploit the induced bias that enrichment (generally, fitness) is a function of genotype. As a result, this parameterization may enable extrapolation of fitness to new genotypes not observed in directed evolution experiments. This may enable further virtual screening, discovery of high-fitness motifs, and in silico optimization. As a result, the embodiments disclosed herein may have the technical advantage of identifying the fitness of biomolecules not present in the original enrichment round because the molecular discovery system 100 utilizes a deep learning model that can infer the fitness of unseen biomolecules.

[0093] In certain embodiments, the molecular discovery system 100 may generate a biomolecular representation of a first biomolecule, which may be a polypeptide corresponding to a first genotype. The generating may include determining a plurality of amino acids of the first biomolecule, applying a function to determine a respective amino acid feature representation for each amino acid of the plurality of amino acids, and generating a genotype representation corresponding to the first genotype based on the plurality of feature representations associated with the plurality of amino acids. To represent the genotype, an amino acid (a j Instead of one-hot encoding the chemical descriptor MF(a j ) through each a j The structure of the MF may be embedded in a hashed Morgan fingerprint function with substructure counts. By doing this, (1) the model can learn generalizable amino acid features, and (2) the system can j Inject an induction bias into the space. MF is genotype g i When applied to, token x i =[MF(a i ,1),MF(a i ,2)],...,MF(a i , L )] can be obtained.

[0094] In one exemplary embodiment, the molecular discovery system 100 used a transformant-like encoder with multi-head attention [9] to map from the embedded amino acid sequence to logw. The transformant-like encoder may have 8 multi-head attention, a 64-dimensional feedforward layer, and finally a single linear layer to predict fitness. The AdamW optimizer with grid search per loss function was used to train the model for 500 epochs on 25,600 samples, selecting the cosine annealing learning rate, batch size, and weight decay.

[0095] The molecular discovery system 100 may optimize hyperparameters of the fitness inference model using a grid search. In certain embodiments, the one or more neural networks may comprise a first neural network trained to predict on-target fitness scores associated with biomolecules and a second neural network trained to predict off-target fitness scores associated with biomolecules. In the case of on-target and off-target fitness models, the molecular discovery system 100 may first optimize the off-target model and then optimize the on-target model using the previously optimized frozen off-target model. By way of example and not limitation, the following grid of hyperparameters may be considered: The learning rate is [10 -4 , 10 -3 ,.., 10 -2 ]. The weight loss can be [0 , 10 -4 , 10 -3.5 , 10 -3 ,.., 10 -1 ]. The batch size can be [64 , 128 , 256]. The weights from each training run may be obtained using early stopping of the validation loss. In an exemplary embodiment, 10% of the training data may be randomly selected as validation data for model selection.

[0096] In one embodiment, to learn the unsupervised genotype space, the molecular discovery system 100 trained a VAE using the same encoder and feature set described for the deep learning fitness model. The fully connected decoder reconstructs both the peptide sequence and the embedding of each amino acid, thus learning amino acid similarity in addition to sequence similarity. Therefore, the final loss is the sum of KL divergence, sequence rearrangement, and amino acid rearrangement. The final representation is obtained from the latent space and shown after 2D projection with UMAP.

[0097] Since equation (3) may only allow prediction between two consecutive rounds, each mini-batch may only contain two consecutive rounds with the same subset of genotypes. To ensure that the loss is defined for each batch, the molecular discovery system 100 may use a simple batch sampling algorithm. Pseudocode for this algorithm is shown below. Def sample_batch( batch_size:int, genotypes:np.Ndarray,#G x genotype embedding dimension count_matrix:np.Ndarray,#G x T )-> tuple(np.Ndarray,np.Ndarray):#X,y T=count_matrix.Shape[1] t = np.Random.Randint(0,T-1) log_enrichment_defined_mask=( (count_matrix[:,t]>0)&(count_matrix[:,t+1]>0) ) log_enrichment_defined_idx=np.Where( log_enrichment_defined_mask, )[0] batch_idx=np.Random.Choice( log_enrichment_defined_idx, batch_size, ) X=genotypes[batch_idx] y=count_matrix[batch_idx,[t,t+1]] return X,y

[0098] 2A illustrates an exemplary deep learning fitness model with an off-target oracle. The on-target neural network (i.e., on-target model 210) and the off-target neural network (i.e., off-target oracle 220) can predict fitness from genotypes (represented by sequence representation 230), which can be used to predict future round genotype frequencies. As previously mentioned, the molecular discovery system 100 calculates the total fitness W total into on-target and one or more off-target contributions. Any combination of fitness parameterization and loss function can be used to calculate W using off-target rounds. on may be inferred, and W can be calculated using all target rounds via equation (4). off may be inferred. In other words, the inferred fitness score for the first biomolecule may include an on-target fitness score associated with the first biomolecule binding to the target protein and an off-target fitness score associated with the first biomolecule binding to the test device instead of the target protein. By inferring fitness contributions separately, the molecular discovery system 100 may select genotypes with better specificity. The molecular discovery system 100 may determine the binding specificity of the first biomolecule based on the ratio of the on-target fitness score to the off-target fitness score. As used herein, "binding specificity" refers to the binding pattern of the first biomolecule to the target protein and the test device. For example, low binding specificity indicates that the first biomolecule binds more strongly to the test device, while high binding specificity indicates that the first biomolecule binds more strongly to the target protein. As a result, the embodiments disclosed herein may have the technical advantage of being able to remove low-specificity binders from the selection results, since the molecular discovery system 100 can determine both on-target and off-target binding.

[0099] 2B shows another exemplary deep learning fitness model. A neural network 240 can predict fitness from genotypes (represented by sequence representation 230) without separating the total fitness into on-target and off-target contributions. The predicted fitness W can be used to predict genotype frequencies in future rounds.

[0100] Training the machine learning model may further include pre-training an off-target model. In practice, the molecular discovery system 100 may first pre-train an off-target fitness model using only off-target rounds. The pre-training may include identifying one or more off-target enrichment rounds from multiple enrichment rounds and pre-training an off-target model based on sequencing time-series data from one or more off-target enrichment rounds. This may provide an effective solution because (1) the molecular discovery system 100 can utilize off-target data measured for different targets, thus potentially increasing data size, and (2) the molecular discovery system 100 can efficiently reuse the pre-trained model for new targets. The pre-trained "off-target oracle" may then be frozen, and an on-target model may be trained using all target rounds via equation (4). In other words, training the machine learning model may further include accessing sequencing time series data from one or more on-target enrichment rounds from multiple enrichment rounds, and generating an on-target model based on the accessed sequencing time series data from the one or more on-target enrichment rounds and an off-target model. Embodiments disclosed herein have compared this approach to learning on-target and off-target fitnesses jointly and observed that the latter results in less stable training and significantly longer training times. [Example]

[0101] The presently disclosed subject matter provides improved fitness-based biomolecule selection. Examples of fitness inference and biomolecule selection are described below.

[0102] Example 1: Directed evolution data The directed evolution experiment analyzed below includes seven rounds (the first round is a random library) and three off-target rounds following rounds 5, 6, and 7. The directed evolution experiment selected for activity against the target, and the library included octamer (cyclic peptides). Embodiments disclosed herein divided the data into training and test sets by genotype, using holdout round 7 and off-target round 7 as holdout rounds. Embodiments disclosed herein then filtered the training and test sets so that each genotype had at least 10 counts in one on-target round and appeared in at least two consecutive rounds. The final number of training / test genotypes is summarized in Table 1. Figure 3 shows an exemplary distribution of enrichment for each pair of rounds. More specifically, Figure 3 shows a kernel density plot of the log enrichment ratios of filtered genotypes by pair of time points. [Table 1]

[0103] Example 2: Fitness inference comparison with baseline The embodiments disclosed herein compare the disclosed fitness inference approach with three baselines: the frequency of the prior round assumes that enrichment is equal to the frequency at the end of the training data, in other words, the fittest genotypes are the most frequent; the enrichment of the prior round assumes that enrichment in the holdout round is equal to enrichment in the penultimate round; enrichment regression introduces a regression loss function that can be used with any fitness parameterization; the loss is the mean squared error between the predicted log enrichment and the actual log enrichment for each pair of rounds; enrichment is calculated based on the first and prior rounds. It is defined as the ratio of genotype frequencies in TIFF2025536942000043.tif5170. This baseline may correspond to a naive extension of enrichment regression to data with multiple time points. It may not take into account the fact that enrichment depends not only on the genotype but also on the fitness of competing genotypes in the same selection round.

[0104] The embodiments disclosed herein evaluated model performance through two generalization tasks. First, the embodiments disclosed herein evaluated the ability of all models trained at time points 1, ..., T-1 to predict enrichment for a holdout final round T for a training sequence (Table 2, confirmed molecules). In this case, the first biomolecule is present in the population of biomolecules in the multiple enrichment rounds. Next, the embodiments disclosed herein evaluated the ability of deep learning fitness models to predict enrichment for round T for a holdout genotype (Table 2, unconfirmed molecules). In this case, the first biomolecule is not present in the population of biomolecules in the multiple enrichment rounds. In certain embodiments, the molecular discovery system 100 may determine whether the biological activity associated with the first biomolecule meets predetermined criteria for selection based on the inferred fitness score for the first biomolecule. Pearson-r is reported to measure the agreement between actual and predicted enrichment. Pearson-r weights highly enriched sequences most strongly, thus reflecting the goal of selecting high-fitness genotypes. When an off-target oracle is used, the individual contributions (on-target and off-target fitness) are aggregated for evaluation. [Table 2]

[0105] The results demonstrate the advantages of using a fitness inference framework compared to all baselines. Prior round frequency and enrichment baselines predict enrichment less well compared to the disclosed fitness-based method, demonstrating that enrichment is driven by the fittest considering all rounds, not just the last. Fitness inference with DM loss also outperformed enrichment regression when either parameterization was used, demonstrating the advantages of the stochastic loss built in the fitness framework. Between the two fitness-based losses, DM loss produced improved results over multinomial loss, which was difficult to parameterize for deep learning fitness models. This may indicate that considering read counts from prior rounds improves model performance. As can be seen, the embodiments disclosed herein may have the technical advantage of more effective selection of discovered hits, as the molecular discovery system 100 infers the fitness of biomolecules that exhibit biological activity and then determines discovered hits based on such biological activity.

[0106] Example 3: On-target and off-target fitness inference The embodiments disclosed herein utilize the individual fitness contributions W on and W off Figure 4 shows exemplary on-target versus off-target fitness for predictions on 100,000 holdout genotypes. Grayscale intensity corresponds to total fitness. As shown, the same total fitness corresponds to different on-target / off-target ratios. The dataset is characterized by high off-target binding, which translates to overall high off-target fitness. Therefore, the disclosed pipeline can be used to filter out low-specificity binders from the selection results. Upon retraining the model, embodiments disclosed herein observed high consistency in the inferred on-target / off-target fitness ratios (pairwise Spearman correlation = 0.66, 3 replicates).

[0107] Embodiments disclosed herein evaluate the usefulness of estimated fitness to highlight novel and distinct genotypes not readily detected through a standard final round frequency pipeline, where genotypes with high frequencies in the final selection rounds are deemed most fit ([12;13;14;15;16;17;18]). In certain embodiments, the molecular discovery system 100 may process a plurality of biomolecular representations associated with a plurality of respective biomolecules through a machine learning model to determine a plurality of inferred fitness scores for the plurality of biomolecules, respectively. The molecular discovery system 100 may further select one or more biomolecules that meet predetermined criteria for selection based on the inferred fitness scores for the plurality of biomolecules. One or more of the selected biomolecules may each be associated with a low relative biomolecular frequency in the final round of the multiple enrichment rounds.

[0108] Example 4: Discovery of novel and distinct genotypes In certain embodiments, the molecular discovery system 100 may select a set of distinct biomolecules as discovered hits based on a fitness score. The biomolecules in the set of distinct biomolecules may meet predetermined criteria for selection. The molecular discovery system 100 may calculate a distance metric between the biomolecules. By way of example and not limitation, the distance metric may be based on an edit distance. By way of another example and not limitation, the distance metric may be based on a distance determined based on a method for identifying distinct peptides. More information regarding identifying distinct peptides may be found in U.S. Patent Application No. 18 / 338,772, filed June 21, 2023, particularly paragraphs 0069-0120, which are incorporated by reference in their entirety, among other discussions in that patent application. By way of yet another example and not limitation, the distance metric may be based on a variational autoencoder (VAE). In one embodiment, training of the VAE may be performed using a reconstruction loss and monomer prediction. For further clarity, the VAE may be trained to reconstruct input biomolecules. The embedding layer in the bottleneck may generate embeddings for biomolecules by inserting the biomolecules into the trained VAE and reading out the vectors in the embedding layer in the bottleneck of the VAE. As a result, the distance between two biomolecules may be calculated based on the Euclidean distance between the two vectors for the two biomolecules thus generated. The molecular discovery system 100 may then cluster the biomolecules based on a distance metric. The molecular discovery system 100 may further select a distinct set of biomolecules based on both the fitness score and the cluster (e.g., selecting biomolecules with high fitness scores from distinct clusters). In other words, the predetermined criteria may be based on both the fitness score and the cluster. By way of example and not limitation, the molecular discovery system 100 may use an algorithm to select a distinct set of biomolecules. The algorithm may be a heuristic that combines high fitness scores and distinct clusters based on its evaluation. For example, the algorithm may select biomolecules from two or more distinct clusters.

[0109] The molecular discovery system 100 may further generate a genotype space based on the biomolecule frequencies and the inferred fitness scores of the multiple biomolecules. The molecular discovery system 100 may then select one or more biomolecules by identifying one or more biomolecules from one or more regions within the genotype space. Each of the one or more regions may be associated with a specific biomolecule frequency range and a specific biomolecule fitness range. Figure 5A shows an exemplary genotype space generated based on final round frequencies and final round frequency winners. Grayscale intensity corresponds to logarithmically scaled frequency and fitness, respectively. Figure 5B shows an exemplary genotype space generated based on fitness and fitness winners. Figure 5C shows an exemplary final round frequency versus fitness. The molecular discovery system 100 unsupervisedly learned via VAE to embed genotypes into a meaningful "genotype space" incorporating UMAP and compared high-frequency and high-fitness regions within this space (Figure 5A). As shown, while some regions share high frequency and high fitness values, regions of genotypes characterized by high fitness but low frequency (or vice versa) may be identified. Such regions may help improve the number and diversity of identified biomolecules. Focusing on the top 50 "winners" across the two criteria ( FIG. 5B ), it can be observed that the inferred fitness helps select genotypes that are distant (i.e., less similar) from those selected through frequency. Indeed, analyzing the agreement between final round frequency and fitness reveals that the estimated fitness spans a wide frequency range ( FIG. 5C ), especially for low frequency values. As can be seen, embodiments disclosed herein may have the technical advantage of the ability to discover novel and distinct genotypes, as the molecular discovery system 100 may identify genotypes with high fitness but low frequency (and vice versa) that may be overlooked through standard approaches for biomolecule discovery.

[0110] Example 5: Efficiency of hit discovery The embodiments disclosed herein can identify hits several months earlier than conventional methods. By inferring fitness, multiple time points can be used to identify hits earlier. For example, in one example case, each time point is separated by up to one month, which means that hits can potentially be identified more than two or three months earlier than the baseline.

[0111] Example 6: Identification of short macrocycles The embodiments disclosed herein can further identify short macrocycle hits without a dedicated short macrocycle library. In conventional protocols, dedicated short macrocycle libraries are often not implemented because they historically have not yielded good results. Fitness inference with sufficiently deep sequencing depth can enable the identification of short macrocycles (6-9 aa) in a general macrocycle library (6-14 aa), despite the dominance of long macrocycles (10-14 aa) in frequency.

[0112] Further noise reduction As previously mentioned, to find hits, biomolecules undergo multiple rounds of selection, including repeated steps of incubation with the target protein, washing away non-binders, amplification, and retranslation of the DNA sequence into macrocyclic peptides. The embodiments disclosed herein take the washing step into account when modeling the absolute fitness or number of offspring of each individual genotype. It may be assumed that each genotype has a probability of binding and then a probability of surviving washing. The binding probability may be exponentially distributed, with few hits having a high binding probability, and most genotypes having a probability near zero. Because binders are rare, washing can be thought of as a bottleneck, narrowing down a large number of candidate genotypes (e.g., 10 13 genotypes) to a much smaller number, e.g., 10 7 genotypes (representing a 1 in 1 million chance of binding, averaged across all genotypes). In the case of washing, it may be reasonable to assume that the probability of each genotype surviving washing is independent, which means that it follows a binomial distribution. Due to the properties of the binomial distribution, washing may not change the fitness estimate, but may add binomial sampling noise, especially when the copy number of the bound genotype is low. Because short macrocycles may be weak binders, the copy number of the bound genotype may be very low after binding. The washing step may then introduce binomial noise into the small number of copies of the bound genotype.

[0113] Binomial sampling may introduce noise into genotype frequencies over time as follows: During each selection step, aliquots may be taken or samples may be collected for DNA sequencing. By way of example and not limitation, frequency counts may be convoluted by ribosomal / translational bias and selection artifacts. The molecular discovery system 100 may utilize a probability model of the observations to calculate a p-value or Bayes factor of genotype activity to determine how much of the time series trajectory is due to random noise versus fitness-based selection. The strength of evidence may be assessed using a baseline model to generate a p-value, false discovery rate, or Bayes factor. For example, the probability of observing three sequences rather than just one with a timeline trajectory suggestive of high fitness may be lower under a noise model.

[0114] In one example, ribosome / translation bias can complicate fitness inference because sequences that ribosomes translate more efficiently may have more "children" per generation. In certain embodiments, the molecular discovery system 100 may address ribosome / translation bias by incorporating a simple model of ribosome translation efficiency into the deep learning fitness model. By way of example and not limitation, one method may include it as a logarithmic additive term. If the r of one genotype is twice the r value of another genotype, then let r be the ribosome translation efficiency of the genotype scaled so that the first genotype has twice the number of translation products. Then, translational control fitness f may be inferred according to the data, such as fr = w, where w is the effective fitness that governs population variation. A simple model of ribosome translation efficiency may be, for example, the average codon score in a sequence, because ribosome efficiency may have codon-dependent bias.

[0115] The deep learning fitness model can better distinguish signal from noise, thereby identifying short macrocycles with meaningful activity. The deep learning fitness model may utilize more information, i.e., aggregate across all time points and use a probabilistic model of observations and noise. The deep learning fitness model can be used to infer the fitness of each peptide by aggregating information across all time points under a probabilistic model. The deep learning fitness model can control ribosome translation efficiency. Furthermore, the molecular discovery system 100 can increase power by aggregating information across peptides within the same family.

[0116] In certain embodiments, the deep learning fitness model may further incorporate family clustering. With access to a pairwise amino acid similarity matrix defined using chemical atom-atom-pathway similarities and expert knowledge, the molecular discovery system 100 may utilize a similarity function between two sequences as the average pairwise similarity in those amino acids. The first step may be to cluster genotypes by similarity and report fitness statistics for the clusters: maximum, median, mean, number of genotypes in the cluster, cluster similarity, etc. Family clustering may strengthen the evidence for individual hits when there are highly related genotypes with loci that also support high fitness estimates. Similarity metrics may also be used to generate unseen sequences from seed input sequences. Experimenters may do this manually, but computational algorithms may find useful for easily designing libraries on a large scale. By way of example and not limitation, an algorithm may generate all mutants within a predetermined Hamming distance from the seed, discard those already observed in the data, rank the remainder by similarity, and return them.

[0117] In one embodiment, the molecular discovery system 100 may utilize a model for sampling macrocycles and interpreting families. Descending a hierarchical clustering tree from root to leaf may be considered a denoising procedure, starting with generality and gradually adding specificity. However, while a single hierarchical clustering tree is a tree, there may be multiple valid methods for clustering. Furthermore, while clustering may typically maximize sequence similarity between child nodes, in certain embodiments, clusters with slightly different sequences but higher fitness may be acceptable. For macrocycles, there may be multiple families to which a single macrocycle peptide can belong. Clustering may assign individuals to only a single family. In contrast, a model for sampling macrocycles and interpreting families may process relationships better described as a directed acyclic graph (DAG). By way of example and not limitation, the initial state of the model may be a maximally general macrocycle containing amino acids. At each state, a set of available actions can transform some abstract symbols into more concrete symbols, for example, any amino acid can be transformed into a symbol representing a set of hydrophobic, polar, or positively / negatively charged amino acids. The final state can be the observed macrocycle peptide sequence. In certain embodiments, a model for sampling macrocycles and interpreting families can be trained with r(x) = fitness, and once trained, new macrocycles can be sampled based on the fitness distribution, benefiting from the generalization ability and potential synthesizing potential of the model. To interpret the clusters, a stream of intermediate states (representing macrocycle families) can be analyzed and compared.

[0118] Consideration This disclosure introduces fitness inference for directed evolution time series data with on-target and off-target selection rounds. The embodiments disclosed herein developed and compared two parameterizations of fitness using two different probabilistic loss functions: per-genotype and deep learning fitness models. With both parameterizations, fitness inference using the DM loss significantly improved results compared to a baseline that did not consider the competition induced in multi-round selection experiments. Compared to the per-genotype parameterization, the deep learning fitness model enabled fitness predictions for novel, unseen genotypes that learned a general "motif-to-fitness" mapping under the assumption that similar genotypes yield similar fitness values.

[0119] The present disclosure demonstrates the ability of the fitness framework to accurately include off-target binding data. Compared to regression modeling, the disclosed fitness framework may not require learning explicit trade-off parameters for off-target binding. It is observed that the on-target and off-target strategies slightly negatively impact performance in terms of directly learning total fitness. This may be caused by the model requiring twice the number of parameters, and weight sharing and other regularization techniques may be explored to improve performance. However, the performance impact may be worth the ability to disentangle total fitness into its contributions for downstream filtering.

[0120] Finally, we evaluated the diversity of the most fit genotypes (estimated by a deep learning fitness model with DM loss) compared to the genotypes with the highest final round frequency. We found that fitness-based selection not only rediscovers genotypes similar to those selected by frequency, but also highlights genotypes in unexplored regions, thus improving overall diversity.

[0121] 6 illustrates an exemplary method 600 for biomolecular fitness inference. The method may begin at step 610, in which the molecular discovery system 100 may access a biomolecular representation of a first biomolecule, the first biomolecule being a macrocycle. In step 620, the molecular discovery system 100 may process the biomolecular representation of the first biomolecule with a machine learning model, the machine learning model comprising one or more neural networks, the machine learning model trained using sequencing time-series data related to the biomolecular frequency of the particular biomolecule, the sequencing time-series data obtained from directed evolution of a population of biomolecules over multiple enrichment rounds, the multiple enrichment rounds including at least three enrichment rounds, at least one of the enrichment rounds being a control round in which the population of biomolecules is analyzed without the presence of the target protein, the population of biomolecules being amplified by polymerase chain reaction (PCR) in each of the multiple enrichment rounds, the population of biomolecules in each enrichment round being a unique set of biomolecules relative to each other enrichment round, and the sequence of each enrichment round being a unique set of biomolecules. The sequence determination time series data includes a biomolecule frequency of each biomolecule in the population of biomolecules in each enrichment round, the sequencing time series data includes DNA sequencing time series data, and the biomolecule frequency of a particular biomolecule indicates a genotype frequency. Training includes learning an inferred fitness score for the population of biomolecules for each enrichment round by predicting the biomolecule frequency of the population of biomolecules in each enrichment round, taking into account the biomolecule frequency of the population of biomolecules in one or more prior enrichment rounds. Learning the inferred fitness score in training the machine learning model includes optimizing a Dirichlet multinomial loss function, which utilizes an overdispersed multinomial distribution to account for the increasing difficulty associated with predicting the biomolecule frequency of the population of biomolecules in each enrichment round, taking into account the biomolecule frequency of the population of biomolecules in prior enrichment rounds.In step 630, the molecular discovery system 100 may output an inferred fitness score for the first biomolecule based on processing the biomolecular representation of the first biomolecule by the machine learning model, the inferred fitness score for the first biomolecule indicating a biological activity of the first biomolecule with respect to the target protein, the inferred fitness score for the first biomolecule including one or more of an on-target fitness score associated with the first biomolecule binding to the target protein or an off-target fitness score associated with the first biomolecule binding to a test device instead of the target protein. In step 640, the molecular discovery system 100 may determine whether a biological activity associated with the first biomolecule meets predetermined criteria for selection based on the inferred fitness score for the first biomolecule. Certain embodiments may repeat one or more steps of the method of FIG. 6 as appropriate. Although the present disclosure describes and illustrates certain steps of the method of FIG. 6 as occurring in a particular order, the present disclosure contemplates that any suitable steps of the method of FIG. 6 may occur in any suitable order. Additionally, although this disclosure describes and illustrates exemplary methods for biomolecular fitness inference that include particular steps of the method of Figure 6, this disclosure contemplates any suitable method for biomolecular fitness inference that includes any suitable steps, which may include all, some, or none of the steps of the method of Figure 6, where appropriate. Additionally, although this disclosure describes and illustrates particular components, devices, or systems that perform particular steps of the method of Figure 6, this disclosure contemplates any suitable combination of any suitable components, devices, or systems that perform any suitable steps of the method of Figure 6.

[0122] Systems and methods 7 illustrates an exemplary computer system 700. In particular embodiments, one or more computer systems 700 perform one or more steps of one or more methods described or illustrated herein. In particular embodiments, one or more computer systems 700 provide functionality described or illustrated herein. In particular embodiments, software running on one or more computer systems 700 performs one or more steps of one or more methods described or illustrated herein or provides functionality described or illustrated herein. Particular embodiments include one or more portions of one or more computer systems 700. As used herein, references to a computer system may encompass computing devices, and vice versa, where appropriate. Furthermore, references to a computer system may encompass one or more computer systems, where appropriate.

[0123] The present disclosure contemplates any suitable number of computer systems 700. The present disclosure contemplates computer system 700 taking any suitable physical form. By way of example and not limitation, computer system 700 may be an embedded computer system, a system-on-chip (SOC), a single-board computer system (SBC) (e.g., a computer-on-module (COM) or system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, an interactive kiosk, a mainframe, a mesh of computer systems, a mobile phone, a personal digital assistant (PDA), a server, a tablet computer system, or a combination of two or more of these. Where appropriate, computer system 700 may include one or more computer systems 700 and may be unitary, distributed, spanning multiple locations, multiple machines, multiple data centers, or in a cloud that may include one or more cloud components in one or more networks. Where appropriate, one or more computer systems 700 may perform one or more steps of one or more methods described or illustrated herein without substantial spatial or temporal limitations. By way of example and not limitation, one or more computer systems 700 may perform one or more steps of one or more methods described or illustrated herein in real time or batch mode. One or more computer systems 700 may perform one or more steps of one or more methods described or illustrated herein at different times or in different locations, where appropriate.

[0124] In particular embodiments, computer system 700 includes a processor 702, memory 704, storage 706, an input / output (I / O) interface 708, a communication interface 710, and a bus 712. Although this disclosure describes and illustrates a particular computer system having a particular number of particular components in a particular arrangement, this disclosure contemplates any suitable computer system having any suitable number of any suitable components in any suitable arrangement.

[0125] In particular embodiments, processor 702 includes hardware for executing instructions, such as those comprising a computer program. By way of example and not limitation, to execute instructions, processor 702 may retrieve (or fetch) instructions from an internal register, an internal cache, memory 704, or storage device 706, decode and execute them, and then write one or more results to an internal register, an internal cache, memory 704, or storage device 706. In particular embodiments, processor 702 may include one or more internal caches for data, instructions, or addresses. This disclosure contemplates processor 702 including any suitable number of any suitable internal caches, where appropriate. By way of example and not limitation, processor 702 may include one or more instruction caches, one or more data caches, and one or more translation lookaside buffers (TLBs). Instructions in an instruction cache may be copies of instructions in memory 704 or storage device 706, and the instruction cache may speed up retrieval of those instructions by processor 702. Data in the data cache may be a copy of data in memory 704 or storage device 706 for the operation of an instruction executed in processor 702, the results of a previous instruction executed in processor 702 for access by a subsequent instruction executed in processor 702 or for writing to memory 704 or storage device 706, or other suitable data. The data cache may speed up read or write operations by processor 702. The TLB may speed up virtual address translation for processor 702. In particular embodiments, processor 702 may include one or more internal registers for data, instructions, or addresses. This disclosure contemplates processor 702 including any suitable number of any suitable internal registers, where appropriate. Where appropriate, processor 702 may include one or more arithmetic logic units (ALUs), may be a multi-core processor, or may include more than one processor 702. Although this disclosure describes and illustrates a particular processor, this disclosure contemplates any suitable processor.

[0126] In particular embodiments, memory 704 includes main memory for storing instructions executed by processor 702 or data for operation of processor 702. By way of example and not limitation, computer system 700 may load instructions into memory 704 from storage device 706 or another source (e.g., another computer system 700, etc.). Processor 702 may then load the instructions from memory 704 into an internal register or internal cache. To execute the instructions, processor 702 may retrieve the instructions from the internal register or internal cache and decode them. During or after execution of the instructions, processor 702 may write one or more results (which may be intermediate or final results) to an internal register or internal cache. Processor 702 may then write one or more of those results to memory 704. In particular embodiments, processor 702 executes only instructions stored in one or more internal registers or caches or in memory 704 (as opposed to storage device 706 or elsewhere) and operates only on data stored in one or more internal registers or caches or in memory 704 (as opposed to storage device 706 or elsewhere). One or more memory buses (which may each include an address bus and a data bus) may connect processor 702 to memory 704. Bus 712 may include one or more memory buses, as described below. In particular embodiments, one or more memory management units (MMUs) reside between processor 702 and memory 704 to facilitate accesses to memory 704 requested by processor 702. In particular embodiments, memory 704 includes random access memory (RAM). This RAM may be volatile memory, where appropriate. This RAM may be dynamic RAM (DRAM) or static RAM (SRAM), where appropriate. Moreover, this RAM may be single-ported RAM or multi-ported RAM, where appropriate. This disclosure contemplates any suitable RAM. Memory 704 may include, where appropriate, one or more memories 704. Although this disclosure describes and illustrates particular memory, this disclosure contemplates any suitable memory.

[0127] In particular embodiments, storage device 706 includes mass storage for data or instructions. By way of example and not limitation, storage device 706 may include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disk, a magneto-optical disk, magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Storage device 706 may include removable or non-removable (i.e., fixed) media, where appropriate. Storage device 706 may be internal or external to computer system 700, where appropriate. In particular embodiments, storage device 706 is non-volatile solid-state memory. In particular embodiments, storage device 706 includes read-only memory (ROM). Where appropriate, this ROM may be mask-programmed ROM, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), electrically alterable ROM (EAROM), or flash memory, or a combination of two or more of these. The present disclosure contemplates mass storage device 706 taking any suitable physical form. Storage 706 may include, where appropriate, one or more storage control units that facilitate communications between processor 702 and storage 706. Where appropriate, storage 706 may include one or more storage devices 706. Although this disclosure describes and illustrates particular storage devices, this disclosure contemplates any suitable storage device.

[0128] In particular embodiments, I / O interface 708 includes hardware, software, or both that provide one or more interfaces for communication between computer system 700 and one or more I / O devices. Computer system 700 may include one or more of these I / O devices, where appropriate. One or more of these I / O devices may enable communication between a person and computer system 700. By way of example and not limitation, the I / O devices may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, stylus, tablet, touch screen, trackball, video camera, other suitable I / O device, or a combination of two or more thereof. The I / O devices may include one or more sensors. This disclosure contemplates any suitable I / O devices and any suitable I / O interface 708 therefor. Where appropriate, I / O interface 708 may include one or more device or software drivers that enable processor 702 to drive one or more of these I / O devices. I / O interface 708 may include, where appropriate, one or more I / O interfaces 708. Although this disclosure describes and illustrates a particular I / O interface, this disclosure contemplates any suitable I / O interface.

[0129] In particular embodiments, communication interface 710 includes hardware, software, or both that provide one or more interfaces for communications (e.g., packet-based communications) between computer system 700 and one or more other computer systems 700 or one or more networks. By way of example and not limitation, communication interface 710 may include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wired-based network, or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network, such as a Wi-Fi network. This disclosure contemplates any suitable network and any suitable communication interface 710 for that network. By way of example and not limitation, computer system 700 may communicate with an ad-hoc network, a personal area network (PAN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), or one or more portions of the Internet, or a combination of two or more of these. One or more portions of one or more of these networks may be wired or wireless. By way of example, computer system 700 may communicate with a wireless PAN (WPAN) (e.g., a BLUETOOTH WPAN, etc.), a Wi-Fi network, a Wi-MAX network, a cellular network (e.g., a Global System for Mobile Communications (GSM) network), or other suitable wireless network, or a combination of two or more of these. Computer system 700 may include any suitable communication interface 710 for any of these networks, where appropriate. Communication interface 710 may include one or more communication interfaces 710, where appropriate. Although this disclosure describes and illustrates a particular communication interface, this disclosure contemplates any suitable communication interface.

[0130] In particular embodiments, bus 712 includes hardware, software, or both that interconnects components of computer system 700. By way of example and not limitation, bus 712 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI Express (PCIe) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or any other suitable bus, or a combination of two or more of these. Bus 712 may include one or more buses 712, where appropriate. Although this disclosure describes and illustrates a particular bus, this disclosure contemplates any suitable bus or interconnect.

[0131] As used herein, one or more computer-readable non-transitory storage media may comprise, where appropriate, one or more semiconductor-based or other integrated circuits (ICs) (such as, for example, field programmable gate arrays (FPGAs) or application-specific ICs (ASICs)), hard disk drives (HDDs), hybrid hard drives (HHDs), optical disks, optical disk drives (ODDs), magneto-optical disks, magneto-optical drives, floppy diskettes, floppy disk drives (FDDs), magnetic tapes, solid-state drives (SSDs), RAM drives, secure digital cards or drives, any other suitable computer-readable non-transitory storage media, or any suitable combination of two or more of these. Non-transitory computer-readable storage media may, where appropriate, be volatile, non-volatile, or a combination of volatile and non-volatile.

[0132] Enumeration of Embodiments Embodiment 1: A method comprising: accessing, by one or more computing systems, a biomolecular representation of a first biomolecule; processing the biomolecular representation of the first biomolecule with a machine learning model, wherein the machine learning model is trained using sequencing time series data related to biomolecular frequencies of a particular biomolecule, the sequencing time series data being obtained from directed evolution of a population of biomolecules over multiple enrichment rounds, the population of biomolecules in each enrichment round being a unique set of biomolecules relative to each other enrichment round, and the sequencing time series data for each enrichment round including a biomolecular frequency of each biomolecule in the population of biomolecules in the respective enrichment round; and outputting, by the machine learning model, an inferred fitness score for the population of biomolecules for each enrichment round by predicting the biomolecular frequency of the population of biomolecules in the respective enrichment round, taking into account the biomolecular frequencies of the population of biomolecules in one or more prior enrichment rounds.

[0133] Embodiment 2: The method of embodiment 1, further comprising determining whether a biological activity associated with the first biomolecule meets predetermined criteria for selection based on the inferred fitness score for the first biomolecule.

[0134] Embodiment 3: The method of embodiment 1 or 2, wherein the multiple enrichment rounds comprise at least three enrichment rounds, and at least one of the enrichment rounds is a control round in which the population of biomolecules is analyzed without the presence of the target protein.

[0135] Embodiment 4: The method of any one of embodiments 1 to 3, wherein the inferred fitness score for the first biomolecule is indicative of a biological activity of the first biomolecule against the target protein.

[0136] Embodiment 5: The method of any one of embodiments 1 to 4, wherein learning the inferential fitness scores in training the machine learning model comprises optimizing a Dirichlet multinomial loss function, the Dirichlet multinomial loss function utilizing an overdispersed multinomial distribution to account for the increasing difficulty associated with predicting biomolecule frequencies of the population of biomolecules in each enrichment round given the biomolecule frequencies of the population of biomolecules in prior enrichment rounds.

[0137] Embodiment 6: The method of any one of embodiments 1 to 5, wherein training the machine learning model further comprises calculating a Dirichlet loss negative log-likelihood between the predicted biomolecule frequencies and the actual biomolecule frequencies as a negative log-likelihood.

[0138] Embodiment 7: The method of any one of embodiments 1 to 6, wherein the inferred fitness score for the first biomolecule comprises an on-target fitness score associated with binding of the first biomolecule to the target protein.

[0139] Embodiment 8: The method of any one of embodiments 1 to 6, wherein the inferred fitness score for the first biomolecule comprises an off-target fitness score associated with the first biomolecule binding to the test device instead of the target protein.

[0140] Embodiment 9: The method of any one of embodiments 1 to 6, wherein the inferred fitness score for the first biomolecule comprises an on-target fitness score associated with the first biomolecule binding to the target protein and an off-target fitness score associated with the first biomolecule binding to the test device instead of the target protein, and the method further comprises determining the binding specificity of the first biomolecule based on a ratio of the on-target fitness score to the off-target fitness score.

[0141] Embodiment 10: The method of any one of embodiments 1 to 9, wherein the machine learning model comprises one or more neural networks.

[0142] Embodiment 11: The method of embodiment 10, wherein the one or more neural networks include a first neural network trained to predict on-target fitness scores associated with the biomolecule and a second neural network trained to predict off-target fitness scores associated with the biomolecule.

[0143] Embodiment 12: The method of any one of embodiments 1 to 11, further comprising generating a biomolecular representation of a first biomolecule, wherein the first biomolecule is a polypeptide corresponding to a first genotype, wherein generating comprises determining a plurality of amino acids of the first biomolecule, and for each amino acid of the plurality of amino acids, applying a function to determine a feature representation for each amino acid, and generating a genotype representation corresponding to the first genotype based on the plurality of feature representations associated with the plurality of amino acids.

[0144] Embodiment 13: The method of any one of embodiments 1 to 12, wherein the first biomolecule is within a population of biomolecules in multiple enrichment rounds.

[0145] Embodiment 14: The method of any one of embodiments 1 to 12, wherein the first biomolecule is not within the population of biomolecules in multiple enrichment rounds.

[0146] Embodiment 15: The method of any one of embodiments 1 to 14, wherein the sequencing time series data comprises DNA sequencing time series data, and the biomolecule frequencies of specific biomolecules indicate genotype frequencies.

[0147] Embodiment 16: The method of any one of embodiments 1 to 15, further comprising: processing a plurality of biomolecular representations associated with a plurality of respective second biomolecules through a machine learning model to determine a plurality of inferred fitness scores for each of the plurality of second biomolecules; and selecting one or more second biomolecules that meet predetermined criteria for selection based on the inferred fitness scores for the plurality of second biomolecules, wherein one or more of the selected second biomolecules each is associated with a low relative biomolecule frequency in a final round of the plurality of enrichment rounds.

[0148] Embodiment 17: The method of embodiment 16, further comprising: generating a genotype space based on biomolecule frequencies and inferred fitness scores for a plurality of second biomolecules; and selecting one or more second biomolecules by identifying one or more second biomolecules from one or more regions in the genotype space, wherein each of the one or more regions is associated with a particular biomolecule frequency range and a particular biomolecule fitness range.

[0149] Embodiment 18: The method of any one of embodiments 1 to 17, wherein the training further comprises pre-training an off-target model, wherein pre-training comprises identifying one or more off-target enrichment rounds from the plurality of enrichment rounds, and pre-training the off-target model based on sequencing time-series data in the one or more off-target enrichment rounds.

[0150] Embodiment 19: The method of embodiment 18, wherein the training further comprises accessing sequencing time series data from one or more on-target enrichment rounds from the plurality of enrichment rounds, and generating an on-target model based on the accessed sequencing time series data from the one or more on-target enrichment rounds and an off-target model.

[0151] Embodiment 20: The method of any one of embodiments 1 to 19, wherein the first biomolecule is a macrocycle.

[0152] Embodiment 21: The method of any one of embodiments 1 to 20, wherein the population of biomolecules is amplified by polymerase chain reaction (PCR) in each of multiple enrichment rounds.

[0153] Embodiment 22: The method of any one of embodiments 1 to 21, further comprising: processing a plurality of biomolecular representations associated with a plurality of respective second biomolecules with a machine learning model to determine a plurality of inferred fitness scores for each of the plurality of second biomolecules; and selecting one or more distinct biomolecules from the plurality of second biomolecules based on the inferred fitness scores for the plurality of second biomolecules, wherein the one or more distinct biomolecules satisfy predetermined criteria for selection, and wherein one or more of the distinct second biomolecules are each associated with a low relative biomolecule frequency in a final round of the plurality of enrichment rounds.

[0154] Embodiment 23: One or more non-transitory computer-readable storage media embodying software, which when executed accesses a biomolecular representation of a first biomolecule and processes the biomolecular representation of the first biomolecule with a machine learning model, the machine learning model being trained using sequencing time series data related to the biomolecular frequency of the particular biomolecule, the sequencing time series data being obtained from directed evolution of a population of biomolecules over multiple enrichment rounds, the population of biomolecules in each enrichment round being a unique set of biomolecules relative to each other enrichment round, and .... one or more non-transitory computer-readable storage media, wherein the sequence determination time series data includes a biomolecule frequency for each biomolecule in the population of biomolecules in each enrichment round, and training includes learning an inferred fitness score for the population of biomolecules for each enrichment round by predicting a biomolecule frequency for the population of biomolecules in each enrichment round given the biomolecule frequencies of the population of biomolecules in one or more prior enrichment rounds, and the machine learning model is operable to output, based on processing a biomolecular representation of the first biomolecule, an inferred fitness score for the first biomolecule.

[0155] Embodiment 24: A method and apparatus for detecting and analysing a biomolecule comprising: one or more processors; and a non-transitory memory, coupled to the processor and comprising instructions executable by the processor, wherein the processor, when executing the instructions, accesses a biomolecular representation of a first biomolecule; processes the biomolecular representation of the first biomolecule with a machine learning model; the machine learning model is trained using sequencing time series data related to biomolecular frequencies of the particular biomolecule; the sequencing time series data is obtained from directed evolution of a population of biomolecules over multiple enrichment rounds, wherein the population of biomolecules in each enrichment round is a unique set of biomolecules relative to each other enrichment round. the sequencing time series data at each enrichment round includes a biomolecule frequency for each biomolecule in the population of biomolecules at the respective enrichment round, the training includes learning an inferred fitness score for the population of biomolecules for each enrichment round by predicting a biomolecule frequency for the population of biomolecules at the respective enrichment round taking into account the biomolecule frequencies of the population of biomolecules at one or more prior enrichment rounds, and the machine learning model is operable to output an inferred fitness score for the first biomolecule based on processing the biomolecular representation of the first biomolecule.

[0156] Embodiment 25: Accessing, by one or more computing systems, a plurality of biomolecular representations of a plurality of respective biomolecules; and processing the plurality of biomolecular representations of the plurality of respective biomolecules with a machine learning model, wherein the machine learning model is trained using sequencing time series data related to biomolecular frequencies of particular biomolecules, the sequencing time series data being obtained from directed evolution of a population of biomolecules over multiple enrichment rounds, the population of biomolecules in each enrichment round being a unique set of biomolecules relative to each other enrichment round, the sequencing time series data in each enrichment round including the biomolecular frequency of each biomolecule in the population of biomolecules in the respective enrichment round; and wherein the training is performed using sequencing time series data related to biomolecular frequencies of particular biomolecules in one or more prior enrichment rounds. a machine learning model for outputting a plurality of inferred fitness scores for the plurality of biomolecules based on the processing of the plurality of biomolecular representations of the plurality of biomolecules; and selecting one or more biomolecules that meet predetermined criteria for selection based on the inferred fitness scores for the plurality of biomolecules, wherein one or more of the selected biomolecules are each associated with a low relative biomolecular frequency in a final round of the plurality of enrichment rounds.

[0157] Embodiment 26: The method of embodiment 25, further comprising generating a genotype space based on biomolecule frequencies and inferred fitness scores for the plurality of biomolecules, wherein selecting one or more biomolecules that meet predetermined criteria for selection comprises identifying one or more biomolecules from one or more regions within the genotype space, each of the one or more regions being associated with a particular biomolecule frequency range and a particular biomolecule fitness range.

[0158] Embodiment 27: The method of embodiment 25 or 26, further comprising determining whether a biological activity associated with the first biomolecule meets predetermined criteria for selection based on the inferred fitness score for the first biomolecule of the plurality of biomolecules.

[0159] Embodiment 28: The method of any one of embodiments 25 to 27, wherein the multiple enrichment rounds comprise at least three enrichment rounds, and at least one of the enrichment rounds is a control round in which the population of biomolecules is analyzed without the presence of the target protein.

[0160] Embodiment 29: The method of any one of embodiments 25 to 28, wherein the inferred fitness score for each of the plurality of biomolecules indicates the biological activity of the corresponding biomolecule against the target protein.

[0161] Embodiment 30: The method of any one of embodiments 25 to 29, wherein learning the inferential fitness score in training the machine learning model includes optimizing a Dirichlet multinomial loss function, the Dirichlet multinomial loss function utilizing an overdispersed multinomial distribution to account for the increasing difficulty associated with predicting the biomolecule frequencies of the population of biomolecules in each enrichment round, given the biomolecule frequencies of the population of biomolecules in previous enrichment rounds.

[0162] Embodiment 31: The method of any one of embodiments 25 to 30, wherein training the machine learning model further comprises calculating a Dirichlet loss negative log-likelihood between the predicted biomolecule frequencies and the actual biomolecule frequencies as a negative log-likelihood.

[0163] Embodiment 32: The method of any one of embodiments 25 to 31, wherein the inferred fitness score for each of the plurality of biomolecules comprises an on-target fitness score associated with binding of the corresponding biomolecule to the target protein.

[0164] Embodiment 33: The method of any one of embodiments 25 to 31, wherein the inferred fitness score for each of the plurality of biomolecules includes an off-target fitness score associated with the corresponding biomolecule binding to the test device instead of the target protein.

[0165] Embodiment 34: The method of any one of embodiments 25 to 31, wherein the inferred fitness scores for each of the plurality of biomolecules include an on-target fitness score associated with the corresponding biomolecule binding to the target protein and an off-target fitness score associated with the corresponding biomolecule binding to the test device instead of the target protein, and the method further includes determining the binding specificity of the corresponding biomolecule based on a ratio of the on-target fitness score to the off-target fitness score.

[0166] Embodiment 35: The method of any one of embodiments 25 to 34, wherein the machine learning model comprises one or more neural networks.

[0167] Embodiment 36: The method of embodiment 35, wherein the one or more neural networks include a first neural network trained to predict on-target fitness scores associated with the biomolecule and a second neural network trained to predict off-target fitness scores associated with the biomolecule.

[0168] Embodiment 37: The method of any one of embodiments 25 to 36, further comprising generating a biomolecular representation of a first biomolecule among the plurality of biomolecules, wherein the first biomolecule is a polypeptide corresponding to a first genotype, and generating the biomolecular representation of the first biomolecule further comprises determining a plurality of amino acids of the first biomolecule, and for each amino acid of the plurality of amino acids, applying a function to determine a feature representation for each amino acid, and generating a genotype representation corresponding to the first genotype based on the plurality of feature representations associated with the plurality of amino acids.

[0169] Embodiment 38: The method of any one of embodiments 25 to 37, wherein each of the plurality of biomolecules is within the population of biomolecules in multiple enrichment rounds.

[0170] Embodiment 39: The method of any one of embodiments 25 to 37, wherein each of the plurality of biomolecules is not within the population of biomolecules in multiple enrichment rounds.

[0171] Embodiment 40: The method of any one of embodiments 25 to 39, wherein the sequencing time series data comprises DNA sequencing time series data, and the biomolecule frequencies of specific biomolecules indicate genotype frequencies.

[0172] Embodiment 41: The method of any one of embodiments 25 to 40, wherein the training further comprises pre-training an off-target model, wherein pre-training comprises identifying one or more off-target enrichment rounds from the plurality of enrichment rounds, and pre-training the off-target model based on sequencing time series data in the one or more off-target enrichment rounds.

[0173] Embodiment 42: The method of embodiment 41, wherein the training further comprises accessing sequencing time series data from one or more on-target enrichment rounds from a plurality of enrichment rounds, and generating an on-target model based on the accessed sequencing time series data from the one or more on-target enrichment rounds and an off-target model.

[0174] Embodiment 43: The method of any one of embodiments 25 to 42, wherein the first biomolecule of the plurality of biomolecules is a macrocycle.

[0175] Embodiment 44: The method of any one of embodiments 25 to 43, wherein the population of biomolecules is amplified by polymerase chain reaction (PCR) in each of multiple enrichment rounds.

[0176] others As used herein, "or" is inclusive and not exclusive, unless clearly indicated otherwise or indicated otherwise by context. Thus, as used herein, "A or B" means "A, B, or both," unless clearly indicated otherwise or indicated otherwise by context. Moreover, "and" is both jointly and severally, unless clearly indicated otherwise or indicated otherwise by context. Thus, as used herein, "A and B" means "A and B, jointly or severally," unless clearly indicated otherwise or indicated otherwise by context.

[0177] The scope of the present disclosure encompasses all changes, substitutions, variations, alterations, and modifications to the exemplary embodiments described or illustrated herein that would be understood by a person skilled in the art. The scope of the present disclosure is not limited to the exemplary embodiments described or illustrated herein. Furthermore, although the present disclosure describes and illustrates each embodiment herein as including particular components, elements, features, functions, operations, or steps, any of these embodiments may include any combination or permutation of any of the components, elements, features, functions, operations, or steps described or illustrated anywhere herein that would be understood by a person skilled in the art. Furthermore, references in the appended claims to a device or system or a component of a device or system that is arranged, arranged, enabled, configured, enabled, enabled, or operates to perform a particular function encompass that device, system, or component, to the extent that the device, system, or component is so arranged, arranged, enabled, configured, enabled, enabled, or operates, regardless of whether it or that particular function is activated, turned on, or released. Furthermore, although the present disclosure describes or illustrates particular embodiments as providing certain advantages, the particular embodiments may provide none, some, or all of these advantages.

Claims

1. by one or more computing systems, accessing a biomolecular representation of a first biomolecule; processing the biomolecular representation of the first biomolecule with a machine learning model; the machine learning model is trained using sequencing time series data related to biomolecular frequencies of specific biomolecules; the sequencing time series data is obtained from directed evolution of a population of biomolecules over multiple rounds of enrichment; the population of biomolecules in each enrichment round is a unique set of biomolecules relative to each other enrichment round; the sequencing time series data for each enrichment round includes a biomolecule frequency for each biomolecule in the population of biomolecules in the respective enrichment round; processing the biomolecular representation of the first biomolecule, wherein the training comprises learning an inferred fitness score for the population of biomolecules for each enrichment round by predicting a biomolecule frequency for the population of biomolecules in the respective enrichment round, taking into account a biomolecule frequency for the population of biomolecules in one or more prior enrichment rounds; outputting, by a machine learning model, an inferred fitness score for the first biomolecule based on the processing of the biomolecular representation of the first biomolecule; A method comprising:

2. determining whether a biological activity associated with the first biomolecule meets predetermined criteria for selection based on the inferred fitness score for the first biomolecule; The method of claim 1 further comprising:

3. 2. The method of claim 1, wherein the plurality of enrichment rounds comprises at least three enrichment rounds, and at least one of the enrichment rounds is a control round in which the population of biomolecules is analyzed without the presence of a target protein.

4. 2. The method of claim 1, wherein the inferred fitness score for the first biomolecule is indicative of a biological activity of the first biomolecule against a target protein.

5. 2. The method of claim 1, wherein learning the inferred fitness scores in the training of the machine learning model comprises optimizing a Dirichlet multinomial loss function, the Dirichlet multinomial loss function utilizing an overdispersed multinomial distribution to account for increasing difficulty associated with predicting biomolecule frequencies of the population of biomolecules in each enrichment round given biomolecule frequencies of the population of biomolecules in prior enrichment rounds.

6. The training of the machine learning model includes: Calculating a Dirichlet loss negative log likelihood between the predicted biomolecule frequencies and the actual biomolecule frequencies as a negative log likelihood; The method of claim 5 further comprising:

7. 2. The method of claim 1, wherein the inferred fitness score for the first biomolecule comprises an on-target fitness score associated with binding of the first biomolecule to a target protein.

8. 2. The method of claim 1, wherein the inferred fitness score for the first biomolecule includes an off-target fitness score associated with the first biomolecule binding to a test device instead of a target protein.

9. wherein the inferred fitness scores for the first biomolecule include an on-target fitness score associated with binding of the first biomolecule to a target protein and an off-target fitness score associated with binding of the first biomolecule to a test device instead of the target protein, and the method comprises: determining the binding specificity of the first biomolecule based on a ratio of the on-target fitness score to the off-target fitness score; The method of claim 1 further comprising:

10. The method of claim 1 , wherein the machine learning model comprises one or more neural networks.

11. The one or more neural networks a first neural network trained to predict an on-target fitness score associated with a biomolecule; a second neural network trained to predict off-target fitness scores associated with the biomolecule; The method of claim 10, comprising:

12. generating the biomolecular representation of the first biomolecule, wherein the first biomolecule is a polypeptide corresponding to a first genotype, wherein the generating comprises: determining a plurality of amino acids of the first biomolecule; applying a function to each amino acid of said plurality of amino acids to determine a characterization for said each amino acid; generating a genotype representation corresponding to the first genotype based on the plurality of feature representations associated with the plurality of amino acids; The method of claim 1 , comprising:

13. 2. The method of claim 1, wherein the first biomolecule is within the population of biomolecules in the multiple enrichment rounds.

14. 2. The method of claim 1, wherein the first biomolecule is not within the population of biomolecules in the multiple enrichment rounds.

15. 2. The method of claim 1, wherein the sequencing time series data comprises DNA sequencing time series data, and biomolecule frequencies of specific biomolecules represent genotype frequencies.

16. processing a plurality of biomolecular representations associated with a plurality of respective second biomolecules through the machine learning model to determine a plurality of inferred fitness scores for the plurality of second biomolecules, respectively; selecting one or more second biomolecules that meet predetermined criteria for selection based on the inferred fitness scores for the plurality of second biomolecules, wherein one or more of the selected second biomolecules are each associated with a low relative biomolecule frequency in a final round of the plurality of enrichment rounds; The method of claim 1 further comprising:

17. generating a genotype space based on the biomolecule frequencies and the inferred fitness scores for the plurality of second biomolecules; selecting the one or more second biomolecules by identifying the one or more second biomolecules from one or more regions in the genotype space, each of the one or more regions being associated with a particular biomolecule frequency range and a particular biomolecule fitness range; 17. The method of claim 16, further comprising:

18. The training further includes pre-training an off-target model, the pre-training comprising: identifying one or more off-target enrichment rounds from said plurality of enrichment rounds; and pre-training the off-target model based on sequencing time-series data from the one or more off-target enrichment rounds; The method of claim 1 , comprising:

19. The training includes: accessing sequencing time series data from one or more on-target enrichment rounds from said plurality of enrichment rounds; generating an on-target model based on the accessed sequencing time-series data from the one or more on-target enrichment rounds and the off-target model; 20. The method of claim 18, further comprising:

20. The method of claim 1 , wherein the first biomolecule is a macrocycle.

21. 2. The method of claim 1, wherein the population of biomolecules is amplified by polymerase chain reaction (PCR) in each of the multiple enrichment rounds.

22. processing a plurality of biomolecular representations associated with a plurality of respective second biomolecules through the machine learning model to determine a plurality of inferred fitness scores for the plurality of second biomolecules, respectively; selecting one or more distinct biomolecules from the plurality of second biomolecules based on the inferred fitness scores for the plurality of second biomolecules, wherein the one or more distinct biomolecules meet predetermined criteria for selection, and wherein one or more of the distinct biomolecules are each associated with a low relative biomolecule frequency in a final round of the plurality of enrichment rounds; The method of claim 1 further comprising:

23. One or more non-transitory computer-readable storage media embodying software that, when executed, accessing a biomolecular representation of a first biomolecule; processing the biomolecular representation of the first biomolecule with a machine learning model; the machine learning model is trained using sequencing time series data related to biomolecular frequencies of specific biomolecules; the sequencing time series data is obtained from directed evolution of a population of biomolecules over multiple rounds of enrichment; the population of biomolecules in each enrichment round is a unique set of biomolecules relative to each other enrichment round; the sequencing time series data in each enrichment round includes a biomolecule frequency for each biomolecule in the population of biomolecules in each enrichment round; the training includes learning an inferred fitness score for the population of biomolecules for each enrichment round by predicting biomolecule frequencies for the population of biomolecules in each enrichment round given biomolecule frequencies for the population of biomolecules in one or more prior enrichment rounds; outputting, by a machine learning model, an inferred fitness score for the first biomolecule based on the processing of the biomolecular representation of the first biomolecule. One or more non-transitory computer-readable storage media operable to:

24. one or more processors; and a non-transitory memory coupled to the processors and including instructions executable by the processors, the processors, when executing the instructions, accessing a biomolecular representation of a first biomolecule; processing the biomolecular representation of the first biomolecule with a machine learning model; the machine learning model is trained using sequencing time series data related to biomolecular frequencies of specific biomolecules; the sequencing time series data is obtained from directed evolution of a population of biomolecules over multiple rounds of enrichment; the population of biomolecules in each enrichment round is a unique set of biomolecules relative to each other enrichment round; the sequencing time series data in each enrichment round includes a biomolecule frequency for each biomolecule in the population of biomolecules in each enrichment round; the training includes learning an inferred fitness score for the population of biomolecules for each enrichment round by predicting biomolecule frequencies for the population of biomolecules in the respective enrichment round given biomolecule frequencies for the population of biomolecules in one or more prior enrichment rounds; outputting, by a machine learning model, an inferred fitness score for the first biomolecule based on the processing of the biomolecular representation of the first biomolecule. A system capable of operating as follows.

25. by one or more computing systems, accessing a plurality of biomolecular representations of a plurality of respective biomolecules; processing the plurality of biomolecular representations of the plurality of respective biomolecules with a machine learning model; the machine learning model is trained using sequencing time series data related to biomolecular frequencies of specific biomolecules; the sequencing time series data is obtained from directed evolution of a population of biomolecules over multiple rounds of enrichment; the population of biomolecules in each enrichment round is a unique set of biomolecules relative to each other enrichment round; the sequencing time series data in each enrichment round includes a biomolecule frequency for each biomolecule in the population of biomolecules in each enrichment round; processing the plurality of biomolecular representations for each of the plurality of respective biomolecules, wherein the training comprises learning an inferred fitness score for the population of biomolecules for each enrichment round by predicting a biomolecule frequency for the population of biomolecules in the respective enrichment round, taking into account a biomolecule frequency for the population of biomolecules in one or more prior enrichment rounds; outputting, by the machine learning model, a plurality of inferred fitness scores for each of the plurality of biomolecules based on the processing of the plurality of biomolecular representations of the plurality of biomolecules; selecting one or more biomolecules that meet predetermined criteria for selection based on the inferred fitness scores for the plurality of biomolecules, wherein one or more of the selected biomolecules are each associated with a low relative biomolecule frequency in a final round of the plurality of enrichment rounds; A method comprising:

26. generating a genotype space based on the biomolecule frequencies and the inferred fitness scores for the plurality of biomolecules; Selecting the one or more biomolecules that meet the predetermined criteria for selection includes identifying the one or more biomolecules from one or more regions within the genotype space, each of the one or more regions being associated with a particular biomolecule frequency range and a particular biomolecule fitness range.

26. The method of claim 25.