Systems and methods for mapping molecules to interfaces
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-13
- Publication Date
- 2026-08-14
Smart Images

Figure CN122580701A_ABST
Abstract
Description
[0001] Cross-referencing
[0002] This application claims priority to U.S. Provisional Patent Application 63 / 609,755, filed December 13, 2023, entitled “SYSTEM AND METHOD FOR MAPPINGMOLECULES INTO INTERFACES”, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This disclosure relates to computing, selective visual display systems, data processing, molecular discovery, machine learning, and device interfaces. Background Technology
[0004] Finding molecules with similar properties and / or characteristics can help identify potential drugs for treating diseases. However, visualizing molecules across multiple properties and / or characteristics can be challenging. This can make it difficult to discover new and effective treatments.
[0005] There is a need for improved or alternative approaches to molecular discovery. Summary of the Invention
[0006] The embodiments described herein relate to systems and methods for mapping molecules to interfaces, generating maps for display and interaction, or providing interfaces or maps. Molecules can be organized into groups and then sorted or sampled for molecular discovery. The systems and methods described herein can be used to map proteins, protein-like molecules such as antibodies or antigens or fragments thereof, small molecule drugs, or biomolecules. Improved or alternative approaches to molecular discovery are needed.
[0007] The systems and methods disclosed herein can be used for various applications, such as, for example, drug discovery, antibody discovery or optimization (e.g., format conversion, humanization), monitoring immune responses (e.g., after immunization or vaccination), diagnosis, monitoring disease progression, etc. The systems and methods disclosed herein may be particularly useful in prospective drug discovery.
[0008] In antibody discovery, understanding the diversity and specificity of immune responses can be important not only for identifying novel conjugates but also for identifying antibodies with better therapeutic potential. This paper describes a visualization method that enables structural and biophysical comparisons of antibody libraries by representing each antibody / antigen or each part of an antibody / antigen as points on a map, where the spatial arrangement reflects their similarities. The systems and methods described in this paper can help analyze the quality of antibody immune responses through complementation site diversity, assess the impact of immunological methods or different genetic backgrounds on antibody complementation site diversity, and sample antibody complementations for therapeutic activity.
[0009] According to one aspect, a computer-implemented system for mapping molecules to an interface is provided. The system has: a processing subsystem including one or more processors and one or more memories coupled to the processors; the processing subsystem being configured such that the system: receives input molecules, wherein the input molecules are a set of one or more sets of one or more molecules, wherein each molecule is defined as a sequence of information, structure, or properties; encodes the sequence of information, structure, or properties; generates a dataset by processing the input molecules, wherein the dataset includes the encoded information sequence, three-dimensional coordinates of the information sequence for each molecule defining the molecular structure, and features; transforms the dataset by feature extraction to generate a molecular map and features, thereby reducing a high number of dimensions to a lower-dimensional data representation that can be indicated by the interface, while capturing valuable data in the lower-dimensional data representation; generates a map user interface including a molecular map as a representation of the lower-dimensional data representation of the dataset; and provides the map user interface.
[0010] In some embodiments, the map user interface includes a visual interface, and the lower-dimensional data representation can be visualized in the visual interface.
[0011] In some embodiments, the input molecule is an antibody or its antigen-binding fragment and / or antigen, and the molecular map is a complementary site map.
[0012] In some embodiments, the input molecule is an antigen, and the molecular map is an epitope map.
[0013] In some embodiments, the features include fingerprints, wherein the processing subsystem uses machine learning or statistical methods involving one or more of dimensionality reduction, reconstructive autoencoders, variational autoencoders, adversarial autoencoders, neural networks, graph neural networks, attention networks, recurrent networks, and generative models to extract features and generate fingerprints.
[0014] In some embodiments, the system has a data storage device for a molecular databank, wherein each molecule is assigned a unique index.
[0015] In some embodiments, the system compares a molecule with a database of molecules and assigns the molecule an index to the closest molecule in the database, wherein molecules assigned the same index have similar sequences, structures, or properties.
[0016] In some embodiments, feature extraction includes layer-by-layer embedding, wherein features of each dataset in the multiple datasets or features of each part of the multiple datasets are extracted together or separately for multiple datasets or multiple parts of a dataset.
[0017] In some embodiments, features can be drawn as overlapping visual layers.
[0018] In some embodiments, the map user interface that includes a visualization of the dataset has control inputs that allow layers to be viewed individually or in relation to other layers.
[0019] In some embodiments, one or more layers are used for training feature extraction, and one or more other layers are used for testing feature extraction.
[0020] In some embodiments, feature extraction includes arranging embeddings in clusters.
[0021] In some embodiments, feature extraction includes embedding individual molecules.
[0022] In some embodiments, feature extraction includes generating clusters around molecules of interest.
[0023] In some embodiments, feature extraction includes sampling from a molecular map.
[0024] In some embodiments, feature extraction includes encoding scores in a molecular map.
[0025] In some embodiments, the map user interface includes visualization of extracted features of the dataset.
[0026] In some embodiments, the system has a user device for displaying a map user interface.
[0027] In some embodiments, the map user interface is one-dimensional, two-dimensional, three-dimensional, four-dimensional, or higher-dimensional.
[0028] In some embodiments, the layer-by-layer embedding is a time series that varies over time, and the map user interface includes a two-dimensional or three-dimensional embedding that varies over time as a time series representing a four-dimensional embedding.
[0029] In some embodiments, the map user interface includes visualization of clusters around molecules of interest.
[0030] In some embodiments, the map user interface includes one or more clusters of molecules.
[0031] In some embodiments, the processing subsystem performs feature extraction on each molecule in the dataset to obtain individual molecule embeddings.
[0032] In some embodiments, layer-by-layer embedding includes individual molecular embedding.
[0033] In some embodiments, the user interface uses extracted features to characterize and / or obtain information about the input molecule, wherein the input molecule may optionally come from the cluster of interest.
[0034] In some embodiments, the information includes the extent, nature, and / or robustness of the immune response (against an antigen such as an immunogen or vaccine).
[0035] In some embodiments, the input molecule includes an antibody or an antigen-binding fragment thereof, and the information includes the amino acid sequence or structure or properties of one or more of the antibody or antigen-binding fragment.
[0036] In some embodiments, the input molecule includes an antigen, and the information includes the amino acid sequence of one or more of the antigens.
[0037] In some embodiments, the system allows a user to modify the amino acid sequence, and the system predicts the effect of the modification on the characteristics of the molecule (binding, function, stability, expressibility, affinity, immunogenicity, etc.).
[0038] In some embodiments, modifications are made to amino acid substitutions, deletions, and / or additions in one or more CDRs, variable regions, frame regions, and / or constant regions of the antibody or its antigen-binding fragment.
[0039] In some embodiments, the modification is the humanization, deimmunization, glycosylation, or deglycosylation of the antibody or its antigen-binding fragment.
[0040] In some embodiments, the system allows users to import other molecules and determine their similarity to the input molecule.
[0041] In some embodiments, the other molecules are other antibodies or their antigen-binding fragments, and the similarity is complementary site similarity.
[0042] In some embodiments, the output includes antibodies or antigen-binding fragments thereof selected from the map or variants thereof.
[0043] In some embodiments, the user synthesizes an input molecule or an output molecule or a variant thereof, or causes an input molecule or an output molecule or a variant thereof to be synthesized.
[0044] In some embodiments, the information provided by the system is used to manufacture molecules.
[0045] In some embodiments, the input molecule includes a single-domain antibody or an antigen-binding fragment thereof, and the output molecule includes an antibody or an antigen-binding fragment thereof selected from conventional antibodies, single-domain antibodies, single-chain variable fragments, humanized antibodies, or chimeric antibodies.
[0046] In some embodiments, the processing subsystem connects the features of the parts of the molecule together to have the overall features of the molecule.
[0047] In some embodiments, the antibody or its antigen-binding fragment includes antibodies derived from a species, including but not limited to mice, cattle, rabbits, camels, llamas, humans, alpacas, and standard species.
[0048] In some embodiments, the encoded information sequence refers to an encoded amino acid sequence or structure or property.
[0049] In some embodiments, the processing subsystem outputs the selected molecule.
[0050] In some embodiments, a computer-implemented system is provided to produce an article of articles obtained from the selected molecular output.
[0051] In some embodiments, a product obtained by a computer-implemented system is provided.
[0052] According to another aspect, a computer-implemented method for mapping molecules to a visual interface is provided. The method involves: receiving input molecules, wherein the input molecules are one or more groups of one or more molecules, each molecule being defined as an information sequence; encoding the information sequence; generating a dataset by processing the input molecules, wherein the dataset includes the encoded information sequence, three-dimensional coordinates of the information sequence for each molecule defining its molecular structure, biophysical properties, features, and fingerprints; generating a transformed dataset through feature extraction and fingerprinting to generate a molecular map, thereby reducing a high number of dimensions to a lower-dimensional data representation that can be indicated by the interface, while capturing valuable data in the lower-dimensional data representation; generating a map user interface, the map user interface including the molecular map as a representation of the lower-dimensional data representation of the dataset; and providing the map user interface.
[0053] In some embodiments, the map user interface includes a visual interface, and the lower-dimensional data representation can be visualized in the visual interface.
[0054] In some embodiments, the input molecule is a protein or protein-like molecule comprising an antibody or its antigen-binding fragment and / or an antigen.
[0055] In some embodiments, the antibody or its antigen-binding fragment includes antibodies derived from a species, including but not limited to mice, cattle, rabbits, camels, llamas, humans, alpacas, and standard species.
[0056] In some embodiments, the input molecule is an antibody or its antigen-binding fragment and / or antigen, and the molecular map is a complementary site map.
[0057] In some embodiments, the input molecule is an antigen, and the molecular map is an epitope map.
[0058] In some embodiments, the features include fingerprints, wherein the processing subsystem uses machine learning or statistical methods involving one or more of dimensionality reduction, reconstructive autoencoders, variational autoencoders, adversarial autoencoders, neural networks, graph neural networks, attention networks, recurrent networks, and generative models to extract features and generate fingerprints.
[0059] In some embodiments, the method involves storing a database of molecules in a data storage device, wherein each molecule is assigned a unique index.
[0060] In some embodiments, the method involves comparing a molecule with a database of molecules and assigning the molecule an index to the closest molecule in the database of molecules, wherein molecules assigned the same index have similar sequences, structures, or properties.
[0061] In some embodiments, the method involves feature extraction with layer-by-layer embedding, wherein features of each dataset in the multiple datasets or features of each part of the multiple datasets are extracted together or separately for multiple datasets or multiple parts of a dataset.
[0062] In some embodiments, features can be drawn as overlapping visual layers.
[0063] In some embodiments, the method involves providing a map user interface that includes a visualization of a dataset, having control inputs that enable viewing layers individually or in relation to other layers.
[0064] In some embodiments, one or more layers are used for training feature extraction, and one or more other layers are used for testing feature extraction.
[0065] In some embodiments, feature extraction includes arranging embeddings in clusters.
[0066] In some embodiments, feature extraction includes embedding individual molecules.
[0067] In some embodiments, feature extraction includes generating clusters around molecules of interest.
[0068] In some embodiments, feature extraction includes sampling from a molecular map.
[0069] In some embodiments, feature extraction includes encoding scores in a molecular map.
[0070] In some embodiments, the map user interface includes visualization of extracted features of the dataset.
[0071] In some embodiments, the method involves using a user device to display a map user interface.
[0072] In some embodiments, the map user interface is one-dimensional, two-dimensional, three-dimensional, four-dimensional, or higher-dimensional.
[0073] In some embodiments, the layer-by-layer embedding is a time series that varies over time, and the map user interface includes a two-dimensional or three-dimensional embedding that varies over time as a time series representing a four-dimensional embedding.
[0074] In some embodiments, the method involves using a map user interface to provide visualization of clusters around molecules of interest.
[0075] In some embodiments, the map user interface includes one or more clusters of molecules.
[0076] In some embodiments, the method involves performing feature extraction on individual molecules of a dataset to obtain individual molecule embeddings.
[0077] In some embodiments, layer-by-layer embedding includes individual molecular embedding.
[0078] In some embodiments, the method involves using extracted features to characterize and / or obtain information about input molecules, wherein the input molecules may optionally come from clusters of interest.
[0079] In some embodiments, the information includes the extent, nature, and / or robustness of the immune response (against an antigen such as an immunogen or vaccine).
[0080] In some embodiments, the input molecule includes an antibody or an antigen-binding fragment thereof, and the information includes an amino acid sequence of one or more of the antibody or antigen-binding fragment.
[0081] In some embodiments, the input molecule includes an antigen, and the information includes the amino acid sequence of one or more of the antigens.
[0082] In some embodiments, the method involves modifying an amino acid sequence, and wherein the system predicts the effect of the modification on the characteristics of the molecule (binding, function, stability, expressibility, affinity, immunogenicity, etc.).
[0083] In some embodiments, modifications are made to amino acid substitutions, deletions, and / or additions in one or more CDRs, variable regions, frame regions, and / or constant regions of the antibody or its antigen-binding fragment.
[0084] In some embodiments, the modification is the humanization, deimmunization, glycosylation, or deglycosylation of the antibody or its antigen-binding fragment.
[0085] In some embodiments, the method allows users to import other molecules and determine their similarity to the input molecule.
[0086] In some embodiments, the molecule is another antibody or its antigen-binding fragment, and the similarity is complementary site similarity.
[0087] In some embodiments, the output includes antibodies or antigen-binding fragments thereof selected from the map or variants thereof.
[0088] In some embodiments, the user synthesizes an input molecule or an output molecule or a variant thereof, or causes an input molecule or an output molecule or a variant thereof to be synthesized.
[0089] In some embodiments, the method involves using information provided by the system to manufacture molecules.
[0090] In some embodiments, the input molecule includes a single-domain antibody or an antigen-binding fragment thereof, and the output molecule includes an antibody or an antigen-binding fragment thereof selected from conventional antibodies, single-domain antibodies, single-chain variable fragments, humanized antibodies, or chimeric antibodies.
[0091] In some embodiments, the method involves linking features of different parts of a molecule together to have the overall features of the molecule.
[0092] In some embodiments, the encoded information sequence refers to the encoded amino acid sequence.
[0093] In some embodiments, the method involves outputting a selected molecule.
[0094] In some embodiments, a method implemented by a computer provides an article of selected molecular output obtained from the article.
[0095] In some embodiments, a product obtained by a computer-implemented method is provided.
[0096] In some embodiments, the method involves the step of generating molecules identified or selected from a map user interface.
[0097] In some embodiments, a product obtained by a computer-implemented method is provided.
[0098] In some embodiments, the product is an antibody or an antigen-binding fragment thereof.
[0099] According to another aspect, one or more non-transitory computer-readable media having stored thereon machine-interpretable instructions, which, when executed by a processing subsystem, cause the processing subsystem to perform a method for mapping molecules to a visual interface, the method comprising: receiving input molecules, wherein the input molecules are one or more groups of one or more molecules, wherein each molecule is defined as a sequence of information, structure, or attribute; encoding the sequence of information, structure, or attribute; generating a dataset by processing the input molecules, wherein the dataset includes the encoded information sequence, three-dimensional coordinates, features, and fingerprints of the information sequence for each molecule used to define the molecular structure; generating a transformed dataset by feature extraction and fingerprinting to generate a molecular map, thereby reducing a higher number of dimensions to a lower-dimensional data representation that can be indicated by the interface, while capturing valuable data in the lower-dimensional data representation; generating a map user interface, the map user interface including the molecular map as a representation of the lower-dimensional data representation of the dataset; and providing the map user interface.
[0100] According to another aspect, a computer-implemented system for a molecule-related interface is provided. The system has: a processing subsystem including one or more processors and one or more memories coupled to the processors; the processing subsystem is configured such that the system: receives input molecules, wherein the input molecules are a set of one or more sets of molecules, each molecule being defined as a sequence of information, structure, or properties; encodes the sequence of information, structure, or properties; generates a dataset by processing the input molecules, wherein the dataset includes the encoded information sequence, three-dimensional coordinates of the information sequence for each molecule defining the molecular structure, features, and fingerprints; generates a transformed dataset by feature extraction and fingerprinting, thereby reducing a higher number of dimensions to a lower-dimensional data representation that can be indicated by the interface, while capturing valuable data in the lower-dimensional data representation; generates one or more metrics from the transformed dataset, wherein the one or more metrics include a lower-dimensional data representation of the dataset and summarize the properties of the input molecules; and provides one or more metrics to the interface.
[0101] According to another aspect, a computer-implemented system for a visual interface for mapping molecules is provided. The system has: a processing subsystem including one or more processors and one or more memories coupled to the processors; the processing subsystem providing a map user interface, wherein the map user interface: receives input molecules, wherein the input molecules are a set of one or more sets of molecules, wherein each molecule is defined as a sequence of information, structure, or properties; and provides a map interface including a molecular map as a representation of a lower-dimensional data representation of the dataset of input molecules, wherein the dataset includes a sequence of information for each molecule, three-dimensional coordinates of a sequence of information for defining the molecular structure of each molecule, and features; wherein the molecular map includes a transformed dataset generated through feature extraction and fingerprinting, thereby reducing the higher number of dimensions to a lower-dimensional data representation that can be indicated by the interface, while capturing valuable data in the lower-dimensional data representation.
[0102] In some embodiments, the input molecule is a protein or protein-like molecule that includes antibodies and antigens.
[0103] In some embodiments, antibodies include species-specific antibodies, including but not limited to mice, cattle, rabbits, camels, llamas, humans, alpacas, and standard species.
[0104] In some embodiments, the input molecules are antibodies and antigens, and the molecular map is a complementary site map.
[0105] In some embodiments, the input molecule is an antigen, and the molecular map is an epitope map.
[0106] In some embodiments, the map user interface includes layer-by-layer embedding to provide layers for map visualization.
[0107] In some embodiments, the map user interface renders features as overlaid map visualization layers.
[0108] In some embodiments, the map user interface has control inputs that enable viewing layers individually or in relation to other layers.
[0109] In some embodiments, the map user interface has control inputs to add or remove layers in the layers used for map visualization.
[0110] In some embodiments, the map user interface includes visualization of extracted features of the dataset.
[0111] In some embodiments, the map user interface receives one or more reference molecules or target molecules, wherein the transformation of the dataset is based on one or more reference molecules or target molecules.
[0112] In some embodiments, the map user interface receives one or more scores for the molecule, wherein the scores include expressibility scores and fuzzy filtering scores.
[0113] In some embodiments, map visualization includes one or more clusters corresponding to molecules, wherein the map user interface receives clustering control commands to update the map visualization with hyperparameters of one or more clusters.
[0114] In some embodiments, the map visualization displays one or more scores about the molecule, including expressibility scores or fuzzy screening scores.
[0115] In some embodiments, the map user interface receives control commands for sampling from at least a portion of the map visualization.
[0116] In some embodiments, the map user interface receives control commands for editing samples extracted from at least a portion of the map visualization.
[0117] In some embodiments, the map user interface receives drawing settings corresponding to the visualization characteristics of the map visualization.
[0118] In some embodiments, the user equipment may display a map user interface.
[0119] In some embodiments, the map user interface is one-dimensional, two-dimensional, three-dimensional, four-dimensional, or higher-dimensional.
[0120] In some embodiments, the layer-by-layer embedding is a time series that varies over time, and the map user interface includes a two-dimensional or three-dimensional embedding that varies over time as a time series representing a four-dimensional embedding.
[0121] In some embodiments, the map user interface includes visualization of clusters around molecules of interest.
[0122] In some embodiments, the map user interface includes one or more clusters of molecules.
[0123] In some embodiments, the map user interface includes individual molecule embeddings.
[0124] In some embodiments, layer-by-layer embedding includes individual molecular embedding.
[0125] In some embodiments, the map user interface receives scores.
[0126] In some embodiments, the map user interface exports a file.
[0127] In some embodiments, the map user interface includes one or more buttons for adding or removing layers, one or more buttons for receiving input molecules, one or more buttons for adding or removing individual molecules, and one or more buttons for importing fractions.
[0128] In some embodiments, the map user interface includes multiple settings selected from the group consisting of: navigation settings for map visualization, drawing settings, settings for encoding scores in map visualization, clustering settings, settings for editing samples, sample settings, reporting settings, map analysis settings, and export settings.
[0129] According to one aspect, a computer-implemented system is provided for mapping proteins, protein-like molecules, or fragments thereof onto a visual interface and generating maps for display and interaction. The system includes a processing subsystem comprising one or more processors and one or more memories coupled to the processors. The processing subsystem is configured such that the system: receives one or more sets of one or more input data of proteins, protein-like molecules, or fragments thereof, wherein the data includes features of one or more proteins, protein-like molecules, or fragments thereof; generates at least one dataset by processing one or more sets of one or more input data of proteins, protein-like molecules, or fragments thereof; and transforms the data and one or more of the at least one dataset by feature extraction or feature selection to generate maps and additional features, thereby reducing a higher-dimensional data representation to a lower-dimensional data representation for visualization via a visual interface, wherein the lower-dimensional data representation captures valuable information from the data and one or more of the at least one dataset, and the lower-dimensional data representation includes one or more clusters of proteins, protein-like molecules, or fragments thereof. The process includes: generating a visual map interface, which comprises a map representing a lower-dimensional data representation of the dataset, the visual representation comprising a visualization of the dataset as one or more layers of proteins, protein-like molecules, or fragments thereof, each layer comprising one or more clusters of proteins, protein-like molecules, or fragments thereof; providing the visual map interface with tools for interacting with the map, wherein the interaction with the map includes one or more of the following: inspecting, searching, sampling, clustering, and analyzing one or more proteins, protein-like molecules, or fragments thereof or newly generated proteins, protein-like molecules, or fragments thereof; receiving commands or detecting interactions with the map at the visual map interface via tools; updating the map based on commands or interactions; and triggering an update of the visual map interface using the updated map.
[0130] In some embodiments, the protein, protein-like molecule or fragment thereof is selected from: antibody, antigen, lectin, receptor, ligand, enzyme or fragment thereof.
[0131] In some embodiments, the protein, protein-like molecule or fragment thereof includes an antibody or fragment thereof and / or an antigen or fragment thereof, and wherein the map is a complementary site map or an epitope map or a map including the protein or protein-like molecule or fragment thereof.
[0132] In some embodiments, proteins, protein-like molecules, or fragments thereof include antibodies or antibody fragments thereof, and wherein the data includes the structure of one or more antibodies or antibody fragments thereof, the amino acid sequence of one or more of the antibodies or antibody fragments thereof, the amino acid atomic or molecular coordinates, and one or more of the biophysical properties of one or more antibodies or antibody fragments thereof.
[0133] In some embodiments, the protein, protein-like molecule, or fragment thereof is selected from conventional antibodies, antibody-like molecules, artificial antibodies, antibody mimics, single-domain antibodies, single-chain antibodies, humanized antibodies, chimeric antibodies, or fragments thereof.
[0134] In some embodiments, the fragment includes an antigen-binding fragment or an antigen-binding domain.
[0135] In some embodiments, the antigen-binding fragment or antigen-binding domain is selected from one or more complementarity-determining regions and / or one or more frame regions, one or more variable domains, or complementary sites.
[0136] In some embodiments, the visual representation of lower-dimensional data includes different colors and / or marker shapes and / or marker sizes and / or color transparency and / or color gradients to indicate one or more layers and one or more clusters of proteins, protein-like molecules or fragments thereof.
[0137] In some embodiments, the lower-dimensional data representation is a one-dimensional, two-dimensional, three-dimensional, or four-dimensional data representation.
[0138] In some embodiments, the features include fingerprints, wherein the processing subsystem uses machine learning or statistical methods involving one or more of dimensionality reduction, reconstructive autoencoders, variational autoencoders, adversarial autoencoders, neural networks, graph neural networks, attention networks, recurrent networks, sequence processing algorithms, image processing algorithms, computer vision algorithms, and identity transformations to extract features and generate fingerprints.
[0139] In some embodiments, the processing subsystem causes the system to cluster data and one or more data sets to generate one or more clusters of proteins, protein-like molecules or fragments thereof.
[0140] In some embodiments, the processing subsystem causes the system to encode the raw data and generate additional features from the encoded data.
[0141] In some embodiments, the visual representation overlays one or more layers of proteins, protein-like molecules, or fragments thereof as part of the visualization of the dataset, wherein tools trigger one or more layers to move to different locations or levels, or remove them from the map, or change the order of the displayed layers, or zoom in or out of one or more layers, or move across layers in the map.
[0142] In some embodiments, the processing subsystem enables the system to perform map analysis, wherein the map analysis includes one or more of the following: generating clusters around proteins, protein-like molecules or fragments of interest, arranging embeddings in the clusters, embedding layer by layer, embedding individual proteins, protein-like molecules or fragments of them, sampling from the map, encoding scores in the map, wherein the map contains one or more clusters and visualizing one or more clusters.
[0143] In some embodiments, the processing subsystem causes the system to generate or compute one or more clusters of proteins, protein-like molecules, or fragments thereof, and wherein the map user interface includes visualization of one or more clusters of proteins, protein-like molecules, or fragments thereof.
[0144] In some embodiments, feature extraction includes extracting useful information from a dataset, and feature selection includes selecting a subset of the dataset of proteins, protein-like molecules, or fragments thereof.
[0145] In some embodiments, the processing subsystem causes the system to transform one or more of the data and datasets to generate a map by one or more of sequencing and clustering, sampling, intersection of data subsets, and subtraction of data subsets.
[0146] In some embodiments, the processing subsystem causes the system to partition or segment a digital map into multiple map tiles, label each of one or more clusters with corresponding map tiles in the multiple map tiles, and display one or more clusters within the multiple map tiles using labels, wherein the visualization indicates the multiple map tiles and one or more clusters.
[0147] In some embodiments, the processing subsystem causes the system to: (i) intersect one or more layers of proteins, protein-like molecules, or fragments thereof, or (ii) subtract one or more layers of proteins, protein-like molecules, or fragments thereof, or (iii) add one or more layers of proteins, protein-like molecules, or fragments thereof to update the map based on commands or interactions.
[0148] In some embodiments, the tools at the visual map interface include sampling tools for sampling proteins, protein-like molecules, or fragments thereof from one or more clusters, wherein the processing subsystem causes the system to update the map by sampling proteins, protein-like molecules, or fragments thereof in response to activation of the sampling tools, and triggers an update of the visual map interface using the updated map to visualize the sampling.
[0149] In some embodiments, the processing subsystem causes the system to subtract the library of non-immunized proteins, protein-like molecules, or fragments thereof from the library of immunized proteins, protein-like molecules, or fragments thereof to filter out non-specific proteins, protein-like molecules, or fragments thereof, and to reduce the search space used for sampling and searching for specific molecular candidates for one or more targets. If there are multiple layers or datasets for the immunized library, the subsystem causes the system to intersect the layers or datasets after the subtraction to further reduce the search space.
[0150] In some embodiments, the processing subsystem causes the system to subtract the library of proteins, protein-like molecules, or fragments immunized against one or more targets from the library of proteins, protein-like molecules, or fragments immunized against the target of interest, in order to filter out non-binding portions of the target of interest and reduce the search space used for sampling and searching for specific molecular candidates against the target of interest. If there are multiple layers or datasets in the immune library against the target of interest, the subsystem causes the system to intersect the layers or datasets after the subtraction, in order to further reduce the search space.
[0151] In some embodiments, the processing subsystem enables the system to export or report inspections, searches, sampling, clustering, and analyses of proteins, protein-like molecules, or fragments thereof through text, tables, graphs, or visualizations.
[0152] According to one aspect, a computer processing method is provided for mapping proteins, protein-like molecules, or fragments thereof onto a visual interface and generating digital maps for display and interaction. The method includes: receiving one or more sets of one or more input sets of data on proteins, protein-like molecules, or fragments thereof, wherein the data includes features of one or more proteins, protein-like molecules, or fragments thereof; generating at least one dataset by processing one or more sets of one or more input sets of data on proteins, protein-like molecules, or fragments thereof; transforming the data and one or more of the at least one dataset by feature extraction or feature selection to generate a map and additional features, thereby reducing a higher-dimensional data representation to a lower-dimensional data representation for visualization via a visual interface, wherein the lower-dimensional data representation captures valuable information from the data and one or more of the at least one dataset, and the lower-dimensional data representation includes one or more clusters of proteins, protein-like molecules, or fragments thereof, wherein the proteins, protein-like molecules, or fragments thereof include one or more input proteins, protein-like molecules, or fragments thereof. Protein molecules or fragments thereof, or generated proteins, protein-like molecules or fragments thereof; generating a visual map interface, the visual map interface comprising a map as a visual representation of a lower-dimensional data representation of a dataset, the visual representation comprising visualizations of the dataset as one or more layers of proteins, protein-like molecules or fragments thereof, each layer comprising one or more clusters of proteins, protein-like molecules or fragments thereof; providing the visual map interface with tools for interacting with the map, wherein the interaction with the map comprises one or more of the following: inspection, searching, sampling, clustering and analysis of one or more proteins, protein-like molecules or fragments thereof or newly generated proteins, protein-like molecules or fragments thereof; receiving commands or detecting interactions with the map at the visual map interface via tools; and triggering updates to the visual map interface and the map based on commands or interactions.
[0153] According to one aspect, a computer-readable medium encoded with instructions is provided, which, when executed by a processor, cause the processor to map proteins, protein-like molecules, or fragments thereof onto a visual interface and generate a digital map for display and interaction. The instructions include instructions for performing the following: receiving one or more sets of one or more input data of proteins, protein-like molecules, or fragments thereof, wherein the data includes features of one or more proteins, protein-like molecules, or fragments thereof; generating at least one dataset by processing one or more sets of one or more input data of proteins, protein-like molecules, or fragments thereof; transforming the data and one or more of the at least one dataset by feature extraction or feature selection to generate a map and additional features, thereby reducing a higher-dimensional data representation to a lower-dimensional data representation for visualization via a visual interface, wherein the lower-dimensional data representation captures valuable information from the data and one or more of the at least one dataset, and the lower-dimensional data representation includes one or more clusters of proteins, protein-like molecules, or fragments thereof, the proteins, protein-like molecules, or fragments thereof including one or more input data of proteins, protein-like molecules, or fragments thereof. The method involves: inputting proteins, protein-like molecules, or fragments thereof, or generating proteins, protein-like molecules, or fragments thereof; generating a visual map interface, which includes a map representing a lower-dimensional data representation of the dataset, the visual representation including visualizations of the dataset as one or more layers of proteins, protein-like molecules, or fragments thereof, each layer including one or more clusters of proteins, protein-like molecules, or fragments thereof; providing tools to the visual map interface for interacting with the map, wherein the interaction with the map includes one or more of the following: inspection, searching, sampling, clustering, and analysis of one or more proteins, protein-like molecules, or fragments thereof, or newly generated proteins, protein-like molecules, or fragments thereof; receiving commands or detecting interactions with the map at the visual map interface via tools; and triggering updates to the visual map interface and the map based on commands or interactions.
[0154] According to one aspect, a computer-implemented system is provided for mapping molecules onto an interface and generating a map for the interface. The system includes a processing subsystem comprising one or more processors and one or more memories coupled to the processors. The processing subsystem is configured such that the system: receives data of one or more sets of one or more input molecules, wherein the input molecules are one or more sets of one or more molecules, and wherein the data includes features of the molecules; generates at least one dataset by processing the data of one or more sets of one or more input molecules, wherein the dataset includes an encoded sequence of information, coordinates for each molecule to define its molecular structure, and features; transforms the data and one or more of the at least one dataset by feature extraction or feature selection to generate a map and additional features, thereby reducing a higher-dimensional data representation to a lower-dimensional data representation that can be indicated, depicted, or visualized by the interface, while capturing valuable information in the lower-dimensional data representation, which includes one or more clusters of the input molecules or newly generated molecules; generates a map user interface, the map user interface including a map as a representation of the lower-dimensional data representation of the dataset, representing the at least one dataset as one or more molecular layers, each layer including one or more of one or more clusters of molecules; and provides the map user interface.
[0155] In some embodiments, the processing subsystem enables the system to: provide the map user interface with tools for interacting with the map and for inspecting, searching, sampling, clustering, and analyzing one or more input molecules or newly generated molecules; receive commands or detect interactions with the map at the visual map interface via the tools; update the map based on the commands or interactions; and trigger an update of the map interface using the updated map.
[0156] In some embodiments, the molecule is a protein, a protein-like molecule, a fragment thereof, a small molecule drug, or a nucleic acid molecule.
[0157] In some embodiments, the input protein, protein-like molecule or fragment thereof includes an antibody or fragment thereof and / or an antigen or fragment thereof, and wherein the map is a complementary site map or an epitope map or a map including a protein or protein-like molecule or fragment thereof.
[0158] In some embodiments, proteins, protein-like molecules or fragments thereof are selected from the group consisting of antibodies, antigen-binding fragments, drug candidates, compounds, candidate conjugates and binding agents.
[0159] In some embodiments, the map user interface includes a visual interface, and the lower-dimensional data representation can be visualized in the visual interface.
[0160] In some embodiments, the input molecule is a protein or protein-like molecule comprising an antibody or its antigen-binding fragment and / or an antigen.
[0161] In some embodiments, the antibody or its antigen-binding fragment includes antibodies derived from a species, including but not limited to mice, cattle, rabbits, camels, llamas, humans, alpacas, and standard species.
[0162] In some embodiments, the input molecule is an antibody or its antigen-binding fragment and / or antigen, and the molecular map is a complementary site map.
[0163] In some embodiments, the input molecule is an antigen, and the molecular map is an epitope map.
[0164] In some embodiments, the features include fingerprints, wherein the processing subsystem uses machine learning or statistical methods involving one or more of dimensionality reduction, reconstructive autoencoders, variational autoencoders, adversarial autoencoders, neural networks, graph neural networks, attention networks, recurrent networks, and generative models to extract features and generate fingerprints.
[0165] In some embodiments, the system further includes a data storage device for a database of molecules, wherein each molecule is assigned a unique index.
[0166] In some embodiments, the system compares a molecule with a database of molecules and assigns the molecule an index to the closest molecule in the database, wherein molecules assigned the same index have similar sequences, structures, or properties.
[0167] In some embodiments, feature extraction includes layer-by-layer embedding, wherein features of each dataset in the plurality of datasets or features of each part of the plurality of datasets are extracted individually for a plurality of datasets or a plurality of parts of a dataset.
[0168] In some embodiments, features can be drawn as overlapping visual layers.
[0169] In some embodiments, the map user interface that includes a visualization of the dataset has control inputs that allow layers to be viewed individually or in relation to other layers.
[0170] In some embodiments, one or more layers are used for training feature extraction, and one or more other layers are used for testing feature extraction.
[0171] In some embodiments, feature extraction includes arranging embeddings in clusters.
[0172] In some embodiments, feature extraction includes embedding individual molecules.
[0173] In some embodiments, feature extraction includes generating clusters around molecules of interest.
[0174] In some embodiments, feature extraction includes sampling from a molecular map.
[0175] In some embodiments, feature extraction includes encoding scores in a molecular map.
[0176] In some embodiments, the map user interface includes visualization of extracted features of the dataset.
[0177] In some embodiments, a user device for displaying a map user interface is also included.
[0178] In some embodiments, the map user interface is one-dimensional, two-dimensional, three-dimensional, four-dimensional, or higher-dimensional.
[0179] In some embodiments, the layer-by-layer embedding is a time series that varies over time, and the map user interface includes a two-dimensional or three-dimensional embedding that varies over time as a time series representing a four-dimensional embedding.
[0180] In some embodiments, the map user interface includes visualization of clusters around molecules of interest.
[0181] In some embodiments, the map user interface includes one or more clusters of molecules.
[0182] In some embodiments, the processing subsystem performs feature extraction on each molecule in the dataset to obtain individual molecule embeddings.
[0183] In some embodiments, layer-by-layer embedding includes individual molecular embedding.
[0184] In some embodiments, the user interface uses extracted features to characterize and / or obtain information about the input molecule, wherein the input molecule may optionally come from the cluster of interest.
[0185] In some embodiments, the information includes the extent, nature, and / or robustness of the immune response (against an antigen such as an immunogen or vaccine).
[0186] In some embodiments, the input molecule includes an antibody or an antigen-binding fragment thereof, and the information includes an amino acid sequence of one or more of the antibody or antigen-binding fragment.
[0187] In some embodiments, the input molecule includes an antigen, and the information includes the amino acid sequence of one or more of the antigens.
[0188] In some embodiments, the system allows a user to modify the amino acid sequence, and the system predicts the effect of the modification on the characteristics of the molecule (binding, function, stability, expressibility, affinity, immunogenicity).
[0189] In some embodiments, modifications are made to amino acid substitutions, deletions, and / or additions in one or more CDRs, variable regions, frame regions, and / or constant regions of the antibody or its antigen-binding fragment.
[0190] In some embodiments, the modification is the humanization, deimmunization, glycosylation, or deglycosylation of the antibody or its antigen-binding fragment.
[0191] In some embodiments, the system allows users to import other molecules and determine their similarity to the input molecule.
[0192] In some embodiments, the other molecules are other antibodies or their antigen-binding fragments, and the similarity is complementary site similarity.
[0193] In some embodiments, the output includes antibodies or antigen-binding fragments thereof selected from the map or variants thereof.
[0194] In some embodiments, the user synthesizes an input molecule or an output molecule or a variant thereof, or causes an input molecule or an output molecule or a variant thereof to be synthesized.
[0195] In some embodiments, information provided by the system is used to manufacture molecules.
[0196] In some embodiments, the input molecule includes a single-domain antibody or an antigen-binding fragment thereof, and the output molecule includes an antibody or an antigen-binding fragment thereof selected from conventional antibodies, single-domain antibodies, single-chain variable fragments, humanized antibodies, or chimeric antibodies.
[0197] In some embodiments, the processing subsystem connects the features of the parts of the molecule together to have the overall features of the molecule.
[0198] In some embodiments, the encoded information sequence refers to the encoded amino acid sequence.
[0199] In some embodiments, the processing subsystem outputs the selected molecule.
[0200] According to one aspect, an article is obtained from the selected molecular output by a computer-implemented system described herein.
[0201] According to one aspect, a product obtained by a system implemented by a computer as described herein is provided.
[0202] According to one aspect, a computer-implemented method is provided for mapping molecules onto an interface and generating a map for the interface. The method includes: receiving data of one or more sets of one or more input molecules, wherein the input molecules are one or more sets of one or more molecules, and wherein the data includes features of the molecules; generating at least one dataset by processing the data of one or more sets of one or more input molecules, wherein the dataset includes an encoded sequence of information, coordinates for each molecule to define its molecular structure, features, and fingerprints; transforming the data and one or more of the at least one dataset by feature extraction, feature selection, or fingerprint generation to generate a map and additional features, thereby reducing a higher-dimensional data representation to a lower-dimensional data representation that can be indicated, depicted, or visualized by the interface, while capturing valuable information in the lower-dimensional data representation, which includes one or more clusters of the input molecules or newly generated molecules; generating a map user interface, the map user interface including a map as a representation of the lower-dimensional data representation of the dataset, representing the at least one dataset as one or more molecular layers, each layer including one or more of one or more clusters of molecules; and providing the map user interface.
[0203] In some embodiments, the map user interface includes a visual interface, and the lower-dimensional data representation can be visualized in the visual interface.
[0204] In some embodiments, the input molecule is a protein or protein-like molecule comprising an antibody or its antigen-binding fragment and / or an antigen.
[0205] In some embodiments, the antibody or its antigen-binding fragment includes antibodies derived from a species, including but not limited to mice, cattle, rabbits, camels, llamas, humans, alpacas, and standard species.
[0206] In some embodiments, the input molecule is an antibody or its antigen-binding fragment and / or antigen, and the molecular map is a complementary site map.
[0207] In some embodiments, the input molecule is an antigen, and the molecular map is an epitope map.
[0208] In some embodiments, the features include fingerprints, wherein the processing subsystem uses machine learning or statistical methods involving one or more of dimensionality reduction, reconstructive autoencoders, variational autoencoders, adversarial autoencoders, neural networks, graph neural networks, attention networks, recurrent networks, and generative models to extract features and generate fingerprints.
[0209] In some embodiments, the method further includes storing a database of molecules in a data storage device, wherein each molecule is assigned a unique index.
[0210] In some embodiments, the method further includes comparing the molecule with a database of molecules and assigning the molecule an index to the closest molecule in the database of molecules, wherein molecules assigned the same index have similar sequences, structures, or properties.
[0211] In some embodiments, the method further includes feature extraction with layer-by-layer embedding, wherein features of each dataset in the plurality of datasets or features of each part of the plurality of datasets are extracted individually for the plurality of datasets or the plurality of parts of the datasets.
[0212] In some embodiments, features can be drawn as overlapping visual layers.
[0213] In some embodiments, the method further includes providing a map user interface that visualizes the dataset and has control inputs to enable viewing of layers individually or in relation to other layers.
[0214] In some embodiments, one or more layers are used for training feature extraction, and one or more other layers are used for testing feature extraction.
[0215] In some embodiments, feature extraction includes arranging embeddings in clusters.
[0216] In some embodiments, feature extraction includes embedding individual molecules.
[0217] In some embodiments, feature extraction includes generating clusters around molecules of interest.
[0218] In some embodiments, feature extraction includes sampling from a molecular map.
[0219] In some embodiments, feature extraction includes encoding scores in a molecular map.
[0220] In some embodiments, the map user interface includes visualization of extracted features of the dataset.
[0221] In some embodiments, the method further includes using a user device to display a map user interface.
[0222] In some embodiments, the map user interface is one-dimensional, two-dimensional, three-dimensional, four-dimensional, or higher-dimensional.
[0223] In some embodiments, the layer-by-layer embedding is a time series that varies over time, and the map user interface includes a two-dimensional or three-dimensional embedding that varies over time as a time series representing a four-dimensional embedding.
[0224] In some embodiments, the method further includes using a map user interface to provide visualization of clusters around molecules of interest.
[0225] In some embodiments, the map user interface includes one or more clusters of molecules.
[0226] In some embodiments, the method further includes performing feature extraction on individual molecules in the dataset to obtain individual molecule embeddings.
[0227] In some embodiments, layer-by-layer embedding includes individual molecular embedding.
[0228] In some embodiments, the method further includes using extracted features to characterize and / or obtain information about input molecules, wherein the input molecules may optionally come from clusters of interest.
[0229] In some embodiments, the information includes the extent, nature, and / or robustness of the immune response (against an antigen such as an immunogen or vaccine).
[0230] In some embodiments, the input molecule includes an antibody or an antigen-binding fragment thereof, and the information includes an amino acid sequence of one or more of the antibody or antigen-binding fragment.
[0231] In some embodiments, the input molecule includes an antigen, and the information includes the amino acid sequence of one or more of the antigens.
[0232] In some embodiments, the method further includes modifying the amino acid sequence, and wherein the system predicts the effect of the modification on the characteristics of the molecule (binding, function, stability, expressibility, affinity, immunogenicity, etc.).
[0233] In some embodiments, modifications are made to amino acid substitutions, deletions, and / or additions in one or more CDRs, variable regions, frame regions, and / or constant regions of the antibody or its antigen-binding fragment.
[0234] In some embodiments, the modification is the humanization, deimmunization, glycosylation, or deglycosylation of the antibody or its antigen-binding fragment.
[0235] In some embodiments, the system allows users to import other molecules and determine their similarity to the input molecule.
[0236] In some embodiments, the other molecules are other antibodies or their antigen-binding fragments, and the similarity is complementary site similarity.
[0237] In some embodiments, the output includes antibodies or antigen-binding fragments thereof selected from the map or variants thereof.
[0238] In some embodiments, the user synthesizes an input molecule or an output molecule or a variant thereof, or causes the input molecule or an output molecule or a variant thereof to be synthesized.
[0239] In some embodiments, the method further includes using information provided by the system to manufacture molecules.
[0240] In some embodiments, the input molecule includes a single-domain antibody or an antigen-binding fragment thereof, and the output molecule includes an antibody or an antigen-binding fragment thereof selected from conventional antibodies, single-domain antibodies, single-chain variable fragments, humanized antibodies, or chimeric antibodies.
[0241] In some embodiments, the method further includes linking features of the portions of the molecule together to have the overall features of the molecule.
[0242] In some embodiments, the encoded information sequence refers to the encoded amino acid sequence.
[0243] In some embodiments, the method further includes outputting the selected molecule.
[0244] According to one aspect, an article is obtained by a computer-implemented method described herein, which yields a selected molecular output.
[0245] According to one aspect, a product obtained by a computer-implemented method described herein is provided.
[0246] In some embodiments, the method includes the step of generating molecules identified or selected from a map user interface.
[0247] According to one aspect, a product obtained by a computer-implemented method described herein is provided.
[0248] In some embodiments, the product is an antibody or an antigen-binding fragment thereof.
[0249] According to one aspect, one or more non-transitory computer-readable media having stored thereon machine-interpretable instructions, which, when executed by a processing subsystem, cause the processing subsystem to perform a method for mapping molecules onto a visual interface and generating a map for the interface. The method includes: receiving data of one or more sets of one or more input molecules, wherein the input molecules are one or more sets of one or more molecules, and wherein the data includes features of the molecules; generating at least one dataset by processing the data of one or more sets of one or more input molecules, wherein the dataset includes an encoded sequence of information, coordinates for each molecule to define its molecular structure, features, and fingerprints; transforming the data and one or more of the at least one dataset by feature extraction, feature selection, or fingerprint generation to generate a map and additional features, thereby reducing a higher-dimensional data representation to a lower-dimensional data representation that can be indicated by the interface, while capturing valuable information in the lower-dimensional data representation, which includes one or more clusters of the input molecules or newly generated molecules; generating a map user interface, the map user interface including a map as a representation of the lower-dimensional data representation of the dataset, representing the at least one dataset as one or more molecular layers, each layer including one or more of one or more clusters of molecules; and providing the map user interface.
[0250] According to one aspect, a computer-implemented system for a molecule-related interface is provided. The system includes: a processing subsystem comprising one or more processors and one or more memories coupled to the processors, the processing subsystem being configured such that the system: receives input molecules, wherein the input molecules are a set of one or more sets of one or more molecules, wherein each molecule is defined as an information sequence; encodes the information sequence; generates a dataset by processing the input molecules, wherein the dataset includes the encoded information sequence, three-dimensional coordinates of the information sequence for each molecule defining the molecular structure, features, and fingerprints; generates a transformed dataset by feature extraction and fingerprinting, thereby reducing a high number of dimensions to a lower-dimensional data representation that can be indicated by the interface, while capturing valuable data in the lower-dimensional data representation; generates one or more metrics from the transformed dataset, wherein the one or more metrics include the lower-dimensional data representation of the dataset and summarize the characteristics of the input molecules; and provides one or more metrics to the interface.
[0251] According to one aspect, a computer-implemented system for a visual interface for mapping molecules is provided. The system includes a processing subsystem comprising one or more processors and one or more memories coupled to the processors. The processing subsystem provides a map user interface, wherein the map user interface: receives input molecules, wherein the input molecules are one or more groups of one or more molecules, wherein each molecule is defined as an information sequence; and provides a map interface comprising a molecular map as a representation of a lower-dimensional data representation of the dataset of input molecules, wherein the dataset includes the information sequence of each molecule, three-dimensional coordinates of the information sequence for defining the molecular structure of each molecule, and features; wherein the molecular map includes a transformed dataset generated through feature extraction and fingerprinting, thereby reducing the higher number of dimensions to a lower-dimensional data representation that can be indicated by the interface, while capturing valuable data in the lower-dimensional data representation.
[0252] In some embodiments, the input molecule is a protein or protein-like molecule that includes antibodies and antigens.
[0253] In some embodiments, antibodies include species-specific antibodies, including but not limited to mice, cattle, rabbits, camels, llamas, humans, alpacas, and standard species.
[0254] In some embodiments, the input molecules are antibodies and antigens, and the molecular map is a complementary site map.
[0255] In some embodiments, the input molecule is an antigen, and the molecular map is an epitope map.
[0256] In some embodiments, the map user interface includes layer-by-layer embedding to provide layers for map visualization.
[0257] In some embodiments, the map user interface renders features as overlaid map visualization layers.
[0258] In some embodiments, the map user interface has control inputs that enable viewing layers individually or in relation to other layers.
[0259] In some embodiments, the map user interface has control inputs to add or remove layers in the layers used for map visualization.
[0260] In some embodiments, the map user interface includes visualization of extracted features of the dataset.
[0261] In some embodiments, the map user interface receives one or more reference molecules or target molecules, wherein the transformation of the dataset is based on one or more reference molecules or target molecules.
[0262] In some embodiments, the map user interface receives one or more scores for the molecule, wherein the scores include expressibility scores and fuzzy filtering scores.
[0263] In some embodiments, map visualization includes one or more clusters corresponding to molecules, wherein the map user interface receives clustering control commands to update the map visualization using hyperparameters of one or more clusters.
[0264] In some embodiments, the map visualization displays one or more scores associated with the molecule, including expressibility scores or fuzzy screening scores.
[0265] In some embodiments, the map user interface receives control commands for sampling from at least a portion of the map visualization.
[0266] In some embodiments, the map user interface receives control commands for editing samples extracted from at least a portion of the map visualization.
[0267] In some embodiments, the map user interface receives drawing settings corresponding to the visualization characteristics of the map visualization.
[0268] In some embodiments, the system also includes a user device for displaying a map user interface.
[0269] In some embodiments, the map user interface is one-dimensional, two-dimensional, three-dimensional, four-dimensional, or higher-dimensional.
[0270] In some embodiments, the layer-by-layer embedding is a time series that varies over time, and the map user interface includes a two-dimensional or three-dimensional embedding that varies over time as a time series representing a four-dimensional embedding.
[0271] In some embodiments, the map user interface includes visualization of clusters around molecules of interest.
[0272] In some embodiments, the map user interface includes one or more clusters of molecules.
[0273] In some embodiments, the map user interface includes individual molecule embeddings.
[0274] In some embodiments, layer-by-layer embedding includes individual molecular embedding.
[0275] In some embodiments, the map user interface receives scores.
[0276] In some embodiments, the map user interface exports a file.
[0277] In some embodiments, the map user interface includes one or more buttons for adding or removing layers, one or more buttons for receiving input molecules, one or more buttons for adding or removing individual molecules, and one or more buttons for importing fractions.
[0278] In some embodiments, the map user interface includes multiple settings selected from the group consisting of: navigation settings for map visualization, drawing settings, settings for encoding scores in map visualization, clustering settings, settings for editing samples, sample settings, reporting settings, map analysis settings, and export settings.
[0279] In some embodiments, the processing subsystem causes the system to selectively subtract non-immunized molecular libraries from immunized molecular libraries to filter out non-specific molecules and reduce the search space for sampling and searching for specific candidate drug molecules for one or more targets; if the immunized library contains multiple layers or datasets, the subsystem causes the system to selectively intersect the layers after subtraction to further reduce the search space.
[0280] In some embodiments, the processing subsystem causes the system to selectively subtract a molecular library immunized against one or more targets from a molecular library immunized against the target of interest to filter out molecules that do not bind to the target of interest and reduce the search space used for sampling and searching for specific candidate drug molecules against the target of interest; if the molecular library immunized against the target of interest contains multiple layers or datasets, the subsystem causes the system to selectively intersect the layers after the subtraction to further reduce the search space.
[0281] Upon reading this disclosure, many further features and combinations thereof relating to the embodiments described herein will become clear to those skilled in the art. Attached Figure Description
[0282] In the attached diagram, Figure 1 This is a block diagram of an example mapping system according to an embodiment.
[0283] Figure 2 This is a block diagram of an example molecular map according to an embodiment.
[0284] Figure 3 This is a flowchart of an example user interface according to an embodiment.
[0285] Figure 4 This is an example user interface according to an embodiment for users to import / export files, visualize maps, and select settings.
[0286] Figure 4A , Figure 4B , Figure 4C , Figure 4D , Figure 4E This is an example map visualization for a user interface according to an embodiment.
[0287] Figure 5 This is a block diagram of another example mapping system according to an embodiment.
[0288] Figure 6 This is a flowchart of an example method for mapping molecules to a visual interface according to an embodiment.
[0289] Figure 7 This is a schematic diagram of a computing device according to an embodiment.
[0290] Figure 8 This is a flowchart of another example method for mapping molecules to a visual interface according to an embodiment.
[0291] Figure 9 This is a schematic diagram illustrating an example representation of antibodies in a complementary site map.
[0292] Figure 10 This is an example schematic diagram of data processing, showing how the features of the various parts of an antibody are linked together to form the overall characteristics of the antibody.
[0293] Figure 11 This is a diagram of an example map user interface according to an embodiment.
[0294] Figure 12 This is a diagram of another example map user interface for a sample according to an embodiment.
[0295] Figure 13 This is a diagram of another example map user interface for selecting candidate antibodies according to an embodiment.
[0296] Figure 14 It is a flowchart used to predict antibody-antigen interactions.
[0297] Figure 15 This is a schematic diagram illustrating an example of antibody fingerprinting.
[0298] Figure 16 This is an example schematic diagram illustrating the generation of complementary bit fingerprints used to calculate metrics.
[0299] Figure 17 An example mosaic tile map according to some embodiments is shown.
[0300] Figure 18 Another example mosaic tile map using fractional color encoding is shown according to some embodiments.
[0301] Figure 19An example intersection of two layers according to some embodiments is shown.
[0302] Figure 20 An example of subtracting one layer from another is shown according to some embodiments.
[0303] Figure 21 Examples of diverse sampling from a map interface are shown according to some embodiments.
[0304] Figures 22A-22G Example 2D probability distributions for sampling are shown, generated using different probability distributions on the interface according to some embodiments.
[0305] Figure 23 The illustration depicts an exploration of exemplary immune responses generated in mice with the same genetic background but immunized with different antigenic forms, according to some embodiments.
[0306] These accompanying drawings depict exemplary embodiments for illustrative purposes, and various changes, alternative configurations, alternative components, and modifications may be made to these exemplary embodiments. Detailed Implementation
[0307] The systems and methods of this disclosure can provide one or more maps. For example, the systems and methods of this disclosure can provide a visual map representing antibodies in an antibody library, as well as an interface including tools for interacting with the map to examine, search, sample, cluster, and analyze antibodies. Advantageously, the map can represent, for example, substantially all and / or each antibody in an antibody library (e.g., sequenced via next-generation sequencing (NGS) or other sequencing platforms). In some embodiments, the systems and methods of this disclosure can also compare different antibody libraries obtained, for example, against different targets or different immunization processes (through layer subtraction and / or map subtraction, layer overlay and / or map overlay) to characterize all and / or each antibody (e.g., for their specificity, representativeness, or other parameters determined by the user).
[0308] Advantageously, the systems and methods of this disclosure can be used to identify antibodies with desired properties. In some embodiments, the selection of an antibody against a given target can be sequence-driven, and in some embodiments, the selection can be driven by its structure or properties or functions rather than its sequence.
[0309] In some embodiments, data on regulatory-approved antibodies can be integrated into the system of this disclosure for comparative purposes. Therefore, in some embodiments, the systems and methods of this disclosure can also be used to discover alternatives to regulatory-approved antibodies (e.g., biosimilars, subsequent biologics, generic biologics).
[0310] Furthermore, the systems and methods disclosed herein can also be used to identify antibody mimics or artificial antibodies having properties similar to those of a given antibody. In some embodiments, alternative antibodies, antibody mimics, or artificial antibodies may be suggested or designed by the system, or may be derived from an antibody library.
[0311] In some embodiments, the systems and methods disclosed herein may also be used to determine the extent, nature, and / or robustness of an immune response against one or more antigens.
[0312] In some embodiments, a protein or protein-like molecule may encompass an antigen or fragment thereof. Accordingly, the systems and methods of this disclosure can be performed to characterize epitopes or domains of an antigen. For example, the systems and methods of this disclosure can be used to identify immunodominant epitopes, hidden epitopes, dimerized or multimerized domains, active sites, etc. Advantageously, the systems and methods of this disclosure can be used to identify epitopes or domains for which no antibody is available, to provide novel antibody candidates for purposes such as therapy or diagnosis.
[0313] In some embodiments, the systems and methods of this disclosure can be used to characterize interactions between biomolecules (e.g., protein-protein interactions, protein-DNA interactions, protein-RNA interactions, etc.). In some embodiments, the systems and methods of this disclosure can be used to characterize interactions between proteins, receptors / ligands, etc., in signaling pathways or cascades. Therefore, in some embodiments, protein mapping can provide in-depth understanding of the structure, function, and interactions of proteins, which may be crucial for understanding biological processes and disease mechanisms.
[0314] Other exemplary embodiments of the molecules may cover small molecules (e.g., small molecule drugs). Accordingly, the systems and methods of this disclosure can be used to discover small molecule drugs that interact with one or more targets of interest (e.g., from de novo synthesis, from in vivo or in vivo analysis, from existing libraries, or generated by artificial intelligence).
[0315] Other exemplary embodiments of molecules include nucleic acid molecules (RNA (mRNA), DNA, etc.).
[0316] In some embodiments, the fragment can play a role in drug interactions, for example, by influencing how drugs interact with their biological targets and how they are processed in vivo. For example, the fragment may contain functional groups that interact with specific sites on biological targets, such as proteins or enzymes. Exemplary embodiments of the fragment include, but are not limited to, antigen-binding fragments, complementary sites, epitopes, domains, peptides, etc.
[0317] In some embodiments, the input molecule is or includes proteins, protein-like molecules or fragments thereof, small molecules or biomolecules such as nucleic acids, complex sugars, lipids, or combinations thereof. Exemplary molecules include proteins, protein-like molecules or fragments thereof.
[0318] In some embodiments, the protein or protein-like molecule encompasses the antibody or a fragment thereof. Accordingly, the systems and methods of this disclosure can be performed to identify antibodies that interact with one or more targets of interest (e.g., from immunized animals or humans, from existing libraries, or generated by artificial intelligence). Therefore, the systems and methods of this disclosure may be useful in identifying therapeutically active antibody candidates.
[0319] Exemplary embodiments of antibody fragments include, but are not limited to, antigen-binding fragments or antigen-binding domains (e.g., complementary sites, one or more CDRs and / or frame regions, variable regions, as described in more detail herein).
[0320] Exemplary embodiments of antigen fragments include, but are not limited to, domains, active sites, epitopes (linear or nonlinear), etc.
[0321] Newly generated data integrated into the system can be part of a main interface that covers information about privately or publicly available molecules (e.g., antibodies), including but not limited to regulatory-approved molecules (e.g., regulatory-approved antibodies), molecules involved in preclinical and clinical trials, and previously generated data. Newly generated data may include, for example, laboratory data, stored data, and data generated using generative AI models (e.g., RFDiffusion; a generative model for proteins).
[0322] Users can query the system and compare the data with the data on the main interface by intersection, subtraction and / or overlay.
[0323] Figure 1 An example pipeline for mapping molecules to a visual interface is shown. The data 102 provided as input to system 100 can be information about a molecule, a group of molecules, or multiple groups of any molecules. In some embodiments, a reference molecule and a target molecule are provided as inputs to 102. The reference molecule and target molecule form the basis of the map. System 100 can involve different types of input molecules. For example, input molecules can be proteins, or any class of protein molecules, including but not limited to antibodies, antigens, lectins, and receptors.
[0324] The term "antibody" encompasses a wide range of antibody forms and structures, including any immunoglobulin, monoclonal antibody, polyclonal antibody, bivalent antibody, monovalent antibody, bispecific antibody, multispecific (multispecific) antibody, conventional or natural antibody (e.g., composed of light and heavy chains), single-domain antibody, single-chain antibody (e.g., single-chain variable fragment, scFv), heavy-chain-only antibody, nanobody, humanized antibody, chimeric antibody, artificial antibody, antibody mimic, any antigen-binding fragment exhibiting desired antigen-binding activity, and any variant of the antibody. Antibodies can be produced after immunization of animals (including transgenic animals), such as any animal species, including but not limited to mice, cows (cattle), rabbits, camels, llamas, humans, and alpacas. Antibodies can be recombinantly produced or chemically synthesized.
[0325] In the context of antibodies, "binding" refers to the interaction between an antibody and an antigen. The binding of an antibody to its target is preferably specific in order to avoid off-target side effects.
[0326] In some embodiments, chimeric antibodies or antigen-binding fragments thereof encompass, but are not limited to, antibodies or antigen-binding fragments thereof comprising a variable region from one species and a constant region from another species. Typically, chimeric antibodies or antigen-binding fragments thereof may comprise a variable region (or a portion thereof) derived from a non-human antibody (e.g., mouse, rat, rabbit, hamster, etc.) and a constant region or a portion thereof from a human antibody (e.g., human Fc, human CH3, or human CH2-CH3 domain).
[0327] In some embodiments, humanized antibodies or antigen-binding fragments thereof encompass, but are not limited to, antibodies or antigen-binding fragments in which amino acid residues are replaced to increase sequence similarity or sequence identity with human antibodies or human antibody common sequences (e.g., germline templates). Typically, non-human antibodies are humanized by replacing one or more amino acid residues in one or more frame regions with corresponding amino acid residues of the most similar or most identical human antibody or human antibody common sequence. An antibody or antigen-binding fragment may be considered fully humanized if the amino acid residues in the frame region are 100% identical to those of the human antibody or human antibody common sequence; or it may be considered partially humanized if the amino acid residues in the frame region are less than 100% identical to those of the human antibody or human antibody common sequence. Humanized antibodies or antigen-binding fragments thereof may also include CDR amino acid residues of human antibodies. Typically, humanized antibodies or antigen-binding fragments thereof (fully or partially humanized) may also include constant regions of human antibodies or portions thereof (e.g., human Fc, human CH3, or human CH2-CH3 domains).
[0328] Exemplary embodiments of artificial antibodies include antibodies generated in computer simulations or generated through machine learning-based methods such as generative artificial intelligence.
[0329] Antibody mimics are generally derived, for example, from non-antibody backbone proteins. Exemplary examples of antibody mimics include, but are not limited to, affibody, adnectin, affilin, affimer, affitin, alphabet, anticalin, aptamer, armadillo repeat protein, atrimer, avimer, designed ankylosing repeat protein (DARPin), fynomer, Kunitz domain peptide, knot peptide, etc. (see Yu, X. et al., Annu Rev Anal Chem, 10(1): 293-320, 2017, the entire contents of which are incorporated herein by reference). Antibody mimics encompass any type of molecule capable of interacting with or binding to antigens. Antibody mimics can specifically interact with antigens through complementary shapes (complementary sites) to antigenic epitopes.
[0330] Other exemplary antibody formats and structures include, but are not limited to, single-chain Fv-CH3 (scFv-CH3) fusions, tandem scFv-CH3 (TaFv-CH3) fusions, dual antibody-CH3 (Db-CH3) fusions, tandem Db-CH3 (TaDb-CH3) fusions, single-chain Db-CH3 fusions (scDb-CH3), Fab-CH3 fusions, single-chain Fab-CH3 fusions, Fab-scFv-CH3 fusions, and biaffinity-based targeted (DART)-CH3 fusions. Fab-DART-CH3 fusion, single-chain Fv-Fc (scFv-Fc) fusion, tandem scFv-Fc (TaFv-Fc) fusion, biantibody-Fc (Db-Fc) fusion, tandem Db-Fc (TaDb-Fc) fusion, single-chain Db-Fc fusion (scDb-Fc), Fab-Fc fusion, single-chain Fab-Fc fusion, Fab-scFv-Fc fusion, biaffinity-based targeted (DART)-Fc fusion, Fab-DART-Fc fusion, etc.
[0331] Examples of antibody-like proteins include, for example, but not limited to, VH-VL, VHH, ScFv (single-chain variable fragment), Fab, HCAb, IgNAR, etc. Input molecules can be antibodies from any species, including but not limited to mice, cows (bovines), rabbits, camels, llamas, humans, alpacas, and standard species. Input molecules can be from any source; for example, they can be natural, synthetic, generated in transgenic animals, generated in computer simulations, etc. In some embodiments, antigens can be used as data, either in conjunction with antibodies or alone. System 100 typically accepts any molecule. If it is used for both antibodies and antigens, the generated maps can be named a complementation site map and an epitope map, respectively. Data 102 can be information about the molecule, including informational sequences (e.g., for antibodies, the informational sequences may include complementation-determining regions), and may also include different properties of the molecule (activity, expressibility, affinity, etc.).
[0332] Exemplary examples of single-domain antibodies include antibodies produced by camelids (dromedary camels, camels, llamas, alpacas, etc.) or sharks. In some cases, "single-domain antibodies" may be produced by transgenic animals modified to express heavy-chain-only antibodies. Exemplary examples of transgenic animals are provided in International Application No. PCT / CA2021 / 050951, filed on July 21, 2021, and published on January 20, 2022, under No. WO2022 / 011457, the entire contents of which are incorporated herein by reference.
[0333] Exemplary embodiments of one or more antigen-binding fragments include fragments of an antibody containing an antigen-binding domain, and the fragments may incorporate or not incorporate other portions of the antibody, such as, for example, amino acid residues in the hinge region, amino acid residues in the constant region, or a portion of the Fc region. Regardless of structure, antigen-binding fragments typically bind to the same antigen recognized by the intact antibody.
[0334] Exemplary embodiments of an antigen-binding domain include a portion of an antibody involved in antigen binding and, for example, include one or more complementarity-determining regions (CDRs), one or more frame regions (FRs), or (one or more) entire variable regions. In the context of a single-domain antibody, an exemplary embodiment of an antigen-binding domain may include a portion of a single-domain antibody involved in antigen binding, such as, for example, one or more CDRs selected from CDRH1, CDRH2, or CDRH3, one or more frame regions FR1, FR2, FR3, FR4, or one or more entire variable regions (VH or VHH). In the context of a natural antibody, the term "antigen-binding domain" refers to a portion of a natural antibody involved in antigen binding and, for example, includes one or more CDRs selected from CDRH1, CDRH2, CDRH3, CDRL1, CDRL2, or CDRL3, one or more light chain or heavy chain frame regions FR1, FR2, FR3, FR4, or one or two entire variable regions (heavy chain variable region (VH) and / or light chain variable region (VL)). Another embodiment of an antigen-binding domain is a complementary site.
[0335] As used herein, the terms “CDRH1,” “CDRH2,” and “CDRH3” refer to CDR1, CDR2, or CDR3 of the antibody heavy chain, respectively. As used herein, the terms “CDRL1,” “CDRL2,” and “CDRL3” refer to CDR1, CDR2, or CDR3 of the antibody light chain, respectively. It should be understood in this document that the position of “CDR” in a conventional antibody can be determined using the Kabat numbering scheme (e.g., Kabat, J Immunol., 147:1709-19 (1991); Chothia C, Lesk AM, J Mol Biol. Aug 20; 196(4):901-17 (1987)), the Chothia numbering scheme, or the IMGT numbering scheme (e.g., Lefranc, M.-P., The Immunologist, 7, 132-136 (1999)).
[0336] In some embodiments, the affinity of an antibody or its antigen-binding fragment can be determined by the strength of the non-covalent interaction between the binder or its antigen-binding domain(s) and the antigen.
[0337] In some embodiments, an epitope encompasses, for example, but not limited to, a specific group of atoms or amino acid residues on an antigen to which an antibody or antigen-binding fragment binds. If two antibodies exhibit competitive binding to an antigen, they may bind to the same or closely related epitopes within the antigen. Epitopes may be linear or conformational (i.e., comprising spaced-apart amino acid residues). For example, if an antibody or antigen-binding fragment blocks the binding of a reference antibody to an antigen by at least 85%, or at least 90%, or at least 95%, then that antibody or antigen-binding fragment may be considered to bind to the same / closely related epitope as the reference antibody. If two antibodies cluster together on a complementary site map, then they may be considered to bind to the same or closely related epitopes.
[0338] In some embodiments, complementary sites encompass, for example, but not limited to, spatial structures arising from specific groups of atoms or amino acid residues on an antibody or antigen-binding fragment that binds to the antigen. Typically, a complementary site corresponds, for example, to the portion of the antibody that binds to the antigen to form an antigen-antibody complex. For instance, the complementary site of a given antibody may be structurally similar to or identical to the complementary site of another antibody without sharing significant amino acid identity.
[0339] As used herein, the term "identity" with respect to sequences refers to the degree of similarity between two or more nucleic acid sequences or two or more amino acid sequences at the time of optimal comparison. Sequence identity can be at least 85%, 90%, or 95%, preferably at least 95%. Non-limiting examples include 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 95%, 96%, 97%, 98%, 99%, and 100%. Generally, identity is determined over the entire length of shorter sequences. However, in some cases, identity may be determined only over a portion of the sequence.
[0340] The term "similarity" in relation to sequences considers the degree of sequence identity and the extent to which amino acid residues are substituted by conserved amino acids. As used herein, the term similarity can refer to different types of similarity measures. Similarity can be in terms of sequence, structure (three-dimensional coordinates), biophysical properties (charge, hydrophobicity, stability (e.g., thermal stability, pH stability, resistance to proteolysis), solubility, aggregation tendency, affinity, specificity, immunogenicity risk, glycosylation, post-translational modifications), expressibility, potency, and / or manufacturing characteristics (e.g., yield). Accordingly, sequence similarity and complementation site similarity are (non-limiting) examples of structural similarity.
[0341] Generally, the Blast2 sequence program (Tatiana A. Tatusova, Thomas L. Madden (1999), "Blast 2 sequences - a new tool for comparing protein and nucleotide sequences", FEMS Microbiol Lett. 174:247-250) is used with the default settings, namely, the blastp program, BLOSUM62 matrix (open vacancy 11 and extended vacancy penalty 1; gapx attenuation 50, expected value 10.0, word length 3) and the filter is activated to determine the degree of sequence similarity and identity between sequences.
[0342] Therefore, variations of this disclosure may include sequences that have at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with a reference sequence or a portion thereof.
[0343] In some embodiments, complementary site similarity can be determined, for example but not limited to, by the similarity or resemblance between the spatial arrangements of atoms in the three-dimensional space of two or more complementary site structures. In an exemplary embodiment, a comparison of protein structures (e.g., complementary site structures) can be performed by superimposing one structure onto another to assess the overlap of atoms. In some embodiments, methods such as root mean square deviation (RMSD) can be used to quantify the structural differences between two protein structures.
[0344] In some embodiments, the structure may include the coordinates of atoms or molecules of a protein, protein-like molecule, or fragment thereof. The structure may also be visualized in a strip or surface manner. The structure may also be a snapshot of a sculpture of a protein or protein-like molecule. The structure may be provided through 3D visualization. The structure may also be provided through exposure (e.g., solvent exposure).
[0345] In some embodiments, the “expressibility” of an antibody or its antigen-binding fragment in the context of an antibody can be determined, for example but not limited to, by the ability of a host organism (such as mammalian cells) to produce and / or secrete antibodies with desired binding and / or functional activity.
[0346] In some embodiments, the stability of an antibody can be determined, for example but not limited to, by means of its structural integrity, functionality, and / or biochemical properties over time and / or its maintenance under various conditions.
[0347] In some embodiments, immunogenicity can be determined, for example but not limited to, by the ability of a substance (such as a protein) to induce an immune response in an organism. In the context of proteins (including therapeutic proteins and biologics), immunogenicity encompasses, for example, the potential of the protein to elicit an immune response in vivo, such as the production of antibodies against the protein (or cells capable of expressing antibodies).
[0348] Exemplary embodiments of protein-like molecules include, but are not limited to, peptide-like molecules and any molecule comprising at least one polypeptide moiety, such as, but not limited to, protein-polymer conjugates, protein-nucleic acid hybrids (aptamers, peptide nucleic acids (PNAs)), and protein-lipid hybrids.
[0349] As used herein, a “fragment” can refer to a portion of a larger molecule. A fragment of a molecule can be continuous or discontinuous. In some embodiments, an epitope (linear or non-linear) can be referred to as a fragment of an antigen. In some embodiments, a complementary site can be referred to as a fragment of an antibody. In some embodiments, fragments can play a role in drug interactions by influencing how a drug interacts with its biological targets and how they are processed in vivo. For example, a fragment can contain functional groups that interact with specific sites on a biological target, such as a protein or enzyme. Exemplary embodiments of fragments include, but are not limited to, antigen-binding fragments, complementary sites, epitopes, domains, peptides, etc. Exemplary embodiments of antibody fragments include, but are not limited to, antigen-binding fragments or antigen-binding domains (e.g., complementary sites, one or more CDRs and / or frame regions, variable regions, as described in more detail herein).
[0350] Input data may include data relating to the sequence, structure, and / or biophysical properties of a molecule. As used herein, a “sequence” refers to a specific order in which amino acids are arranged in a polypeptide or protein. This sequence can be determined by the genetic code and governs the structure and function of the protein. For example, a linear sequence of amino acids linked by peptide bonds in a protein. This sequence is unique to each protein and determines how the protein will fold and function. The structure of a protein or molecule can be described at several levels, including primary structure (e.g., the linear sequence of amino acids), secondary structure (e.g., local folding patterns within a protein, such as α-helices and β-sheets stabilized by hydrogen bonds), tertiary structure (e.g., the overall three-dimensional shape of a single polypeptide chain formed by interactions between the side chains (R groups) of amino acids), and quaternary structure (e.g., the arrangement of multiple polypeptide chains (subunits) in a multi-subunit protein).
[0351] As used herein, “biophysical properties” refers to the physical and / or chemical characteristics of a molecule or protein. Biophysical properties can affect the behavior and interactions of a molecule or protein. Examples include molecular weight, hydrophobicity and hydrophilicity, binding affinity, affinity, specificity, thermal stability, pH stability, resistance to proteolysis, isoelectric point, solubility, aggregation tendency, immunogenicity risk, glycosylation, and / or post-translational modifications.
[0352] As used in this article, a “map” is a spatial representation that links molecules within a space according to one or more models or clusters. Example maps can indicate the spatial distribution of molecules. Specific example map types mentioned in this article include epitope maps, complementation site maps, protein maps, etc.
[0353] As used herein, an "epitope map" can be a detailed representation of a specific region (epitope) on an antigen that is recognized and bound by an antibody. In some embodiments, the map can identify antibody binding sites on the antigen, which can help determine the antibody's mechanism of action and its potential therapeutic uses.
[0354] As used in this article, a “complementary site map” can be a detailed representation of a specific region (complementary site) on an antibody that binds to an antigen.
[0355] As used in this article, a “protein map” can be a comprehensive representation of proteins expressed in a specific organism, tissue, or cell type.
[0356] As used in this article, "clustering" refers to a group of proteins, protein-like molecules, or fragments thereof (and related data elements) that are "similar" to each other within a dataset. Clustering is a technique for organizing data points into meaningful groups based on characteristics.
[0357] The data file 102 can be in different formats. The data 102 can originate from different sources. In some implementations, the data may include sequence and / or structural information obtained from NGS (Next-Generation Sequencing), hybridoma, or single B-cell sorting. For example, a Protein Data Library (PDB) file or FASTA format can be used for proteins, antibodies, and antigens. The PDB format contains structural and sequence information of proteins, while the FASTA format includes the informational sequence of proteins.
[0358] Data file 102 may include molecules represented as information sequences. For example, system 100 may store information sequences representing molecules in a database. In some embodiments, each information sequence of a molecule may be indexed by an identifier. Example identifiers and information sequences are shown below:
[0359] The sequences above are merely examples. These are framework 1 sequences for heavy chain antibodies. These sequences can be used by system 100 to identify antibody sequences during NGS bioinformatics processing (which may be referred to as NGS preprocessing).
[0360] Furthermore, some or all of Data 102 can be in other formats, such as spreadsheet files, data frames, comma-separated values (CSV) format, text files, hexadecimal files, tables, and visualizations of any format. Moreover, biophysical properties of all origins—high-volume data, such as NGS—can be used as part of Data 102. For example, each amino acid or residue in a protein or antibody can have biophysical properties. These properties include, but are not limited to, pH, hydrophobicity and hydrophilicity, negative and positive charge, and solvent exposure, i.e., the degree to which amino acids are present on the surface of the antibody.
[0361] Data 102 can also be obtained from nucleic acid sequencing methods such as PacBio, Illumina, and Nanopore. Several preparation and preprocessing steps, such as stop codons, amber codons, and frameshifts, can also be utilized. During the preprocessing stage, data 102 can also undergo any transformations, such as matrix transformations. For example, NGS preprocessing can be performed. As another example, preprocessing can involve eliminating redundancy in data 102.
[0362] In system 100, dataset 106 can be generated from data 102 through data analysis and preparation 104. The three-dimensional coordinates of the atoms of each molecule can be considered as part of the features of the dataset. For example, if the molecule is a protein, then the three-dimensional coordinates of the atoms of amino acids can be extracted from a PDB file. In some embodiments, the average coordinates of the atoms in each amino acid can be used to have the average three-dimensional coordinates of each amino acid in the protein. The three-dimensional coordinates of the atoms or amino acids contain structural information of the molecule. In some embodiments, system 100 can use a generative model on dataset 106.
[0363] System 100 can generate dataset 106 from data 102 through data analysis and preparation 104. This can involve training data and preprocessing data for training. For example, system 100 can use a precision pipeline to process NGS data. The data can be preprocessed before training, and at step 104, system 100 can prepare data for training. Sequences are obtained by system 100 from the NGS data. For example, system 100 can preprocess the NGS data to obtain a complementary bit map.
[0364] System 100 enables an NGS pipeline for preprocessing data. It aligns paired-end DNA sequencing reads of nucleotide sequences. The nucleotide sequences are analyzed to identify open reading frames encoding proteins and determine the amino acid sequences. The amino acid sequences are then filtered for proteins matching the antibody profile. A list of unique antibody sequences with prevalence data can then be saved.
[0365] Molecules can be viewed as information sequences. For example, if a molecule is a protein, then it is a sequence of amino acids. Information sequences can also be used for dataset generation. Any encoding method, such as one-hot encoding, binary encoding, indexed encoding, or Gray code, can be used to encode information sequences.
[0366] The features extracted from data 102 can be combined to prepare dataset 106. The original dataset 106 may include features. As will be discussed in this paper... Figure 2 As described, dataset 106 also undergoes feature extraction at position 202 to generate a map. Accordingly, features can exist in dataset 106, and feature extraction can also occur during map generation.
[0367] For example, the information encoding sequence, the three-dimensional coordinates of the sequence, and some other properties—such as properties including solvent exposure, hydrophobicity, and charge—can be used as features of dataset 106.
[0368] Figure 9 and Figure 10 This is an example schematic diagram of data processing according to an example embodiment. Figure 9 This is a schematic diagram illustrating an example representation of antibodies in a complementary site map. Figure 10 This is an example diagram illustrating the features of the various parts of an antibody connected together to have the overall characteristics of the antibody. System 100 can implement different data transformations, and system 100 can also reconstruct each part and transformation.
[0369] In some implementations, features of each part of the molecule can be considered. For example, if the molecule is a VHH antibody with multiple parts, namely frame 1, CDR1, frame 2, CDR2, frame 3, CDR3, and frame 4, where each part contains several amino acids of the sequence. System 100 can acquire features of each part individually and then concatenate them. For example, the amino acid encoding, three-dimensional coordinates, solvent exposure, hydrophobicity, and amino acid charge of each part can be used to train and test the reconstructed autoencoder. Features of each part of the antibody can be obtained in the latent space of the reconstructed autoencoder for that part. The features of the parts of the antibody can then be concatenated to have the total features of the antibody. System 100 can concatenate the features to generate a feature vector. The parts have different weights, and system 100 can generate a weighted reconstruction. In order for system 100 to generate a weighted reconstruction, these features are separated. Dataset 106 includes features, and these features are then extracted at 202.
[0370] Another example of an extracted feature can be referred to as a fingerprint, which can be extracted from the original dataset described above. Feature extraction can reduce the dimensionality of dataset 106, which can help generate visualizations efficiently. Feature extraction can involve extracting features that still capture essential information from the original dataset 106. Features can be individual measurable attributes or properties about one or more molecules. Features can also be referred to as parameters or variables and their associated values. For example, numerical features and categorical features can exist. Features can be represented as feature vectors, which are n-dimensional vectors representing the numerical features of molecular attributes or properties. Fingerprint generation is the process of mapping large data items to shorter strings or other values that can uniquely identify the original data. Features and fingerprints can support efficient use of resources because they can, for example, compress large blocks of data to efficiently utilize resources for computation and transmission. In biotechnology, a feature can be referred to as a fingerprint. For example, these terms can be used interchangeably for the example embodiments. For the description herein, a fingerprint is an example of a feature. Original features can be processed using feature extraction to generate fingerprints.
[0371] Various machine learning or statistical methods can be used to extract features from the original dataset and generate fingerprints. Some examples include dimensionality reduction, reconstructive autoencoders, variational autoencoders, adversarial autoencoders, neural networks, graph neural networks, attention networks, recurrent networks, and molecular surface interaction fingerprint generation (MaSIF).
[0372] In some implementations, the characteristics of each part of the molecule can be considered. For example, if the molecule is a VHH antibody, it has multiple parts, namely frame 1, CDR1, frame 2, CDR2, frame 3, CDR3, and frame 4, where each part contains several amino acids of a sequence. The characteristics of each part can be obtained individually and then linked together. For example, the amino acid encoding, three-dimensional coordinates, solvent exposure, hydrophobicity, and amino acid charge of each part can be used to train and test the reconstructed autoencoder. The characteristics of each part of the antibody can be obtained in the latent space of the reconstructed autoencoder of that part. The characteristics of the various parts of the antibody can then be linked together to have the overall characteristics of the antibody. Figure 10 The example process is depicted in the diagram, which is an example diagram showing the features of the parts of an antibody linked together to have the overall features of the antibody. In some embodiments, the method involves linking the features of the parts of a molecule together to have the overall features of the molecule.
[0373] In some embodiments, system 100 may use dimensionality reduction to process the dataset. System 100 may use dimensionality reduction to reduce the dimensionality of the data (e.g., to two dimensions) so that the data can be visualized while still capturing valuable data relationships in the visualization generated by the reduced dimensionality representation.
[0374] Alternatively, an available molecular database can be used, where each molecule is assigned a unique index. Any other molecule can then be compared to this database, and each molecule can be assigned the index of the closest or most similar molecule in the database. In this way, molecules with the same index will have similar sequences, structures, or properties.
[0375] like Figure 1 As shown, in some embodiments, after preparing the dataset 106, it is fed as input to the molecular map 200. In some embodiments, the dataset 106 can be transformed into one or more metrics about molecules. The molecular map 200 transforms the dataset 106 into a map that can be analyzed, examined, and visualized at an interface. This transformation can be any linear or nonlinear transformation. The map can be defined as a transformation of molecules into some feature vectors with a certain dimension, which can also be visualized or analyzed.
[0376] Examples of Molecular Map 200 are in Figure 2The following describes the process. Feature extraction 202 is used to extract features from dataset 106. Feature extraction 202 can also be used to generate fingerprints from dataset 106. Fingerprints are example features. Different methods can be used for feature extraction and fingerprint generation, including but not limited to statistical and machine learning methods. For example, dimensionality reduction or manifold learning algorithms can be used for feature extraction. In some implementations, UMAP, t-SNE, or neural networks can be used for feature extraction. Any dimensionality reduction, feature extraction, or feature selection method can be used to reduce a high number of dimensions to lower-dimensional data, making it visual—while still capturing valuable data in the lower-dimensional representation. For example, valuable information can be embedded by embedding instances that are similar in terms of pattern, label, or feature as close to each other, and instances that are dissimilar in terms of pattern, label, or feature as far apart from each other. If machine learning algorithms are used, then different types of machine learning algorithms can be used—whether unsupervised, supervised, or semi-supervised.
[0377] After feature extraction, the features of the dataset are obtained and can be used for inspection, analysis, and visualization. In some embodiments, this feature extraction can be performed layer by layer and can be referred to as layer-by-layer embedding 202. In layer-by-layer embedding 202, there can be multiple datasets or multiple parts of a dataset, where features are extracted individually for each dataset or each part of a dataset. For example, if the features are to be used for visualization (e.g., visual elements for a user interface), they can be drawn as visualization layers superimposed on each other in a graphical representation. If layer-by-layer embedding 202 is for inspection and analysis, these layers can be analyzed individually or in relation to each other. One or more of these layers can be used for training in the feature extraction algorithm, and other layers can be used in the testing (out-of-sample) phase of the algorithm.
[0378] The extracted features can be represented as a map, which can be analyzed and visualized at a user interface displayed on the device. This map can be of any dimension. For example, it can be one-dimensional, two-dimensional, three-dimensional, four-dimensional, or higher. In some implementations, if the map's embedding varies over time as a time series, it can be a time-varying three-dimensional embedding, representing a four-dimensional embedding. That is, time can provide additional dimensional embeddings. Different combinations of any number of dimensions can also be captured and visualized as a map.
[0379] Embeddings can also be arranged as clusters in a table, as in step 206 of molecular map 200. Tables can be stored and represented in any format, such as data frames, spreadsheet files, CSV files, SQL files, tables, etc.
[0380] In addition to batches of molecules, feature extraction 202 can also be applied to individual molecules to obtain individual molecular embeddings 208. These individual molecules can be considered as additional layers in the layer-by-layer embedding 202. Individual molecular embeddings 208 can be investigated, for example, based on their position in a map compared to other molecules. For example, in drug discovery applications, individual molecules can be individual reference antibodies or potential target antibodies. In this example, the antibody of interest can be considered a reference antibody, and its spatial position in the map can be investigated compared to other antibodies or antigens.
[0381] Clusters around some molecules of interest can be considered and analyzed (210). The size of the cluster around each molecule of interest can be fixed or adjustable and can be determined by the user. In the map, each cluster can be centered on the molecule of interest, or it can contain the molecule of interest without necessarily being centered on it. Clusters can be of any shape, such as spheres, squares, rectangles, hypercubes, hyperspheres, etc. The purpose of clustering around the molecule of interest can be for various reasons. For example, clustering around a molecule can be used to filter and analyze molecules similar to that molecule of interest. Another example is for sampling or comparing molecules close to or near the molecule of interest. In some implementations, the molecule of interest can be a target antibody, where clustering around it is for filtering, analyzing, comparing, and sampling antibodies or antigens that are similar to or complementary to it. Similarity can be about any property, such as sequence, structure, or biophysical properties (solvent exposure, hydrophobicity, charge, etc.). Similarity can be any similarity measure, such as inner product, kernel function, cosine similarity, and negative distance or the inverse of distance. Furthermore, any distance metric can be used to measure dissimilarity or similarity. For example, Euclidean distance, Mahalanobis distance, generalized Mahalanobis distance, and distance metric learning can be used. For instance, System 100 can measure similarity based on the proximity of molecules in a map. This can be in terms of sequence, structure, and physical properties. The term "similarity," as used herein, can refer to different types of similarity measures. Similarity can be in terms of sequence, structure (3D coordinates), and biophysical properties (charge, hydrophobicity). Accordingly, sequence similarity and complementary site similarity are (non-limiting) examples of structural similarity.
[0382] In some embodiments, depending on the embedding of molecules, the map contains one or more clusters or clouds of molecules. Clusters in the map can be analyzed and compared (212). The number of clusters can be determined to analyze the overall structure of the map. Determining the number of clusters is a poorly defined problem, as one might consider a cluster to be two smaller clusters; however, the most natural number of clusters can be determined based on consensus among (most) humans and the visualization of data points in the map. The reference to consensus among humans, as used herein, can mean which points in the map most humans would consider to be clusters. For example, most humans reviewing a visualization with multiple data points might see two clusters, but a minority might see four clusters, another minority might see three clusters, and so on. None of these are incorrect, as there may be many data points on the map, and clusters can be interpreted differently by different humans. The consensus among most humans can be the number of clusters visualized. Any method can be used to determine the number of clusters. In some implementations, different clustering algorithms such as K-means, K-medoids, fuzzy C-means, DBSCAN, hierarchical clustering, etc., can be used to find clusters in the map. Different numbers of clusters can exist. The number of clusters can be an adjustable parameter, and a default value can be set as the initial setting. In some embodiments, a default number of clusters that is generally agreed upon among humans can be to treat large islands as a single cluster. This default value can be set to the adjustable hyperparameter initially, and it can be changed by the user to allow for more or fewer clusters.
[0383] In cluster analysis 212, meta-clustering can also be analyzed, that is, clustering the clusters so that similar clusters are closer together and dissimilar clusters are further apart.
[0384] Indexes can also be used on representative molecules for clustering and cluster analysis.212 A database of molecules can be used for indexing. Molecules that are close to those molecules in the database (e.g., a database storing molecules) within the same cluster can be assigned the same index. The average of all or some antibodies in the same cluster, or any summary statistic, can also be used. Molecules in one or more molecular databases can be indexed. Assuming the number of molecules in the database(s) is n; then the indexes would be 1, 2, ..., n. Each other molecule can then be embedded into the same map as the database. In the embedding space, molecules can be indexed using K-nearest neighbors (e.g., 1-nearest neighbors).
[0385] In some embodiments, system 100 assigns an index to each molecule in the database. System 100 can use this map to map both the database and its antibodies. For example, system 100 can assign an index to an antibody based on which antibody in the database is closest. For example, the database can have n data points indexed as 1...n. If a new antibody is found, system 100 can classify it with an index based on which molecule in the database is the closest.
[0386] In some embodiments, the molecular map 200 is configured to sample (214) molecules from the map. This sampling can come from all different parts of the map, or from a specific part of the map or a cluster. Different sampling techniques can be used, such as simple random sampling, stratified sampling, cluster sampling, bootstrap sampling, etc. Sampling can be with or without replacement. For example, stratified sampling can be performed to sample in proportion to the cluster size. In some implementations, clustering algorithms (such as K-means) can be used with sampling to sample in proportion to the cluster size.
[0387] In map sampling (214), iterative sampling may also be performed, where samples can be drawn from different layers of the layer-by-layer embedding 202. Furthermore, samples can be drawn from specific portions of the map. For example, samples can be drawn within clusters surrounding specific molecules in the map. In some implementations, a reference molecule can be considered in the map, with user-defined clusters considered around it. Molecules in clusters containing the reference molecule can be sampled to obtain similar sampled molecules from the map. The reference molecule is also received as one of the inputs 102, and the reference molecule data undergoes preprocessing and feature extraction methods for use with other data. In some embodiments, system 100 can sample over time. For example, suppose system 100 receives data 102 gradually as online or streaming data (e.g., streaming data, online data, or time data). System 100 can then represent the map as it changes over time by the completion of the data. Furthermore, sampling can be performed at each time step. With further completion of the map, system 100 can better form clusters and perform better sampling. An example use case is that samples are also available in earlier time slots, especially when the rate of streaming data is very low. System 100 can help gain deeper insights into disease progression by sampling antibodies or complementary sites over time. For example, System 100 compares a normal, healthy antibody library with early, intermediate, and late-stage disease. Certain disease targets may emerge during the disease process. Novel complementary sites that appear during the disease's timeline can represent potential diagnostic biomarkers and therapeutic interventions. In vaccine development, the map can indicate the effect of booster doses to understand how antibody production in vivo improves (or does not improve) with consecutive booster vaccinations.
[0388] Sampled molecules can be exported in any format for further analysis. For example, sampled molecules can be exported in PDB, FASTA, CSV, spreadsheet, text, table, data frame, and other formats. In addition to the sequence and structural information of the molecules, other properties and characteristics of the sampled proteins can also be exported; for example, hydrophobicity, charge, and solvent exposure can be exported.
[0389] In map visualization, in some embodiments, system 100 can generate one or more measures about molecules, such as by encoding any fraction of any type as a characteristic of the map visualization (216). For example, one or more fractions can be encoded by marker shape, marker size, color, or color transparency. Fractions can be encoded as one or more characteristics of the visualization. Moreover, multiple fractions can be encoded as one or different characteristics of the visualization in the same map. Fractions can also appear with legend labels, such that the legend on the visualization determines how the fractions appear in the map.
[0390] Fractions can be discrete or continuous. For encoding fractions as some visual characteristic 216, the fractions can be discrete and finite because those visual characteristics may have a limited number of options. In these cases, discrete fractions can be used, or if the fractions are continuous, they can be quantized. For example, to encode fractions as marker shapes, the fractions can be quantized. Any quantization method can be used for quantizing fractions.
[0391] In some embodiments, the scores can be preprocessed using different methods before being encoded 216 in the map. Some example preprocessing methods include histogram equalization, quantization, adjustment, centering, standardization, normalization, Z-score normalization, min-max normalization, transforming scores to fall within a specific range, clipping, saturation, etc. Any transformation can be applied to the scores, and the transformation can be any type of transformation, such as linear, affine, and nonlinear transformations.
[0392] If the molecule is a protein or antibody, then the possible scores that can be encoded in the map (216) could be an expressibility (expression) score. Expressibility scores, or yields, can be measured experimentally in a laboratory setting, or they can be predicted using any machine learning algorithm. For example, a neural network can be trained to regress the yield value for each protein.
[0393] If the molecule is a protein or antibody, another possible score that can be encoded in the map (216) is the screening score. The protein screening process can include multiple rounds, in which the number of target proteins and control proteins is counted in each screening round. In some implementations, two types of control proteins may be present, namely, internal and external control proteins. Therefore, each screening round can contain information on the number of target proteins, internal control proteins, and external control proteins. The ratio of target to internal control and target to external control can be obtained using any formula. Example formulas for the ratio of target to internal control and target to internal control can be:
[0394] TG, IC, and EC represent the counts of target proteins, internal control proteins, and external control proteins, respectively. These scores range from zero to one. For example, these scores can be referred to as screening scores. Generally, if a protein has a large number of target proteins and a small number of internal and / or external control proteins across different screening rounds, then it is a good protein to be further processed. Therefore, in some implementations, screening scores can be used, where a larger screening score is expected across different screening rounds. In each experiment, expected rules can be collected from screening experts. In some implementations, simple ranking can be used, while in others, fuzzy logic can be used, where fuzzy rules can be constructed based on expert rules. Any fuzzy reasoning method and fuzzy reasoning system can be used. For example, the Mamdani or Sugeno (Takagi-Sugeno-Kang) or Tsukamoto fuzzy systems and any T-norm and S-norm can be used. Rules can be based on any information, such as screening scores and screening information across different screening rounds. Antibodies can be ranked based on the output scores of a fuzzy system called fuzzy scores. Proteins ranked higher have better fuzzy scores, which can indicate that they are more expected to be further processed.
[0395] Fuzzy inference can be used for any score, not just for filtering scores. Fuzzy rules and fuzzy inference systems can be constructed using different types of scores and a set of expert rules to obtain fuzzy scores. Fuzzy scores obtained from any type of score and rule can be encoded in a map as any visualization feature.216
[0396] The molecular map 200 can be provided to the map user interface 300 and / or implemented in the map user interface 300. Figure 3The illustration shows an example diagram of a map user interface 300. The map user interface 300 can be any type of interface or platform. For example, it can be a graphical user interface (GUI) for display on a device. The map user interface 300 can be implemented as a standalone visualization, a web-based platform, a computer program, an executable (.exe) file, a portable file, a mobile application, etc. The map user interface 300 can run on any platform, any operating system, or any server. For example, it can run on computers, laptops, personal computers, cellular phones, cloud servers, web servers, web browsers, and various operating systems of computers, cellular phones, and tablets. Figure 7 An example computing device 700 is shown. In some embodiments, the map user interface 300 may be a complementary bitmap for visualizing immune response data. The map user interface 300 may use dimensionality reduction processing to generate visualizations from the dataset 106 to provide improved data visualization for system 100.
[0397] For example, system 100 can use one or more generative models, such as generative machine learning or generative artificial intelligence, to generate interpolated or extrapolated molecules in a map. For instance, a user can select or click on a portion of the map where there are no molecules. Then, if an actual molecule exists in that portion of the map, the system can generate the molecule or its properties. This molecule generation, detached from the map, can have various applications. For example, artificial molecules that mimic real molecules can be generated. Furthermore, generative machine learning can be used to analyze different parts of the map to examine which properties are dominant or more important in that part of the map. In some embodiments, system 100 can, for example, use generative models to perform the processing of transforming map data into protein data.
[0398] Example user interface 400 in Figure 4The user interface 400 is interactive by updating the visualization (e.g., map visualization 310) in response to received control commands. The user interface 400 can have any type of button to provide control commands for updating the map visualization 310. For example, the user interface 400 can have regular buttons, radio buttons, checkboxes, pop-up menus, tables, scan bars, browse buttons, upload / download buttons, submit buttons, run buttons, zoom / move buttons, selection options, visualization windows, etc., to provide different values associated with the control commands, thereby exploring, analyzing, and examining the map visualization 310. Each point in the map can represent a molecule or a group of molecules. In some embodiments, each point in the map is for a molecule. Input controls can be hovered over or clicked on points on the map for selection, interaction, etc. The visualization of each dot representing an individual molecule can be used to visually indicate the relevance of each dot to the distance between clusters. Dots or data elements are "clickable" to enable access to information about the molecule. In some embodiments, there may be associated e-commerce services such that clicking a dot on the map adds an antibody to a shopping cart. This could result in the antibody being manufactured and shipped to the customer, or if an additional test is ordered as an "add-on," the antibody could be shipped to a service provider for testing, and the results sent to the customer.
[0399] In some embodiments, user interface 300 (including example user interface 400) may have buttons for adding and removing layers embedded in 302. An "Add Layer" button and a trash can icon in the embedded 302 of 400 are used for adding and removing layers, respectively. Layers can be moved up and down by the user to change their order. Layers in a later order can be visualized on top of previous layers. For example, up / down arrows in the embedded 302 of example user interface 400 can be used to move layers embedded in 302. Each layer can also be hidden or shown in user interface 400 in response to receiving a control command from, for example, a user. An eye icon in the embedded 302 of 400 can be used to hide or show layers embedded in 302.
[0400] User interface 300 (including, for example, example user interface 400) may have a button for importing file 304. In some embodiments, system 100 can receive data 102 via import file 304, which may be referred to herein as imported data file 102. Imported data file 102 can be in any format, such as PDB, FASTA, text, spreadsheet, CSV, data frame, table, etc. For example, a PDB or FASTA file of proteins can be uploaded to the system. In the background of user interface 300, data preparation 104 is performed by mapping system 100 to prepare dataset 106. System 100 can then prepare molecular map 200.
[0401] User interface 300, exemplified by user interface 400, may have buttons for adding or removing individual molecules 306 of interest (e.g., reference or target molecules). Imported data files 102 for the individual molecules 306 of interest can be in any format, such as PDB, FASTA, text, spreadsheet, CSV, data frame, table, etc. In the background of user interface 300, system 100 can perform data preparation 104 to prepare dataset 106. System 100 can prepare a molecular map 200 of the molecules 306 of interest to be included in the remainder of the map. User interface 300 may have settings and options for reporting by hovering over or clicking on molecules or points in the map. User interface 300 can receive selections about which information should be reported or displayed, for example, by hovering over or clicking on a molecule.
[0402] User interface 300, exemplified by user interface 400, may have a button for importing fraction 308, which is an example metric about a molecule. The fraction can be any type of fraction introduced in step 216 of system 100. Some example fractions are protein expressibility fractions and fuzzy fractions. Fractions can be imported in any format, such as spreadsheets, data frames, tables, CSV files, text files, etc.
[0403] User interface 300, exemplified by interface 400, may have a window for visualization of map 310, as implemented by system 100. Figure 2 As explained in steps 202 and 202, map visualization can be of different dimensions. For example, it can be one-dimensional, two-dimensional, three-dimensional, or higher-dimensional. It can also vary over time to become a time series of visualizations with any number of dimensions. That is, time can provide additional dimensions to the visualization. The visualization window 312 of map 310 can be of any shape, such as a rectangle, square, sphere, etc.
[0404] User interface 300, exemplified by user interface 400, may have buttons for zooming in or out within the visualization window 312. It may also have buttons for moving within the visualization window 312. By clicking the zoom in / out buttons, system 100 can identify or select an area in the map to zoom in / out within that area (e.g., by receiving selection input at user interface 300). User interface 300 can also be used to receive control commands to drag the map, moving it from one part of the map to another for better and more detailed examination of the map and visualization.
[0405] User interface 300, exemplified by interface 400, may have buttons, pop-up menus, and selections for drawing settings 314. For example, it may have a pop-up menu for drawing settings. In some implementations, the pop-up menu may have options for marker shape, marker size, color, color transparency, and legend. By selecting each of these settings, the user can enter / type the value of that setting in a text box provided in the interface. Alternatively, another pop-up menu may be displayed to the user to select one of the possible values for that setting. The user may also have options to select a score imported by system 100 or a calculated score, which is then encoded as any of the selected drawing settings in the pop-up menu. By doing so, the corresponding drawing setting can have the score encoded as a corresponding visualization characteristic. A single score can be used to encode multiple drawing settings, such as marker shape, marker size, color, color transparency, and legend. Multiple scores can also be used to be encoded as different drawing settings and visualization characteristics. One or more scores may be imported by the user (316) or may be calculated by system 100. Some example scores are expressiveness scores and fuzzy filtering scores.
[0406] User interface 300, exemplified by user interface 400, may have settings, buttons, and pop-up menus for selecting clusters around individual molecules 318 of interest, such as reference or target molecules. User interface 300 may receive selections of cluster shape and cluster size along each dimension. User interface 300 may also receive selections of the location of the individual molecule of interest within a cluster, whether at its center or elsewhere within the cluster. User interface 300 may have options to receive selections of multiple molecules of interest, where the cluster encompasses them all, or each molecule of interest may have clusters surrounding it.
[0407] User interface 300, exemplified by user interface 400, may have settings for map analysis 320. Any button, pop-up menu, or scan bar can be used for these settings. The map analysis 320 module can analyze the embedding of maps and molecules in any way. For example, map analysis 320 can cluster the map and return the clusters and the number of clusters as part of the visualization of map 310. Map analysis 320 can have hyperparameters to scan and change the number of clusters (because clustering is an undefined problem and requires hyperparameters). As another example, map analysis 320 can detect large and small clusters, or inliers and outliers (anomalous) molecules, for further examination.
[0408] User interface 300 (e.g., user interface 400) may have settings and options (322) for reporting by hovering over molecules or points in a map. For example, by moving the cursor over the map, any desired or selected information—corresponding to a molecule in the map—can be reported as a box near the cursor. User interface 300 may receive selections of which information will be reported or displayed, for example, by receiving hover input over a molecule. For example, scores (one or more) selected by the user (e.g., received by user interface 300 as selection input), such as expressive yield or fuzzy screening scores, may be reported. Moreover, for example, the distance / difference or similarity of a molecule to its nearest molecule may be reported. As another example, the average distance or similarity to other molecules in the map or in the same cluster may also be reported.
[0409] User interface 300 (e.g., user interface 400) may have setting options (324) for clustering and mosaic tiles. For example, a user can adjust clustering and mosaic tiles based on their preferences, or better display molecules based on features or aspects of the molecules.
[0410] Figure 17 An example mosaic tile illustration is shown according to some embodiments.
[0411] The interfaces of molecules can be divided into multiple clusters (e.g., to provide a direct way to interact with the interfaces in step 322). The clustering of partitioned spaces can resemble mosaic tiles. Each mosaic tile can represent a cluster of molecules. Figure 17 The image shows an example of a mosaic tile, where the mosaic tile is illustrated in the interface along with molecules (displayed by dots). Mosaic tiles or clusters can be obtained using any method of clustering and spatial partitioning.
[0412] In some implementations, any clustering algorithm (e.g., DBSCAN, HDBSCAN, or K-means) can be used to cluster molecules in the interface. The advantage of DBSCAN and HDBSCAN over K-means may be that they do not require prior knowledge of the number of clusters; instead, by setting and tuning a few parameters, they can correctly cluster molecules regardless of the number of clusters. These parameters can be set depending on the desired minimum cluster size. For example, a minimum cluster size can be set such that each outlier becomes part of a cluster.
[0413] Once the molecules on the interface have been partitioned into mosaic tiles, they can be identified using cluster labels. This can be done using either of the following two exemplary methods. In one method, the cluster label for each molecule can be determined by the label of the one or more molecules(s) closest to it in space. When considering multiple nearest molecules, majority voting can be used. In another method, the cluster label for each molecule can be determined by the label of the cluster center closest to it in space. For example, Figure 17 The example in the example uses the first method. Labels can include strings, numbers, barcodes, and names. Labels can be descriptions of one or more molecules found within the corresponding cluster, or they can be descriptions of the methodology used to cluster the molecules.
[0414] Partitioning the space may require traversing the space with a certain step size / resolution and using the algorithm mentioned above to find the cluster label for each point in this space. A finer step size / resolution can make the mosaic tile boundaries smoother and more accurate, but this may make the algorithm run slower.
[0415] After finding the cluster labels for each point in the space, a mosaic tile is calculated, and each molecule falling within a mosaic tile can be assigned a cluster label to that mosaic tile. The molecules to be clustered using the mosaic tile may or may not be in the training data.
[0416] Figure 18 Another example mosaic tile illustration using fractional color encoding according to some embodiments is shown.
[0417] Mosaic tiles can be colorless, or they can be colored in various ways. For example, they can look like... Figure 17 In this case, the mosaic tiles are randomly colored (or patterned). Another method for coloring mosaic tiles is to use fractions or quantities to color-code the tiles, such as protein yield, enrichment quality (e.g., the occurrence rate of a particular sequence in the target screening data compared to control screening, and how it changes across multiple screening rounds), binding energy, and the group / density of the cluster to which it belongs. An example of coloring mosaic tiles by fraction is shown in [the original text]. Figure 18 The image in the middle illustrates this. Using color-coded mosaic tiles with fractions allows for better inspection of clusters.
[0418] Different techniques can be used to achieve greater contrast between the colors of mosaic tiles. For example, when randomly coloring the tiles, graph coloring algorithms can be used to ensure that the mosaic tiles have as many different colors as possible, thus providing contrast between the tiles. Another way to enhance the contrast between tiles is to use histogram equalization or other transformations of the tile colors.
[0419] Mosaic tiles can be displayed in two different exemplary ways: grids and polygons. A grid in interface space can be used to display mosaic tiles, and each point in the grid can be colored with the color of the tile it belongs to. The mosaic tiles can be colored appropriately if the grid's step size is small enough compared to the scaling of the interface. A grid may appear as a mesh when zoomed out and as small dots when zoomed in. Another approach is to calculate the outline of the mosaic tile and visualize it as polygons with different numbers of angles.
[0420] Mosaic tiles can have a variety of applications. For example, they can be used for clustering or grouping molecules, or partitioning space and assigning new molecules to clusters within an interface. They can also be used for summarizing statistics of interfaces. Another example use of mosaic tiles is sampling from mosaic tiles, such as stratified sampling. Moreover, two clustering tiles that may come from the same interface or two different interfaces can be subtracted or intersected.
[0421] User interface 300 (e.g., user interface 400) may have options (326) for settings of subtraction and intersection. For example, molecules can be intersected and subtracted as described above to produce a new set of molecules for review.
[0422] Figure 19 An example intersection of two layers according to some embodiments is shown.
[0423] It is possible to make multiple interfaces of a molecule or multiple layers within a molecular interface intersect, where "multiple" means two or more. The following explanation refers to the intersection of multiple layers, but it can also be applied to multiple interfaces, where interfaces should be used instead of layers.
[0424] Intersection and subtraction of interfaces or layers have numerous applications and use cases and can be performed as part of step 326 below. One example is that intersection can be used to filter out shared molecules between two layers or interfaces. Conversely, subtraction can be used, for example, to filter out non-shared molecules between two or more layers or interfaces. For example, molecules with cross-reactivity can be removed. Another example is that antibodies that do not bind to a specific target antigen can be removed by subtracting the antibodies of unimmunized animals from the antibodies of animals immunized with the target. Another example use case for subtraction and intersection is sampling from shared or non-shared regions of layers / interfaces.
[0425] Following the example method for layer intersection, let n represent the number of layers to be intersected, where n can be an integer greater than or equal to two. For the intersection of multiple layers, these layers can be considered together, and a hypersphere or hypercube with a certain radius / length can be considered around each molecule in each layer. Then, for each layer, the method can iterate over all (n-1) other layers. A hypersphere / hypercube can be considered around the molecules in each other layer. In each iteration, molecules of the layer that fall into the hypersphere / hypercube can be retained and recorded, and the rest can be discarded. Performing this process over all n layers can provide n interfaces, each interface corresponding to the layer after the intersection. An example of two-layer intersection is shown in... Figure 19 As shown in the diagram. Increasing the radius can include more results in the intersection, while decreasing the radius can reduce the number of results in the intersection.
[0426] Multiple interfaces or layers can also be intersected with at least a few levels. For example, consider four layers. The intersection of each layer with at least one other layer, two other layers, or all three other layers can be calculated. Intersecting with at least a larger number of layers may be a more stringent condition, resulting in fewer molecules remaining after the intersection.
[0427] One possible application of interfacial or layer intersections is to identify non-specific antibodies present at interfaces or molecular layers corresponding to multiple targets. For example, consider the main interface of antibodies targeting several targets, where each layer corresponds to an antibody targeting a specific target. If an antibody is present at the intersection of multiple layers, it suggests that the antibody is non-specific to the target corresponding to that layer. The higher the number of layers an antibody resides in, the less specific it is likely to be. For example, if an antibody is present in multiple layers, it is more likely to cross-react with multiple targets.
[0428] Figure 20 An example of subtracting one layer from another layer is shown according to some embodiments.
[0429] Multiple interfaces or layers within a molecular interface can also be subtracted from each other, where "multiple" refers to two or more. This operation provides the opposite function to intersecting interfaces or layers. Subtraction of multiple interfaces is generally the same, where interfaces should be used instead of layers.
[0430] Following the example approach of layer subtraction, suppose we need to subtract (n-1) layers called L_1, L_2, ..., L_{n-1} from a layer called L_0, where n is an integer greater than or equal to two. For subtracting layers L_1, L_2, ..., L_{n-1} from layer L_0, consider a hypersphere or hypercube with a certain radius / length around each molecule in each of the layers L_1, L_2, ..., L_{n-1}. Then, for each molecule in layer L_1, remove all molecules in layer L_0 present in the hypersphere / hypercube surrounding that molecule. This process can be performed on all molecules in each layer L_1, L_2, ..., L_{n-1}. Therefore, some molecules can be removed from layer L_0 for each of the (n-1) layers. The resulting interface can be the subtraction of layers L_1, L_2, ..., L_{n-1} from layer L_0. Figure 20 The diagram shows an example of subtracting two layers. Increasing the radius allows you to subtract more molecules from the result, while decreasing the radius increases the number of molecules in the result.
[0431] Intersection and subtraction of interfaces or layers can be based on any feature of the input data for the interface to find or remove similarities between interfaces or layers. For example, if the input data is an antibody sequence, then intersection and subtraction can select or remove antibodies with similar sequences. Similarly, if the data is based on the structure or biophysical properties of the antibody, then intersection and subtraction of layers can select or remove antibodies with similar antibody structures or biophysical properties.
[0432] User interface 300 (e.g., user interface 400) may have settings (328) for sampling from the map. Sampling can be performed from the entire map or from specific portions of the map. Samples can also be drawn from clusters of interest or from clusters 318 surrounding individual molecules of interest. As explained in step 214 of system 200, different sampling methods can be used, which can be selected by the user. For example, simple random sampling with or without replacement or stratified sampling can be used. Sampling can also be scaled to the size of the clusters in the map. Figure 4 As shown, the user interface 300 can receive the following as input: sampling method, total number of samples, and whether sampling is proportional to the cluster size. The user interface 300 can also receive drawing settings as input for how the extracted samples are visualized in the map visualization window. For example, the shape, size, color, transparency, and legend labels of the extracted samples can all be selected by the user.
[0433] User interface 300 (e.g., user interface 400) may have options (330) for editing extracted samples in the map. Different editing methods may be used to edit the samples. For example, a user can hover the cursor over a sample to move it by dragging and dropping it in the map visualization window. Alternatively, a user can edit the sample in a table, in a pop-up menu, or via one or more buttons.
[0434] This interface can be used to sample molecules. Molecular sampling can be performed for various reasons. For example, sampling can be used to select molecules for further research in the laboratory.
[0435] Different sampling methods can be used to sample molecules. Various sampling methods can be used, such as simple random sampling, bootstrap sampling, sampling with or without replacement, hierarchical sampling, cluster sampling, multi-stage sampling, network sampling, snowball sampling, and Monte Carlo sampling. For example, the simplest sampling algorithm might be simple random sampling, where molecules are randomly sampled. The random seed used in the computer to generate randomness can be set or not set, thus allowing or preventing the reproduction of the results, respectively.
[0436] Another possible sampling algorithm that can be used is hierarchical sampling, which samples from clusters or mosaic tiles proportionally to the size of the clusters or tiles. In this approach, the larger the cluster or tile, the more samples are taken from it. In this approach, the interfaces of the molecules can first be clustered into multiple clusters or tiles using any clustering algorithm (such as DBSCAN, HDBSCAN, K-means, or hierarchical clustering). The parameters of the clustering algorithm can be set based on whether it is desired to consider small clusters as separate clusters. The number of samples drawn from each cluster can be calculated based on the relative size of its cluster or tile to the entire interface and other clusters or tiles. For example, the sample size of each cluster can be calculated as follows: round((cluster_population / all_population) n_samples), Where `cluster_population` is the size of the clusters or tiles, `all_population` is the total number of molecules in the interface, and `n_samples` is the desired total number of samples. If the calculated number of samples selected from small clusters or tiles is zero, then the number of samples from each of these clusters can be set to one or a larger number. This can be done when the user expects samples (one or more) from outliers or small clusters in addition to samples from larger clusters.
[0437] Figure 21Examples of diverse sampling from the interface of molecules are shown according to some embodiments.
[0438] Samples can be randomly drawn or diversified from clusters or tiles. Diversified sampling from clusters or tiles can be achieved by applying another cluster to each cluster, resulting in sub-clusters within each cluster. The number of sub-clusters within each cluster or tile can be equal to the number of samples to be drawn from that cluster. For example, K-means can be used, where K equals the number of samples to be drawn from the cluster. Samples can then be randomly drawn from each sub-cluster in the cluster / tile (random method), or the centers / centroids of the sub-clusters can be considered as samples from the cluster / tile (deterministic method). Figure 21 The image shows examples of diverse sampling from the interfaces of molecules.
[0439] In stratified sampling from clusters, either hard or soft clustering can be used. Hard clustering assigns each molecule completely to a cluster, while soft clustering provides a fractional allocation of molecules to clusters. Different methods can be used for soft clustering, such as stratified clustering. In stratified clustering, cutoff values can be used to define the height within a stratum, and each cutoff value determines the cluster. By changing the cutoff values, some clusters can be merged into one cluster, or clusters can be divided into smaller clusters.
[0440] Sampling can also be performed from mosaic tiles at the molecular interfaces. To do this, first, the mosaic tiles described earlier are applied to the molecular interfaces. Then, samples can be drawn proportionally to the tile size; this is equivalent to stratified sampling as described above. Alternatively, the distance of each molecule from the center of its mosaic tile / cluster representation can be calculated. Based on this distance, the probability of sampling can be determined. For example, molecules closer to the tile center may have a higher probability of being sampled.
[0441] Depending on the sequence portion of the molecule, the extracted samples can also be diverse. For example, when the molecule is an antibody, the sample can be diverse in terms of the sequence portions of (one or more) antibody chains, such as frame regions or complementarity-determining regions (CDRs). For instance, the sample can be diverse in terms of CDR3, which can be an important region in the antibody's complementary site used to determine the binding specificity to the antigen.
[0442] For diverse sampling of molecular sequence portions, numbering methods such as IMGT numbering can be applied to the sequences, allowing residues at the same position to correspond across different sequences. Then, the portion of interest in the molecule's sequence, such as CDR3, can be considered at the interface. The sequence portions can be classified into multiple unique categories. One or more molecules can then be sampled from the sequence portions of each category.
[0443] Classifying or clustering parts of a sequence into groups (where each group contains a unique part) can be hard classification. Soft classification can also be used, where each part of the sequence can be assigned a score to these categories. This can be done using different techniques, such as hierarchical clustering. For example, in hierarchical clustering, a cutoff value can be used on the height of the category stratification, and depending on that cutoff value, a part of the sequence can belong to one of the categories.
[0444] Sampling can also be diverse in terms of the structure of a molecule or parts thereof. To achieve this, the input data to the interface can be only a specific structural part of the molecule, or the three-dimensional positions of atoms in a specific part of the molecule can be considered during sampling. Alternatively, all or some specific parts of the molecular structure can be considered to achieve sample diversity. For example, the atomic structure at complementary sites in the molecule, or at the top of the heavy chain, can be diverse across samples.
[0445] Another approach to sampling molecules is to draw samples from these molecules using probability scores. Any score can be viewed as a probability of sampling from a molecular interface. For example, if the molecule is an antibody, these scores can be determined by metrics such as protein yield, enrichment, binding energy, and the group or density of its cluster. These scores can be normalized so that they fall between 0 and 1 and sum to 1, thus behaving like probabilities. The higher a molecule's sampling probability score, the more likely it is to be drawn as a sample. Optionally, multiple probability scores can be weighted (e.g., by a weighted average) to reach a consensus on sampling with multiple scores.
[0446] Descriptors or scores can also be added to the data used to show, hide, or toggle the molecular interface. By doing so, some molecules with low scores can be removed from the molecular interface, and sampling can then be performed on the reduced interface. For example, if the molecule is an antibody, antibodies with low enrichment scores, which bind to or do not bind to the target, can be removed before sampling.
[0447] Sampling can be focused around one or more reference molecules at the interface of the molecules. Reference molecules can be any molecule of interest or important molecule. For example, it can be an antibody previously shown to bind to a specific target. Sampling for additional antibodies near known binding sites increases the likelihood of identifying additional antibodies with similar properties. Conversely, sampling can also be performed at a distance from some molecules.
[0448] Several methods can be used to sample around one or more reference molecules. For example, samples can be randomly drawn from clusters or mosaic tiles containing reference molecules. Alternatively, higher sampling probabilities can be assigned to molecules closer to the reference molecule. Another approach is to deterministically obtain the (nearest) molecule closest to the reference molecule at the interface. Moreover, using the molecular interface / layer intersection algorithm described earlier, the intersection of the interfaces between the reference molecule and the molecules can be calculated with a certain radius. Samples can then be drawn randomly or deterministically from the result of the intersection.
[0449] Another possible approach for sampling is to use probability distributions or a superposition of multiple probability distributions. Each probability distribution can provide the opportunity to sample molecules at different locations on the interface. For example, if the interface is a 2D map, then the probability distributions are 2D distributions in space, which can be visualized through spatial colors or heatmaps. For example, in Figures 22A-22G The paper uses color coding for some 2D probability distributions, where the color map from blue to red represents the probability from smallest to largest.
[0450] This sampling type, which uses the superposition of probability distributions, allows for iterative sampling of batches of samples. The batch size can be any integer between a given number and the desired number of samples (inclusive). In each iteration, a batch of samples can be drawn. Samples drawn from all previous iterations can be accumulated and used to update the probability distribution in the next iteration. The probability distribution may change across iterations. Smaller batch sizes can lead to more accurate sampling because the samples affect the probability distribution more gradually and accurately. However, this can slow down the sampling process, as smaller batches increase the number of iterations required to reach the desired number of samples. There may be a trade-off between sampling speed and accuracy.
[0451] In sampling via probability superposition, a weighted average of the probabilities can be used. The weights can be based on the importance of each probability distribution in the sampling. Any weighting technique can be used. For example, in some implementations, a power transformation can be applied to each probability distribution, where the weights serve as the power. By changing the power, it can tune how much each probability distribution is concentrated on its mode (so that sampling can focus on its high-probability numerators), or how much it becomes similar to a uniform distribution.
[0452] Any superposition of probability distributions can be used for sampling. For example, in some implementations, the probability distribution can be the density of clusters, the diversity of samples, proximity to cluster or patch centers, the diversity of sequence portions, input scores, proximity to (or more) references, and any other criteria that can be added. The probability distribution for cluster density can be calculated using kernel density estimation of molecules or any other density estimation method (see, for example...). Figure 22A In calculating the probability distribution of sample diversity (see, for example...), Figure 22B In this algorithm, the distance between each molecule and the samples already sampled in previous iterations can be calculated. For each molecule, the k nearest neighbors can be considered, where k can be any positive integer, such as one, two, or three. The greater the (average) distance between a molecule and its nearest neighbor(s), the higher its probability of being used for sampling.
[0453] In calculating the probability distribution of proximity to cluster or tile centers (see, for example) Figure 22C In this framework, clustering tiles or any clustering algorithm can be used to cluster the spatial distribution of molecules. The cluster centers can then be calculated as a summary statistic of the cluster, such as the mean or weighted average of the molecules in the cluster, or as tile centers. The distance of each molecule to its cluster center can then be calculated. The smaller this distance, the higher the probability of sampling. This can be used to sample more molecules from the core or center of the cluster.
[0454] The probability distribution of diversity in the calculated sequence portion (see, for example) Figure 22D In this context, any sequence portion can be considered. For example, if the molecule is an antibody, then the CDR portion can be considered, and any technique can be used, such as BLAST (Basic Local Alignment Search Tool) or the minimum Hamming distance between two strings obtained through dynamic programming, to calculate the difference between the CDR portion of each molecule and the CDR portion from the sampled samples from previous iterations.
[0455] In calculating the probability distribution of the input scores (see, for example) Figure 22E In this context, any input score(s) obtained from the user can be used. Some examples of this score when the molecule is an antibody could be an expressibility yield or enrichment score. The scores can be normalized to sum to one, thus behaving like a probability distribution. This normalization can be done using any technique, such as dividing by the sum of the scores or using a softmax function to have a Gaussian-like distribution. Depending on whether a higher or lower score is better, the probabilities derived from the scores can be used as is or vice versa.
[0456] In calculating the probability distribution of proximity to (one or more) references (see, for example) Figure 22F In this context, any one or more user-provided input references can be used. The one or more reference molecules can be any molecule, such as those previously characterized. The distance of each molecule to all references or to its k nearest references can be calculated, where k can be any positive integer. If sampling is desired at locations close to (or respectively far from) the one or more references, it can be configured such that the smaller the distance (or respectively, the larger), the higher the probability of sampling.
[0457] Any other probability distribution can be added to the probability distribution used for sampling. Finally, the superposition of the above probability distributions can be used as the overall probability of sampling (see example...). Figure 22G Any technique, such as roulette wheel, the inverse of the cumulative distribution function, or Monte Carlo sampling, can be used to randomly (e.g., randomly) sample from the distribution. Alternatively, samples can be drawn deterministically (e.g., non-randomly) from the distribution by drawing the molecule with the highest probability.
[0458] Figures 22A-22G Example 2D probability distributions for sampling are shown, generated using different probability distributions on the same interface according to some embodiments. Figure 22A An example 2D probability distribution is shown, generated using a probability distribution of cluster density on the same interface according to some embodiments. Figure 22B An example 2D probability distribution is shown, generated using a probability distribution of sample diversity on the same interface according to some embodiments. Figure 22C An example 2D probability distribution is shown, generated on the same interface using a probability distribution of proximity to cluster or tile centers, according to some embodiments. Figure 22D An example 2D probability distribution is shown, generated using a probability distribution of sequence partial diversity on the same interface according to some embodiments. Figure 22E An example 2D probability distribution is shown, generated using the probability distribution of input scores on the same interface according to some embodiments. Figure 22F An example 2D probability distribution is shown, generated on the same interface using a probability distribution of proximity to (one or more) references according to some embodiments. Figure 22G An example 2D probability distribution generated using the overall sampling probability on the same interface according to some embodiments is shown.
[0459] In another sampling method, molecules can be sampled using the intersection and subtraction of interfaces or layers introduced earlier. For example, sampling can be performed on the result of subtracting one layer or interface from another, or samples can be drawn from the intersection of two layers or interfaces.
[0460] When the molecule is an antibody, cross-section and / or subtraction can be used (as described above) to generate a good set of antibodies for sampling from. A useful implementation of sampling via cross-section and subtraction of interfaces or layers can be sampling from a library of immune antibodies from one or more immunized laboratory animals. This can also be subtraction / cross-section of a main interface (which may include several targets). Consider one or more unimmunized animals (e.g., laboratory mice) whose bodies have not been injected with a specific target antigen. Alternatively, assume there are one or more immunized animals that have been injected with the target antigen. Antibodies (which may be candidates for drugs) can also be sampled from immunized animals without in-laboratory screening and enrichment. For this purpose, one or more layers or interfaces of unimmunized antibodies can be subtracted from one or more layers or interfaces of immunized antibodies. Doing so reduces antibodies that are not relevant to the target antigen, as the remainder after subtraction is expected to be added after injection of the target.
[0461] Optionally, after subtracting the non-immune library from the immunized library, the intersection of one or more layers or interfaces can be considered. By doing so, antibodies shared among multiple animals after subtracting the non-immune library can be considered. After such subtraction, when an antibody is present in multiple animals immunized with the same target antigen, that antibody may have a higher probability of binding to the target of interest.
[0462] When the molecule is a protein, sampling can also be performed based on germline genes, either by sequence or structure. For example, antibodies can be classified into multiple groups, each derived from the same germline gene or the same molecular parent mutation. Then, any of the sampling algorithms mentioned above can be used, where clusters are groups of antibodies derived from parental mutations (e.g., each cluster shares the same parent, and diversity between parents is taken into account). To classify antibodies into these groups, any technique can be used, such as BLAST (Basic Local Alignment Search Tool), BLOSUM (Block Substitution Matrix), PAM (Point Acknowledged Mutation) matrix, substitution matrix, PSSM (Position Specific Scoring Matrix), or Hamming distance, to calculate the similarity between the antibody and its parent.
[0463] The sampling algorithm can be a combination or integration of any of the sampling methods mentioned above. For example, stratified sampling or mosaic sampling can be performed while considering different sequence parts. Another example is performing stratified sampling or mosaic sampling while using probability scores for sampling. In some implementations, stratified sampling can be used while treating the distance to the cluster or tile center / representation as a probability score for sampling.
[0464] Sampling algorithms can also be used back-to-back in a connected manner. In this connection, the next sampling stage can sample from the sample obtained by the previous stage or add additional samples. Examples of the former could be performing stratified sampling or mosaic sampling, followed by sampling from the obtained sample using multiple sequence parts. Examples of the latter could be first performing stratified sampling from the entire interface of the molecule, then using another sampling method to sample additional molecules from the entire interface, and combining them to obtain the complete sample.
[0465] Molecules can be compared with each other at an interface. For this purpose, all or one or more parts of a molecule can be compared with each other, either whole or partially. This comparison can also be used to compare samples of molecules. The comparison can be an equivalence comparison, or it can use any technique to measure their similarity or difference. If the comparison is structural, their 3D structures can be compared using any technique, such as template matching, 3D matching, or any other comparison method. If the comparison is sequence-related, various techniques can be used, such as BLAST (Basic Local Alignment Search Tool) search, BLOSUM (Block Substitution Matrix) matrix, PAM (Point Accepted Mutation) matrix, substitution matrix, PSSM (Position Specific Scoring Matrix), Hamming distance, or available codebases. One possible application of this comparison is to compare a molecule with other previously characterized molecules to find additional molecules.
[0466] After the user performs sampling, the interface can also recommend new supplementary samples, which may be of interest or missing from the samples obtained through the already performed sampling. For example, the user can use a sampling technique, and the interface can then use unused sampling techniques to generate recommendations.
[0467] User interface 300 (e.g., user interface 400) may have a button (332) for exporting files. Any file(s) can be exported. For example, data 102 or dataset 106, or extracted samples (214 and 324), or scores (216 and 316) can be exported. Different methods can be used to export files; for example, files can be downloaded or moved from one folder / directory to another, or they can be displayed / reported in the user interface using any method such as tables, visualizations, pop-ups, etc.
[0468] Furthermore, the interface can export reports on the sampling methods and statistics of the samples. Statistical analysis and visualization of the samples can be reported in various formats, such as plots, graphs, tables, text, and descriptions.
[0469] Figure 4A , Figure 4B , Figure 4C , Figure 4D , Figure 4E This is an example map visualization 310 based on user interfaces 300, 400 according to an embodiment. Map visualization 310 can visually indicate different layers. Figure 4A Layer 1 is shown. Figure 4B Layer 1 and Layer 2 are shown. Figure 4C Layers 1 and 2, the reference molecule, and the clusters surrounding the reference molecule are shown. Figure 4D The diagram shows layers 1 and 2, a reference molecule, clusters around the reference molecule, and sampling from layers 1 and 2. Figure 4E Layer 1 and Layer 2 are shown, with scores encoded by the tag size.
[0470] The user interface 300 can report the map not only as a visual map, but also in any format. For example, it can report the map via sound, smell, or touch. In this way, the user interface can also be useful to visually impaired users. In some embodiments, the user interface 300 can display the map of molecules as a hologram, where the user can walk between molecules and touch them or interact with them using different gestures. In some other embodiments, the user interface 300 can generate different odors for the molecules in the map. In some embodiments, this odor can also be used to encode scores in the map, where each category of score has a different odor, or an odor gradient can be used for continuous scores. In some other embodiments, the user interface can report the map via sound, where the location and characteristics of molecules in the map are reported through some generated sound. Reports from different senses can also be combined, giving the user a variety of exploration options.
[0471] The mapping system 100 can be used for any reason, application, and purpose. For example, it can be used for antibody selection, immune response (high or low response), tracking cluster size during subsequent immunizations (for vaccine development), associating cluster(s) with(s) functions, secondary assays (e.g., binding, signaling, etc.), deimmunizing antigens or antibodies or proteins, finding common complementary sites (e.g., common epitopes) among different variants of antigens, assessing the robustness of immune responses (more clusters or greater cluster diversity), etc.
[0472] System 100 can be used for different applications and use cases. Example use cases are provided below.
[0473] Antibody discovery
[0474] Complementary site maps can be used to efficiently sample antibodies from antibody libraries. Antibodies can be selected from a single target cluster. Selecting multiple antibodies from a single complementary site cluster increases the number of antibodies with similar biophysical characteristics. Antibodies can be broadly selected from multiple clusters to increase the diversity of the biophysical properties of the sampled antibodies.
[0475] In some embodiments, the antibody or its antigen-binding fragment includes antibodies derived from a species, including but not limited to mice, cattle, rabbits, camels, llamas, humans, alpacas, and standard species.
[0476] Humanization
[0477] The complementation site map contains the original antibody and humanized variants. Antibodies are humanized to increase human content while retaining their complementary sites. The complementation site map can be used to visualize complementation site changes unintentionally caused during humanization. The humanized variant closest to the original antibody on the complementation site map has the highest probability of maintaining the original antibody's binding and functional properties.
[0478] vaccine development
[0479] Complementary site maps can be used to assess immune responses following vaccination. They can also be used to identify and compare immune responses from different vaccine formulations.
[0480] In some embodiments, the map can identify specific binding sites on antibodies to understand how antibodies neutralize pathogens and to design effective vaccines and therapies.
[0481] Format conversion
[0482] Complementation site maps can be generated from different antibody formats, such as single-chain and VH-VL antibodies. Complementation site maps with multiple antibody types will allow conversion from one antibody form to another. Antibodies in the same cluster will bind to shared epitopes.
[0483] Diagnosis or monitoring
[0484] Complementary site maps can be used to identify and monitor changes in disease-associated antibody libraries.
[0485] In some embodiments, system 100 may suggest amino acid variations to improve protein / antibody properties.
[0486] Figure 5 This is another example diagram of the mapping system 100 according to an embodiment.
[0487] Mapping system 100 maps molecular data to a visual interface. Mapping system 100 has a processing subsystem including one or more processors and one or more memories coupled to the processors. The processing subsystem is configured to receive input molecules. Input molecules can be one or more groups of molecules. Each molecule can be defined as an information sequence. Mapping system 100 can encode the information sequence. Mapping system 100 can generate a dataset by processing the input molecules. The dataset can include the encoded information sequence, three-dimensional coordinates of the information sequence for each molecule to define its molecular structure, features, and fingerprints. Mapping system 100 can transform the dataset to generate a molecular map through feature extraction and fingerprint generation, thereby reducing the high number of dimensions to a lower-dimensional data representation that can be visualized, while capturing valuable data in the lower-dimensional data representation.
[0488] Mapping system 100 can generate a map user interface that includes a molecular map (a visualization of a lower-dimensional data representation of a dataset). Mapping system 100 can provide the map user interface to user device 114 via network 116.
[0489] Mapping system 100 can be used to map proteins, protein-like molecules, or fragments thereof onto a visual interface and generate maps for display and interaction. In some embodiments, mapping system 100 can be used to map non-protein molecules, such as lipids, complex sugars, nucleic acids, etc. Mapping system 100 may include a processing subsystem comprising one or more processors and one or more memories coupled to the processors.
[0490] The processing subsystem can be configured to receive data (e.g., raw data, preprocessed data, input data, etc.) for one or more sets of one or more input proteins, protein-like molecules, or fragments thereof. Raw data (e.g., source data or primary data) is data that has been collected from a source but has not yet been processed, cleaned, or analyzed. Raw data can be unprocessed and in its initial state, and may contain errors, outliers, or inconsistencies. Raw data can take many forms and can include numbers, text, images, audio, or any other type of data. While raw data itself may not be immediately useful, once processed and analyzed, it can have the potential to provide valuable insights.
[0491] Data can include features of one or more proteins, protein-like molecules, or fragments thereof. Data can include features such as sequence, structure, biophysical properties, etc. Features can be various measurable attributes or characteristics that describe entities and can be used to analyze these entities (e.g., binding, function, stability, expressibility, affinity, immunogenicity, etc.). Features can be considered as dimensions of data. For example, features can be used as input to train models for machine learning. There can be original dimensions and extracted dimensions. Features can also include, for example, primary structure (e.g., the amino acid sequence in a protein, which determines its unique properties and functions; this linear sequence is crucial because even a single change can affect protein function), secondary structure (e.g., local folding patterns within a protein, such as α-helices and β-sheets stabilized by hydrogen bonds; these structures contribute to the overall shape and stability of the protein), tertiary structure (e.g., the three-dimensional shape of a single polypeptide chain formed by interactions between the side chains (R groups) of amino acids; this structure is essential for protein functionality and its interactions with other molecules), quaternary structure (e.g., the arrangement of multiple polypeptide chains (subunits) in a multi-subunit protein; this structure is important for the function of proteins that act as complexes (such as hemoglobin), and binding sites (e.g., specific regions on a protein where ligands (such as substrates, inhibitors, or other proteins) can bind; these sites...). The properties of proteins include: hydrophobicity and hydrophilicity (e.g., the distribution of hydrophobic (water-repelling) and hydrophilic (water-attracting) regions within a protein; this characteristic affects protein folding, stability, and interactions with other molecules), molecular weight (e.g., the mass of a protein, which can affect its migration rate in techniques such as gel electrophoresis and its behavior in solution), isoelectric point (p1) (e.g., the pH at which a protein carries no net charge; this property is important for protein purification and characterization techniques), and functional domains (e.g., specific regions in a protein that have unique functional roles, such as DNA-binding domains, catalytic domains, or transmembrane regions; these domains are often conserved among different proteins with similar functions).
[0492] The processing subsystem can be configured to generate at least one dataset by processing one or more sets of data containing one or more input proteins, protein-like molecules, or fragments thereof. Processing may involve cleaning, organizing, and transforming the data into a more usable format. In some embodiments, processing may include steps such as filtering data or removing errors, outliers, or inconsistencies.
[0493] The processing subsystem can be configured to transform data and one or more of at least one dataset to generate maps (e.g., spatial representations that associate protein-like molecules or fragments in space based on one or more models or clusters) and additional features through feature extraction (e.g., transforming raw data into informative features), feature selection (e.g., identifying the most relevant features for a model), or feature creation (e.g., creating new features from existing features, such as combining or splitting features) and additional features, thereby reducing the higher-dimensional data representation to a lower-dimensional data representation that can be depicted and / or visualized through a visual interface, while capturing valuable information in the lower-dimensional data representation.
[0494] Visualizing high-dimensional data can be challenging, and transformed, lower-dimensional representations can make the data easier to interpret and present. Lower-dimensional data representations can improve the use of computing and network resources because they reduce transmission size, make interface handling and rendering more efficient, and improve memory usage on display devices. By reducing the number of features, dimensionality reduction can lower the computational resources required for data processing and analysis. By eliminating redundant and irrelevant features, dimensionality reduction can help reduce noise in the data. This can lead to more accurate and robust data representations. It can also improve data analysis by making it easier to identify and focus on the most significant variables, thus simplifying the analysis and interpretation of complex datasets. The user interface can be displayed on devices with different screen sizes and resolutions, which can affect the visualization rendered within the interface. Transformed representations can help improve UI performance for these complex datasets.
[0495] Dimensionality reduction reduces the number of input variables or features in a dataset while preserving relevant information. This process can help simplify models, reduce computation time, and improve performance. Example types of dimensionality reduction techniques include feature selection (selecting a subset of features) and feature extraction (transforming data to a lower-dimensional space). Features can also be considered as dimensions of the data. Extracting or selecting features reduces dimensionality. Some example considerations for feature reduction may include, for example: Choosing the right approach: It is important to select the appropriate dimensionality reduction technique based on the nature of the data and the specific analytical objectives. For example, principal component analysis (PCA) is often used due to its ability to preserve variance, while t-SNE is useful for preserving local structure; Balancing dimensionality and information: The goal is to reduce dimensionality while preserving the most salient features. Techniques such as PCA transform data into a new set of variables (principal components) that capture the maximum variance, thereby ensuring that the most informative aspects of the data are preserved. Regularization and supervised methods: Methods such as linear discriminant analysis (LDA) and partial least squares (PLS) use class labels to maintain relevant information associated with a specific outcome or classification. These supervised techniques ensure that the reduced dimensionality still reflects important biological signals; Combining multiple techniques: Sometimes, combining dimensionality reduction with other methods such as transfer learning can enhance the robustness and interpretability of data. For example, autoencoders can be used to learn compact representations of data, which can then be tuned with additional data to improve performance; and Validation and cross-validation: Regularly validating the reduced data against independent datasets helps ensure that critical information is not lost during dimensionality reduction. Cross-validation can also be used to assess the stability and reliability of the reduced dimensionality.
[0496] Lower-dimensional data representations can include one or more clusters of proteins, protein-like molecules, or fragments thereof. Proteins, protein-like molecules, or fragments thereof can include one or more input proteins, protein-like molecules, or fragments thereof, or newly generated proteins, protein-like molecules, or fragments thereof. A cluster is a group of proteins, protein-like molecules, or fragments thereof (and related data elements) that are "similar" to each other within a dataset. Data clustering can be a subgroup of a large dataset, where each data point is closer to a cluster center than other cluster centers. This proximity can be determined by minimizing the squared distance between a data point and its corresponding cluster center. Clustering helps identify patterns, trends, and relationships within data. By using clustering techniques, mapping system 100 can simplify complex datasets and reveal hidden structures that may not be immediately apparent. Clustering methods include: K-Means clustering (which divides data into (k) clusters by minimizing the distance between data points and cluster centroids), hierarchical clustering (which builds a cluster tree by merging or splitting existing clusters based on their proximity), density-based clustering (DBSCAN) (which forms clusters based on the density of data points, identifying high-density areas as clusters and low-density areas as noise), etc.
[0497] The processing subsystem can be configured to generate a visual map interface comprising a map as a visual representation of the dataset's lower-dimensional data. This visual representation can include visualizations of the dataset as one or more layers representing proteins, protein-like molecules, or fragments thereof, each layer comprising one or more clusters of proteins, protein-like molecules, or fragments thereof. Layers can be used to organize and display different data points of proteins, protein-like molecules, or fragments thereof on the map. Each layer can represent a specific set of data, and multiple layers can be overlaid to combine different sets and generate different visualizations. The map can have a base layer and data layers on which additional data is overlaid. These layers can be interactive, and clicking on features or data points can lead to processing, filtering, or displaying more / less data. These layers can be overlaid on each other (e.g., like layer addition). Map generation or analysis can include generating clusters, arranging embeddings within clusters, layer-by-layer embedding, embedding individual molecules, sampling from the map, or encoding scores in the map. Typically, a map is generated followed by cluster generation (e.g., clustering the data is required to detect clusters). The map can contain clusters and visualize them.
[0498] The processing subsystem can be configured to provide the visual map interface with tools for interacting with the map and for examining, searching, sampling, clustering, and analyzing one or more proteins, protein-like molecules, or fragments thereof, or newly generated proteins, protein-like molecules, or fragments thereof. This can aid in forward-looking decision-making because these tools can be used to proactively interact with the data. Tools may include, for example, UI tools. UI tools are specialized software applications that help create, modify, and explore the visual map interface.
[0499] The processing subsystem can be configured to enable the system to receive commands or detect interactions with the map at the visual map interface via the tool, update the map based on the commands or interactions, and enable the mapping system 100 to trigger an update of the visual map interface with the updated map.
[0500] In some embodiments, the input molecule is a protein or protein-like molecule that includes both an antibody and an antigen. For example, the antibody may be derived from a species including, but not limited to, mice, cattle, rabbits, camels, llamas, humans, alpacas, and standard species.
[0501] In some embodiments, the input molecules are antibodies and antigens, and the molecular map is a complementary site map. In some embodiments, the input molecule is an antigen, and the molecular map is an epitope map.
[0502] In some embodiments, the processing subsystem (e.g., server 112) uses machine learning or statistical methods involving one or more of dimensionality reduction, reconstructive autoencoders, variational autoencoders, adversarial autoencoders, neural networks, graph neural networks, attention networks, and recurrent networks to extract features and generate fingerprints.
[0503] In some embodiments, system 100 has a data storage device for a molecular database 118, wherein each molecule is assigned a unique index.
[0504] In some embodiments, system 100 compares a molecule with a database of molecules 118 and assigns the molecule an index to the closest molecule in the database of molecules, wherein molecules assigned the same index have similar sequences, structures, or properties.
[0505] In some embodiments, feature extraction is performed using layer-by-layer embedding, wherein features are extracted individually for each of the multiple datasets or multiple parts of the datasets. In some embodiments, features can be plotted as overlaid visual layers. In some embodiments, the map user interface includes visualization of the datasets and also has control inputs to allow viewing of these layers individually or in relation to other layers. In some embodiments, one or more layers are used to train the feature extraction, and one or more other layers are used to test the feature extraction.
[0506] In some embodiments, system 100 has a user device 114 for displaying a map user interface. In some embodiments, the map user interface at user device 114 includes a visualization of extracted features of a dataset. In some embodiments, the map user interface is one-dimensional, two-dimensional, three-dimensional, four-dimensional, or higher-dimensional. In some embodiments, the layer-wise embedding is a time series that varies over time, and the map user interface includes a three-dimensional embedding that varies over time as a time series representing a four-dimensional embedding. In some embodiments, the map user interface includes a visualization of clusters around molecules of interest. In some embodiments, the map user interface includes one or more clusters of molecules. In some embodiments, the layer-wise embedding includes individual molecule embeddings.
[0507] In some embodiments, server 112 performs feature extraction on each molecule in the dataset to obtain each molecule embedding.
[0508] Figure 5A network diagram depicting the network environment of mapping system 100 and machine learning system 510 according to an embodiment is shown, along with multiple user devices 104 interconnected via a communication network 116. These systems and devices collaborate in a manner disclosed herein to map molecules into a visual interface. Each user device 114 is a device operable by a user to interact with mapping applications provided by system 100.
[0509] Server 112 and the mapping application can provide one or more interactive user interfaces that can be accessed by a user operating user device 114. The interactive user interface provided at user device 114 includes a user interface for providing input to the user.
[0510] The machine learning system 510 is configured to perform various machine learning processes. The machine learning system 510 can generate datasets by processing input molecules. The datasets may include encoded information sequences, three-dimensional coordinates of the information sequence for each molecule used to define its molecular structure, features, and fingerprints. The machine learning system 510 can transform the datasets through feature extraction and fingerprint generation to generate molecular maps, thereby reducing the high number of dimensions to a lower-dimensional data representation that can be visualized, while capturing valuable data within this lower-dimensional representation.
[0511] In some embodiments, the machine learning system 510 uses one or more machine learning or statistical methods involving dimensionality reduction, reconstructive autoencoders, variational autoencoders, adversarial autoencoders, neural networks, graph neural networks, attention networks, and recurrent networks to extract features and generate fingerprints.
[0512] In some embodiments, system 100 has a data storage device 118 for a database of molecules, wherein each molecule is assigned a unique index. In some embodiments, machine learning system 510 compares a molecule with a database of molecules 119 and assigns the molecule an index to the closest or most similar molecule in the database of molecules, wherein molecules assigned the same index have similar sequences, structures, or properties.
[0513] In some embodiments, the machine learning system 510 uses layer-by-layer embedding to perform feature extraction, wherein features of each of the multiple datasets or features of each of the multiple parts of a dataset are extracted individually for multiple datasets or multiple parts of a dataset. In some embodiments, features may be plotted as visual layers superimposed on each other. In some embodiments, the map user interface includes visualization of the dataset and also has control inputs to enable viewing of these layers individually or in relation to other layers. In some embodiments, one or more layers are used for training the feature extraction, and one or more other layers are used for testing the feature extraction. In some embodiments, layer-by-layer embedding includes individual molecular embeddings.
[0514] In some embodiments, the machine learning system 510 performs feature extraction on each molecule in the dataset to obtain each molecule embedding.
[0515] Figure 6 This is a flowchart of an example method 600 for mapping molecules to a visual interface according to an embodiment. In some embodiments, method 600 relates to mapping molecular data to a visual interface using a mapping system 100. In some embodiments, method 600 is a computer-implemented method for mapping molecules to a visual interface. The embodiments provide improved visualization of molecular data. For example, visualization can provide visual identifiers for molecular data. Visualization can provide clustering of similar molecules. The interface is interactive to update the visualization in response to commands. In some embodiments, method 600 relates to one or more non-transitory computer-readable media storing machine-interpretable instructions that, when executed by the mapping system 100, cause the mapping system 100 to perform method 600 for mapping molecules to a visual interface.
[0516] At 602, the mapping system 100 receives input molecules. The mapping system 100 is configured to receive input molecule data 102. The input molecules can be one or more groups of one or more molecules. Each molecule can be defined as an information sequence. In some embodiments, the input molecules are protein or protein-like molecules comprising antibodies and antigens. Antibodies include species-specific antibodies, including but not limited to mice, cattle, rabbits, camels, llamas, humans, alpacas, and standard species.
[0517] The input molecular data 102 can come from any source. Part or all of the input molecular data 102 can be in other formats. Furthermore, biophysical properties—high-volume data—from all origins can be used as part of the input molecule. These properties include, but are not limited to, pH, hydrophobicity and hydrophilicity, negative and positive charge, and solvent exposure. Part of the input molecular data 102 can also be obtained from nucleic acid sequencing methods. The input molecular data 102 can also undergo some preparation and preprocessing steps, such as transformations.
[0518] Mapping system 100 encodes the information sequence in each input molecule. Each molecule can be viewed as an information sequence by mapping system 100. For example, if the molecule is a protein, then mapping system 100 can model the molecule as an amino acid sequence. This model can be tailored to multiple dimensions of data related to the molecule. The information sequence can also be used for dataset generation. Different encoding methods can be used to encode the information sequence.
[0519] At 604, the mapping system 100 generates a dataset 106 by processing the input molecular data 102. Dataset 106 includes an encoded information sequence and three-dimensional coordinates of the information sequence for each molecule used to define its molecular structure. The three-dimensional coordinates of the atoms in each input molecule can be considered as part of the features of dataset 106. The three-dimensional coordinates of the atoms contain structural information about the molecule. Features extracted from the input molecular data 102 can be used to prepare dataset 106. For example, the encoded information sequence, the three-dimensional coordinates of the sequence, and other properties (including but not limited to solvent exposure, hydrophobicity, and charge) can all be used as features of dataset 106.
[0520] At 606, the mapping system 100 transforms the dataset 106 to generate a molecular map 200 through feature extraction and fingerprint generation. This transformation can be any linear or nonlinear transformation. The fingerprint can be extracted from the original dataset mentioned above. In some embodiments, the mapping system 100 can use machine learning or statistical methods involving one or more of dimensionality reduction, reconstructive autoencoders, variational autoencoders, adversarial autoencoders, neural networks, graph neural networks, attention networks, and recurrent networks to extract features and generate fingerprints. If a machine learning algorithm is used, then any algorithm can be used, whether unsupervised, supervised, or semi-supervised. The mapping system 100 can use one or more dimensionality reduction methods to generate data for improved visualization of the interface.
[0521] In some embodiments, the mapping system 100 generates a complementary bit map from input molecular data 102 containing antibodies and antigens. In some embodiments, the mapping system 100 generates an epitope map from input molecular data 102 containing antigens.
[0522] In some embodiments, feature extraction can be performed layer by layer using layer-by-layer embedding 202. In layer-by-layer embedding 202, multiple datasets or multiple parts of a dataset may exist, wherein features are extracted individually for each dataset or each part of a dataset. In some embodiments, layer-by-layer embedding includes individual molecular embeddings. Mapping system 100 can perform feature extraction on individual molecules of dataset 106 to obtain individual molecular embeddings.
[0523] In some embodiments, features are used for visualization and can be plotted as overlapping visualization layers. In some embodiments, one or more layers are used to train the feature extraction, and one or more other layers are used to test the feature extraction.
[0524] In some embodiments, the mapping system 100 may use dimensionality reduction to process the dataset 106. The mapping system 100 may use dimensionality reduction to reduce data with a high number of dimensions to a lower-dimensional data representation that can be visualized, while capturing valuable data relationships in the lower-dimensional data representation.
[0525] In some embodiments, the mapping system 100 may include a data storage device for a molecular database, wherein each molecule is assigned a unique index. Molecules assigned the same index have similar sequences, structures, or properties. The mapping system 100 may compare a molecule to the molecular database and assign that molecule an index to the closest or most similar molecule in the database.
[0526] In some embodiments, at 608, the mapping system 100 generates a map user interface 300 containing the molecular map 200 as a lower-dimensional data representation of the dataset 106 and a visualization of the extracted features. This user interface can be any type of interface or platform. It can be implemented as a standalone visualization, a web-based platform, a computer program, an executable (.exe) file, a portable file, or a mobile application. It can run on any platform, any operating system, or any server.
[0527] In some embodiments, the map user interface 300 is one-dimensional, two-dimensional, three-dimensional, four-dimensional, or higher-dimensional. In some embodiments, the embedding of the layer-by-layer embedding 202 is a time series that changes over time. In some embodiments, the three-dimensional embedding may be a time series representing the four-dimensional embedding that changes over time.
[0528] In some embodiments, the user interface 300 may have buttons of any type, including but not limited to regular buttons, radio buttons, check boxes, pop-up menus, tables, scan bars, browse buttons, upload and download buttons, submit buttons, run buttons, zoom and move buttons, selection options, and visualization windows. In some embodiments, the user interface 300 containing the visualization of dataset 106 has control inputs to enable viewing of these layers individually or in relation to other layers. The user interface 300 may have buttons for adding and removing layers embedded 302. Layers can be moved up and down by the user to change the order of the layers, and layers in later orders will be visualized on top of previous layers. Each layer can be hidden or shown by the user. The user interface 300 may have a window for visualization of map 310, and buttons for zooming in or out 312 within the visualization window 310.
[0529] In some embodiments, the map user interface includes a visualization of clusters surrounding molecules of interest. In some embodiments, the map user interface includes one or more clusters of molecules. The user interface 300 may have buttons for adding or removing individual molecules 306 of interest. The user interface 300 may have settings, buttons, and pop-up menus for selecting clusters surrounding individual molecules 318 of interest (such as reference or target molecules). Users can select the shape and size of the clusters along each dimension. Users can also select the location of individual molecules of interest within a cluster. Users may have the option to select multiple molecules of interest. The visualization may indicate clusters that may contain molecules of interest. Molecular visual elements may have surrounding cluster(s).
[0530] Figure 7 This is a schematic diagram of a computing device 700 that can be used to implement various elements of infrastructure system 100 and machine learning system 210. In another example, computing device 700 can be used to implement user equipment 114. As depicted, computing device 700 includes at least one processor 702, memory 704, at least one I / O interface 706, and at least one network interface 708.
[0531] Each processor 702 can be, for example, any type of general-purpose microprocessor or microcontroller, digital signal processing (DSP) processor, integrated circuit, field-programmable gate array (FPGA), reconfigurable processor, programmable read-only memory (PROM), or any combination thereof.
[0532] The memory 704 may include any suitable combination of computer memories, whether internal or external, such as, for example, random access memory (RAM), read-only memory (ROM), optical disc read-only memory (CDROM), electro-optical memory, magneto-optical memory, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), ferroelectric RAM (FRAM), etc.
[0533] Each I / O interface 706 enables the computing device 700 to interconnect with one or more input devices (such as a keyboard, mouse, camera, touchscreen, and microphone) or with one or more output devices (such as a display and speakers).
[0534] Each network interface 708 enables the computing device 700 to communicate with other components, exchange data with other components, access and connect to network resources, service applications, and perform other computing applications by connecting to a network (or multiple networks) capable of carrying data, including the Internet, Ethernet, Common Old-Style Telephone Service (POTS) lines, Public Switched Telephone Network (PSTN), Integrated Services Digital Network (ISDN), Digital Subscriber Line (DSL), coaxial cable, fiber optic, satellite, mobile networks, wireless networks (e.g., Wi-Fi, WiMAX), SS7 signaling networks, fixed lines, local area networks, wide area networks, and other networks, including any combination of these.
[0535] For simplicity, only one computing device 700 is shown, but one or both of system 100 and system 210 may include multiple computing devices 700. The computing devices may be the same or different types of devices. The computing devices 700 may be connected in various ways, including direct coupling, indirect coupling via a network, and distributed over a wide geographical area and connected via a network (which may be referred to as "cloud computing").
[0536] For example, and without limitation, computing device 700 may be a server, network device, embedded device, computer expansion module, personal computer, laptop computer, smartphone device, or any other computing device that can be configured to perform the methods described herein.
[0537] Figure 8 This is a flowchart of another example method 800 for mapping molecules to a visual interface according to an embodiment. At 802, system 100 receives input from various sources. For example, system 100 processes the data to generate dataset 106. At 804, the system generates a complementary bit map. At 806, system 100 receives a selection of one or more molecules or data points on the map. At 810, system 100 may receive a manual selection based on a control command received from the visualization. At 812, system 100 may receive commands that are automatically generated using machine learning and prediction. At 814, an article is generated based on the selected molecule, referred to as an antibody in this example processing.
[0538] Figure 9 This is a schematic diagram illustrating an example representation of antibodies in a complementary site map. Figure 10 This is an example schematic diagram of data processing, showing how the features of the various parts of an antibody are linked together to form the overall characteristics of the antibody. The example visualization is for an antibody. Other molecules can also be represented by visualizations.
[0539] Figure 11This is a diagram of an example map user interface according to an embodiment. The interface has buttons that provide control commands to visualize different layers of the map. The visualization can also depict different fractions or measures of molecules.
[0540] Figure 12 This is a diagram of another example map user interface for a sample according to an embodiment. The example interface illustrates different tools for selecting, editing, deleting, or adding points on a map visualization. This interface can be used for the sample.
[0541] Figure 13 This is a diagram of another example map user interface for selecting candidate antibodies according to an embodiment. This example interface illustrates the selection of a reference antibody and its associated cluster. The interface shows the different groups of molecules and reference antibodies. The interface also shows sampled antibodies.
[0542] Figure 14 This is a flowchart for predicting antibody-antigen interactions. The diagram illustrates an example method for molecular mapping. The process may include a training phase, transfer learning, and a validation phase. This process can generate predictions and learning results.
[0543] Figure 15 This is an example schematic diagram illustrating antibody fingerprint generation. Data associated with different molecules can be processed to generate unique fingerprints.
[0544] Figure 16 This is an example schematic diagram illustrating the generation of complementary bit fingerprints used to compute metrics. This example visualizes complementary bit geometry.
[0545] One can consider patches on the surface of a molecule, where each molecule contains multiple overlapping or non-overlapping patches. For this purpose, for example, one can consider grid vertices on the surface of the molecule, and each vertex can be the center of a patch. Then, for each patch, a feature vector and a complementary feature vector can be obtained using any method such as a machine learning algorithm. In some embodiments, MASIF can be used for these feature vectors. Then, to find molecules similar to a specific molecule based on one or more patches, a search can be performed in a database of molecules, comparing the feature vectors of patches on molecules in the database with the feature vectors of patches(s) of that specific molecule. Similarly, to find complementary molecules to a specific molecule based on one or more patches, a search can be performed in a database of molecules, comparing the feature vectors of patches on molecules in the database with the complementary feature vectors of patches(s) of that specific molecule(s). This can have various applications, such as finding molecules similar to or complementary to a specific molecule. For example, finding complementary molecules to a specific molecule can be useful for docking or binding projects in antibody-antigen co-crystals.
[0546] Figure 10 This is a schematic diagram illustrating data processing.
[0547] This document provides different example use cases. As another example, System 100 could be used in a molecule (e.g., antibody) marketplace where different suppliers can supply or sell related services. Third-party customers can receive selected molecules (e.g., antibodies), and supplier customers can receive a percentage of the associated price. Different companies can sell their products or services based on a complementary bitmap. For example, other antibody suppliers can upload their antibody sequence information to be placed on the complementary bitmap, and customers can choose to test, order, or license it. Bioassay providers can offer their services as an "add-on" to the ordered antibodies. Manufacturers can offer options to produce antibodies in different sizes and qualities. For example, service providers in System 100 can receive a percentage of the transaction fee.
[0548] Figure 23 The illustration depicts an exploration of exemplary immune responses generated in mice with the same genetic background but immunized with different antigenic forms, according to some embodiments.
[0549] Different antigenic forms can include, for example, full-length antigens with regions of interest. Maps generated using the systems and methods described herein can reveal unique and overlapping complementary site reactions, potentially demonstrating how immunization strategies influence the structural diversity and specificity of the generated antibodies, which would otherwise go undetected using conventional methods. This analysis can influence sampling results.
[0550] Complementation site mapping can be used to compare immune responses in mouse strains with different genetic makeup immunized with the same antigen. Although mice may exhibit similar serum antibody titers, the diversity of their complementation site map distributions can indicate differences in immune response patterns. Some mouse strains may be characterized by a broader complementation site distribution compared to others, potentially suggesting unique and distinct repertoire diversity across different strains. This differentiation can highlight how a mouse's genetic background can influence the pattern of generated antibodies, even under similar immunization conditions and drafts.
[0551] The system and method described in this paper enable the identification of structurally similar candidates with potential for similar functional activities. By including reference antibodies with known functions in a complementation site map and sampling an antibody library mapped to the vicinity of these reference antibodies, this platform facilitates the discovery of candidates with comparable properties. Furthermore, diverse sampling within the map facilitates the exploration of antibody efficacy against lower immunogenic epitopes. Using the complementation site map, antibodies can be sampled from novel complementary site clusters as well as from clusters near HER2 reference antibodies. In in vitro binding assays, several antibodies demonstrated superior binding compared to reference antibodies in both high- and low-HER2 expressing cancer cell lines. This sampling strategy can accelerate the identification of novel antibodies that may not have been seen in previous discovery activities.
[0552] These findings highlight the platform's ability to visualize and compare antibody responses across various immunization strategies and genomic profiles. By providing a detailed view of the diversity of the immune repertoire, the platform can support the discovery of novel target conjugates and facilitate the selection of the optimal antibody for clinical applications.
[0553] Example implementation method.
[0554] According to one aspect, a computer-implemented system 100 is provided for mapping molecules to interfaces 300, 400. System 100 includes: a processing subsystem comprising one or more processors and one or more memories coupled to the processors, the processing subsystem being configured such that the system: receives input molecules (e.g., data 102), wherein the input molecules are a set of one or more sets of molecules, each molecule being defined as an information sequence; encodes the information sequence; generates a dataset 106 by processing the input molecules (e.g., using data analysis and preparation 104), wherein the dataset 106 includes the encoded information sequence, three-dimensional coordinates of the information sequence for each molecule defining the molecular structure, and features; transforms the dataset by feature extraction to generate a molecular map 200 and features, thereby reducing a higher number of dimensions to a lower-dimensional data representation that can be indicated by the interface, while capturing valuable data in the lower-dimensional data representation; generates a map user interface 300, 400, which includes the molecular map 200 as a representation of the lower-dimensional data representation of the dataset; and provides the map user interface.
[0555] In some embodiments, map user interfaces 300, 400 include a visual interface, and lower-dimensional data representations can be visualized in the visual interface.
[0556] In some embodiments, the input molecule is an antibody or its antigen-binding fragment and / or antigen, and the molecular map is a complementary site map.
[0557] In some embodiments, the input molecule is an antigen, and the molecular map 200 is an epitope map.
[0558] In some embodiments, the features include fingerprints, wherein the processing subsystem uses machine learning or statistical methods involving one or more of dimensionality reduction, reconstructive autoencoders, variational autoencoders, adversarial autoencoders, neural networks, graph neural networks, attention networks, and recurrent networks to extract features and generate fingerprints.
[0559] In some embodiments, system 100 has a data storage device 118 for a database of molecular data, wherein each molecule is assigned a unique index.
[0560] In some embodiments, system 100 compares a molecule with a database of molecules and assigns the molecule an index to the closest molecule in the database of molecules, wherein molecules assigned to the same index have similar sequences, structures, or properties.
[0561] In some embodiments, feature extraction 202 includes layer-by-layer embedding 204, wherein features of each dataset in the plurality of datasets or features of each part of the plurality of datasets are extracted individually for a plurality of datasets or a plurality of parts of a dataset.
[0562] In some embodiments, features can be drawn as overlapping visual layers.
[0563] In some embodiments, the map user interface 300, 400, which includes the visualization of the dataset 310, has control inputs to enable viewing of layers individually or in relation to other layers.
[0564] In some embodiments, one or more layers are used for training feature extraction 202, and one or more other layers are used for testing feature extraction.
[0565] In some embodiments, feature extraction 202 includes embedding in clusters 206.
[0566] In some embodiments, feature extraction 202 includes embedding individual molecules 208.
[0567] In some embodiments, feature extraction 202 includes generating clusters 210 around molecules of interest.
[0568] In some embodiments, feature extraction 202 includes sampling 214 from a molecular map.
[0569] In some embodiments, feature extraction 202 includes encoding the fractions in the molecular map 216.
[0570] In some embodiments, the map user interface 300, 400 includes visualization of extracted features of the dataset.
[0571] In some embodiments, system 100 has a user device 114 for displaying a map user interface.
[0572] In some embodiments, the map user interface 300, 400 is one-dimensional, two-dimensional, three-dimensional, four-dimensional, or higher-dimensional.
[0573] In some embodiments, the layer-by-layer embedding is a time series that varies over time, and the map user interface includes a two-dimensional or three-dimensional embedding that varies over time as a time series representing a four-dimensional embedding.
[0574] In some embodiments, the map user interface 300, 400 includes visualization of clusters around molecules of interest.
[0575] In some embodiments, the map user interface 300, 400 includes one or more clusters of molecules.
[0576] In some embodiments, the processing subsystem performs feature extraction 202 on each molecule of the dataset to obtain each molecule embedding.
[0577] In some embodiments, layer-by-layer embedding includes individual molecular embedding.
[0578] In some embodiments, user interfaces 300, 400 use the extracted features to characterize and / or obtain information about input molecules, wherein the input molecules may optionally come from clusters of interest.
[0579] In some embodiments, the information includes the extent, nature, and / or robustness of the immune response (against an antigen such as an immunogen or vaccine).
[0580] In some embodiments, the input molecule includes an antibody or an antigen-binding fragment thereof, and the information includes an amino acid sequence of one or more of the antibody or antigen-binding fragment.
[0581] In some embodiments, the input molecule includes an antigen, and the information includes the amino acid sequence of one or more of the antigens.
[0582] In some embodiments, system 100 allows a user to modify the amino acid sequence, and the system predicts the effect of the modification on the characteristics of the molecule (binding, function, stability, expressibility, affinity, immunogenicity, etc.).
[0583] In some embodiments, modifications are made to amino acid substitutions, deletions, and / or additions in one or more CDRs, variable regions, frame regions, and / or constant regions of the antibody or its antigen-binding fragment.
[0584] In some embodiments, the modification is the humanization, deimmunization, glycosylation, or deglycosylation of the antibody or its antigen-binding fragment.
[0585] In some embodiments, system 100 allows users to import other molecules and determine their similarity to the input molecule.
[0586] In some embodiments, the other molecules are other antibodies or their antigen-binding fragments, and the similarity is complementary site similarity.
[0587] In some embodiments, the output includes antibodies or antigen-binding fragments thereof selected from the map or variants thereof.
[0588] In some embodiments, the user synthesizes an input molecule or an output molecule or a variant thereof, or causes an input molecule or an output molecule or a variant thereof to be synthesized.
[0589] In some embodiments, the information provided by the system is used to manufacture molecules.
[0590] In some embodiments, the input molecule includes a single-domain antibody or an antigen-binding fragment thereof, and the output molecule includes an antibody or an antigen-binding fragment thereof selected from conventional antibodies, single-domain antibodies, single-chain variable fragments, humanized antibodies, chimeric antibodies.
[0591] In some embodiments, the processing subsystem connects the features of the parts of the molecule together to have the overall features of the molecule.
[0592] In some embodiments, the antibody or its antigen-binding fragment includes antibodies derived from a species, including but not limited to mice, cattle, rabbits, camels, llamas, humans, alpacas, and standard species.
[0593] In some embodiments, the encoded information sequence refers to the encoded amino acid sequence.
[0594] In some embodiments, the processing subsystem outputs the selected molecule.
[0595] In some embodiments, a system 100 implemented by a computer provides an article obtained from the selected molecular output.
[0596] In some embodiments, a product obtained by a computer-implemented system 100 is provided.
[0597] According to another aspect, a computer-implemented method 600 for mapping molecules to a visual interface is provided. Method 600 involves: receiving input molecules (box 602), wherein the input molecules are one or more groups of one or more molecules, each molecule being defined as an information sequence; encoding the information sequence; generating a dataset 106 by processing the input molecules (box 604), wherein the dataset 106 includes the encoded information sequence, three-dimensional coordinates of the information sequence for each molecule to define the molecular structure, features, and fingerprints; generating a transformed dataset 106 by feature extraction and fingerprinting to generate a molecular map 200, thereby reducing a higher number of dimensions to a lower-dimensional data representation that can be indicated by the interface, while capturing valuable data in the lower-dimensional data representation (box 606); generating map user interfaces 300, 400, which include the molecular map 200 as a representation of the lower-dimensional data representation of the dataset; and providing the map user interfaces 300, 400 (box 608).
[0598] In some embodiments, map user interfaces 300, 400 include a visual interface, and lower-dimensional data representations can be visualized in the visual interface.
[0599] In some embodiments, the input molecule is a protein or protein-like molecule comprising an antibody or its antigen-binding fragment and / or an antigen.
[0600] In some embodiments, the antibody or its antigen-binding fragment includes antibodies derived from a species, including but not limited to mice, cattle, rabbits, camels, llamas, humans, alpacas, and standard species.
[0601] In some embodiments, the input molecule is an antibody or its antigen-binding fragment and / or antigen, and wherein the molecular map 200 is a complementary site map.
[0602] In some embodiments, the input molecule is an antigen, and the molecular map 200 is an epitope map.
[0603] In some embodiments, the features include fingerprints, wherein the processing subsystem uses machine learning or statistical methods involving one or more of dimensionality reduction, reconstructive autoencoders, variational autoencoders, adversarial autoencoders, neural networks, graph neural networks, attention networks, and recurrent networks to extract features and generate fingerprints.
[0604] In some embodiments, method 600 involves storing a database of molecules in a data storage device 118, wherein each molecule is assigned a unique index.
[0605] In some embodiments, method 600 involves comparing a molecule with a database of molecules and assigning the molecule an index to the closest molecule in the database of molecules, wherein molecules assigned the same index have similar sequences, structures, or properties.
[0606] In some embodiments, method 600 involves feature extraction 202 with layer-by-layer embedding 204, wherein features of each of the multiple datasets or features of each of the multiple parts of the dataset are extracted individually for multiple datasets or multiple parts of the dataset.
[0607] In some embodiments, features can be drawn as overlapping visual layers.
[0608] In some embodiments, method 600 involves providing a map user interface 300, 400 that visualizes a dataset 106, having control inputs that enable viewing layers individually or in relation to other layers.
[0609] In some embodiments, one or more layers are used for training feature extraction, and one or more other layers are used for testing feature extraction.
[0610] In some embodiments, feature extraction 202 includes arranging embeddings in clusters 206.
[0611] In some embodiments, feature extraction 202 includes embedding individual molecules 208.
[0612] In some embodiments, feature extraction 202 includes generating clusters 210 around molecules of interest.
[0613] In some embodiments, feature extraction 202 includes sampling 214 from a molecular map.
[0614] In some embodiments, feature extraction 202 includes encoding the fractions in the molecular map 216.
[0615] In some embodiments, the map user interface 300, 400 includes a visualization of the extracted features of the dataset 106.
[0616] In some embodiments, method 600 involves using user device 114 to display map user interface 300, 400.
[0617] In some embodiments, the map user interface 300, 400 is one-dimensional, two-dimensional, three-dimensional, four-dimensional, or higher-dimensional.
[0618] In some embodiments, the layer-by-layer embedding is a time series that changes over time, and the map user interface 300, 400 includes a two-dimensional or three-dimensional embedding that changes over time as a time series representing a four-dimensional embedding.
[0619] In some embodiments, method 600 involves using map user interfaces 300, 400 to provide visualization of clusters around molecules of interest.
[0620] In some embodiments, the map user interface 300, 400 includes one or more clusters of molecules.
[0621] In some embodiments, method 600 involves performing feature extraction 202 on individual molecules of dataset 106 to obtain individual molecule embeddings.
[0622] In some embodiments, layer-by-layer embedding includes individual molecular embedding.
[0623] In some embodiments, method 600 involves using extracted features to characterize and / or obtain information about input molecules, wherein the input molecules may optionally come from clusters of interest.
[0624] In some embodiments, the information includes the extent, nature, and / or robustness of the immune response (against an antigen such as an immunogen or vaccine).
[0625] In some embodiments, the input molecule includes an antibody or an antigen-binding fragment thereof, and the information includes an amino acid sequence of one or more of the antibody or antigen-binding fragment.
[0626] In some embodiments, the input molecule includes an antigen, and the information includes the amino acid sequence of one or more of the antigens.
[0627] In some embodiments, method 600 involves modifying an amino acid sequence, and system 100 predicts the effect of the modification on the characteristics of the molecule (binding, function, stability, expressibility, affinity, immunogenicity, etc.).
[0628] In some embodiments, modifications are made to amino acid substitutions, deletions, and / or additions in one or more CDRs, variable regions, frame regions, and / or constant regions of the antibody or its antigen-binding fragment.
[0629] In some embodiments, the modification is the humanization, deimmunization, glycosylation, or deglycosylation of the antibody or its antigen-binding fragment.
[0630] In some embodiments, method 600 allows a user to import other molecules and determine their similarity to the input molecule.
[0631] In some embodiments, the molecule is another antibody or its antigen-binding fragment, and the similarity is complementary site similarity.
[0632] In some embodiments, the output includes antibodies or antigen-binding fragments thereof selected from the map or variants thereof.
[0633] In some embodiments, the user synthesizes an input molecule or an output molecule or a variant thereof, or causes an input molecule or an output molecule or a variant thereof to be synthesized.
[0634] In some embodiments, method 600 involves using information provided by the system to manufacture molecules.
[0635] In some embodiments, the input molecule includes a single-domain antibody or an antigen-binding fragment thereof, and the output molecule includes an antibody or an antigen-binding fragment thereof selected from conventional antibodies, single-domain antibodies, single-chain variable fragments, humanized antibodies, chimeric antibodies.
[0636] In some embodiments, method 600 involves linking features of portions of a molecule together to have the overall features of the molecule.
[0637] In some embodiments, the encoded information sequence refers to the encoded amino acid sequence.
[0638] In some embodiments, method 600 involves outputting the selected molecule.
[0639] In some embodiments, a method 600 implemented by a computer provides an article obtained from the output of selected molecules.
[0640] In some embodiments, a product obtained by a computer-implemented method 600 is provided.
[0641] In some embodiments, method 600 involves the step of generating molecules identified or selected from map user interfaces 300, 400.
[0642] In some embodiments, a product obtained by a computer-implemented method 600 is provided.
[0643] In some embodiments, the product is an antibody or an antigen-binding fragment thereof.
[0644] According to another aspect, one or more non-transitory computer-readable media having machine-interpretable instructions stored thereon are provided. When executed by a processing subsystem, the machine-interpretable instructions cause the processing subsystem to perform a method 600 for mapping molecules to a visual interface. Method 600 includes: receiving input molecules (box 602), wherein the input molecules are one or more groups of one or more molecules, each molecule being defined as an information sequence; encoding the information sequence; generating a dataset 106 by processing the input molecules (box 604), wherein the dataset 106 includes the encoded information sequence, three-dimensional coordinates of the information sequence for each molecule defining the molecular structure, features, and fingerprints; generating a transformed dataset 106 by feature extraction and fingerprinting to generate a molecular map 200, thereby reducing a higher number of dimensions to a lower-dimensional data representation that can be indicated by the interface, while capturing valuable data in the lower-dimensional data representation (box 606); generating a map user interface 300, 400, which includes the molecular map 200 as a representation of the lower-dimensional data representation of the dataset (box 608); and providing the map user interface.
[0645] According to another aspect, a computer-implemented system 100 for a molecule-related interface is provided. System 100 includes: a processing subsystem comprising one or more processors and one or more memories coupled to the processors, the processing subsystem being configured such that the system: receives input molecules (e.g., data 102), wherein the input molecules are one or more groups of one or more molecules, each molecule being defined as an information sequence; encodes the information sequence; generates a dataset 106 by processing the input molecules, wherein the dataset 106 includes the encoded information sequence, three-dimensional coordinates of the information sequence for each molecule defining the molecular structure, features, and fingerprints; transforms the dataset 106 by feature extraction and fingerprint generation, thereby reducing a higher number of dimensions to a lower-dimensional data representation that can be indicated by the interface, while capturing valuable data in the lower-dimensional data representation; generates one or more metrics from the transformed dataset 106, wherein the one or more metrics include the lower-dimensional data representation of the dataset 106 and summarize the characteristics of the input molecules; and provides one or more metrics to interfaces 300, 400.
[0646] According to another aspect, a computer-implemented system 100 for a visual interface 300, 400 for mapping molecules is provided. System 100 includes: a processing subsystem comprising one or more processors and one or more memories coupled to the processors; the processing subsystem providing a map user interface 300, 400, wherein the map user interface 300, 400: receives input molecules (e.g., data 102), wherein the input molecules are one or more groups of one or more molecules, wherein each molecule is defined as an information sequence; and provides a representation of a lower-dimensional data representation of a dataset 106 containing a molecular map 200 as the input molecules, wherein the dataset 106 includes the information sequence of each molecule, three-dimensional coordinates of the information sequence for defining the molecular structure of each molecule, and features; wherein the molecular map 200 includes a transformation of the dataset 106 through feature extraction 202 and fingerprint generation, thereby reducing a higher number of dimensions to a lower-dimensional data representation that can be indicated by the interface, while capturing valuable data in the lower-dimensional data representation.
[0647] In some embodiments, the input molecule is a protein or protein-like molecule that includes antibodies and antigens.
[0648] In some embodiments, antibodies include species-specific antibodies, including but not limited to mice, cattle, rabbits, camels, llamas, humans, alpacas, and standard species.
[0649] In some embodiments, the input molecules are antibodies and antigens, and the molecular map is a complementary site map.
[0650] In some embodiments, the input molecule is an antigen, and the molecular map is an epitope map.
[0651] In some embodiments, the map user interface 300, 400 includes layer-by-layer embedding to provide layers for map visualization.
[0652] In some embodiments, the map user interface 300, 400 renders features as overlaid map visualization layers.
[0653] In some embodiments, the map user interface 300, 400 has control inputs that enable viewing layers individually or in relation to other layers.
[0654] In some embodiments, the map user interface 300, 400 has control inputs to add or remove layers in the layers used for map visualization.
[0655] In some embodiments, the map user interface 300, 400 includes a visualization of the extracted features of the dataset 106.
[0656] In some embodiments, map user interfaces 300, 400 receive one or more reference molecules or target molecules, wherein the transformation of dataset 106 is based on one or more reference molecules or target molecules.
[0657] In some embodiments, map user interfaces 300, 400 receive one or more scores of molecules, wherein the scores include expressibility scores and fuzzy filtering scores.
[0658] In some embodiments, map visualization includes one or more clusters corresponding to molecules, wherein map user interfaces 300, 400 receive clustering control commands to update the map visualization with hyperparameters of one or more clusters.
[0659] In some embodiments, the map visualization displays one or more scores about the molecule, including expressibility scores or fuzzy screening scores.
[0660] In some embodiments, map user interfaces 300, 400 receive control commands for sampling from at least a portion of a map visualization.
[0661] In some embodiments, map user interfaces 300, 400 receive control commands for editing samples extracted from at least a portion of a map visualization.
[0662] In some embodiments, the map user interface 300, 400 receives drawing settings corresponding to the visualization characteristics of the map visualization.
[0663] In some embodiments, user equipment 114 may display a map user interface.
[0664] In some embodiments, the map user interface 300, 400 is one-dimensional, two-dimensional, three-dimensional, four-dimensional, or higher-dimensional.
[0665] In some embodiments, the layer-by-layer embedding is a time series that varies over time, and the map user interface includes a two-dimensional or three-dimensional embedding that varies over time as a time series representing a four-dimensional embedding.
[0666] In some embodiments, the map user interface 300, 400 includes visualization of clusters around molecules of interest.
[0667] In some embodiments, the map user interface 300, 400 includes one or more clusters of molecules.
[0668] In some embodiments, the map user interface 300, 400 includes various molecular embeddings.
[0669] In some embodiments, layer-by-layer embedding includes individual molecular embedding.
[0670] In some embodiments, the map user interface 300, 400 receives scores.
[0671] In some embodiments, the map user interface 300, 400 exports files.
[0672] In some embodiments, the map user interface 300, 400 includes one or more buttons for adding or removing layers 302, one or more buttons for receiving input molecules 304, one or more buttons for adding or removing individual molecules 306, and one or more buttons for importing fractions 308.
[0673] In some embodiments, the map user interface 300, 400 includes multiple settings selected from the group consisting of: navigation settings for map visualization 312, drawing settings 314, settings 316 for encoding scores in map visualization, clustering settings 324, settings 330 for editing samples, sample settings 328, reporting settings 322, map analysis settings 320, and export settings.
[0674] According to one aspect, a computer-implemented system 100 is provided for mapping proteins, protein-like molecules, or fragments thereof onto a visual interface and generating maps for display and interaction. System 100 includes a processing subsystem comprising one or more processors and one or more memories coupled to the processors. The processing subsystem is configured such that the system: receives one or more sets of one or more sets of input data 102 of proteins, protein-like molecules, or fragments thereof, wherein the data includes features of one or more proteins, protein-like molecules, or fragments thereof; generates at least one dataset 106 by processing one or more sets of one or more sets of input data of proteins, protein-like molecules, or fragments thereof (e.g., using data analysis and preparation 104); and transforms the data and one or more of the at least one dataset by feature extraction 202 or feature selection to generate maps and additional features, thereby reducing a higher-dimensional data representation to a lower-dimensional data representation for visualization via a visual interface, wherein the lower-dimensional data representation captures valuable information from the data and one or more of the at least one dataset 106, and the lower-dimensional data representation includes one or more clusters of proteins, protein-like molecules, or fragments thereof. Proteins, protein-like molecules, or fragments thereof include one or more input proteins, protein-like molecules, or fragments thereof, or generated proteins, protein-like molecules, or fragments thereof; generating a visual map interface, the visual map interface including a map 200 as a visual representation of the lower-dimensional data representation of the dataset, the visual representation including a visualization of the dataset 106 as one or more layers of proteins, protein-like molecules, or fragments thereof, each layer including one or more of one or more clusters of proteins, protein-like molecules, or fragments thereof; and providing tools to the visual map interfaces 300 and 400 for interacting with the map 200, wherein the interaction with the map 200 includes one or more of the following: inspection, search, sampling, clustering, and analysis of one or more proteins, protein-like molecules, or fragments thereof, or newly generated proteins, protein-like molecules, or fragments thereof; receiving commands or detecting interactions with the map at the visual map interface via tools; updating the map 200 based on commands or interactions; and triggering an update of the visual map interface using the updated map 200.
[0675] In some embodiments, the protein, protein-like molecule or fragment thereof is selected from: antibody, antigen, lectin, receptor, ligand, enzyme or fragment thereof.
[0676] In some embodiments, the protein, protein-like molecule or fragment thereof includes an antibody or fragment thereof and / or an antigen or fragment thereof, and wherein map 200 is a complementary site map or an epitope map or a map 200 including a protein or protein-like molecule or fragment thereof.
[0677] In some embodiments, proteins, protein-like molecules, or fragments thereof include antibodies or antibody fragments thereof, and wherein the data includes the structure of one or more antibodies or antibody fragments thereof, the amino acid sequence of one or more of the antibodies or antibody fragments thereof, the amino acid atomic or molecular coordinates, and one or more of the biophysical properties of one or more antibodies or antibody fragments thereof.
[0678] In some embodiments, the protein, protein-like molecule, or fragment thereof is selected from conventional antibodies, antibody-like molecules, artificial antibodies, antibody mimics, single-domain antibodies, single-chain antibodies, humanized antibodies, chimeric antibodies, or fragments thereof.
[0679] In some embodiments, the fragment includes an antigen-binding fragment or an antigen-binding domain.
[0680] In some embodiments, the antigen-binding fragment or antigen-binding domain is selected from one or more complementarity-determining regions and / or one or more frame regions, one or more variable domains, or complementary sites.
[0681] In some embodiments, the visual representation of lower-dimensional data includes different colors and / or marker shapes and / or marker sizes and / or color transparency and / or color gradients to indicate one or more layers and one or more clusters of proteins, protein-like molecules or fragments thereof.
[0682] In some embodiments, the lower-dimensional data representation is a one-dimensional, two-dimensional, three-dimensional, or four-dimensional data representation.
[0683] In some embodiments, the features include fingerprints, wherein the processing subsystem uses machine learning or statistical methods involving one or more of dimensionality reduction, reconstructive autoencoders, variational autoencoders, adversarial autoencoders, neural networks, graph neural networks, attention networks, recurrent networks, sequence processing algorithms, image processing algorithms, computer vision algorithms, and identity transformations to extract features and generate fingerprints.
[0684] In some embodiments, the processing subsystem causes system 100 to cluster one or more of the data and dataset 106 to generate one or more clusters of proteins, protein-like molecules or fragments thereof.
[0685] In some embodiments, the processing subsystem causes system 100 to encode the raw data and generate additional features from the encoded data.
[0686] In some embodiments, the visual representation overlays one or more layers of proteins, protein-like molecules, or fragments thereof as overlays as part of the visualization of the dataset 106, wherein tools trigger one or more layers to move to different locations or levels, or remove them from the map, or change the order of the displayed layers, or zoom in or out of one or more layers, or move across layers in the map.
[0687] In some embodiments, the processing subsystem enables system 100 to perform map analysis, wherein map analysis includes one or more of the following: generating clusters around proteins, protein-like molecules or fragments of interest, arranging embeddings in the clusters, embedding layer by layer, embedding individual proteins, protein-like molecules or fragments of them, sampling from the map, encoding scores in the map, wherein the map contains one or more clusters and visualizing one or more clusters.
[0688] In some embodiments, the processing subsystem causes system 100 to generate or calculate one or more clusters of proteins, protein-like molecules or fragments thereof, and wherein the map user interface includes visualization of one or more clusters of proteins, protein-like molecules or fragments thereof.
[0689] In some embodiments, feature extraction 202 includes extracting useful information from a dataset, and feature selection includes selecting a subset of the dataset of proteins, protein-like molecules, or fragments thereof.
[0690] In some embodiments, the processing subsystem causes system 100 to transform one or more of data and datasets 106 to generate map 200 by one or more of sequencing and clustering, sampling, intersection of data subsets, and subtraction of data subsets.
[0691] In some embodiments, the processing subsystem causes the system 100 to partition or segment the digital map 200 into multiple map tiles, label each of one or more clusters with corresponding map tiles in the multiple map tiles, and display one or more clusters within the multiple map tiles using labels, wherein the visualization indicates the multiple map tiles and one or more clusters.
[0692] In some embodiments, the processing subsystem causes system 100 to: (i) intersect one or more layers of proteins, protein-like molecules, or fragments thereof; or (ii) subtract one or more layers of proteins, protein-like molecules, or fragments thereof; or (iii) add one or more layers of proteins, protein-like molecules, or fragments thereof to update the map based on commands or interactions.
[0693] In some embodiments, the tools at the visual map interface 300, 400 include sampling tools for sampling proteins, protein-like molecules or fragments of proteins from one or more clusters of proteins, protein-like molecules or fragments thereof, wherein the processing subsystem causes the system to update the map by sampling proteins, protein-like molecules or fragments thereof in response to activation of the sampling tools, and triggers an update of the visual map interface using the updated map to visualize the sampling.
[0694] In some embodiments, the processing subsystem causes system 100 to subtract a library of non-immunized proteins, protein-like molecules, or fragments thereof from an immunized library of proteins, protein-like molecules, or fragments thereof to filter out non-specific proteins, protein-like molecules, or fragments thereof, and to reduce the search space used for sampling and searching for specific molecular candidates for one or more targets. If multiple layers or datasets exist for the immunized library, the subsystem causes the system to intersect the layers or datasets after the subtraction to further reduce the search space.
[0695] In some embodiments, the processing subsystem causes system 100 to subtract a library of proteins, protein-like molecules, or fragments thereof immunized against one or more targets from a library of proteins, protein-like molecules, or fragments thereof immunized against a target of interest, in order to filter out non-binding portions of the target of interest and reduce the search space used for sampling and searching for specific molecular candidates against the target of interest. If there are multiple layers or datasets in the immune library against the target of interest, the subsystem causes the system to intersect the layers or datasets after the subtraction, in order to further reduce the search space.
[0696] In some embodiments, the processing subsystem enables system 100 to export or report inspections, searches, sampling, clustering, and analyses of proteins, protein-like molecules, or fragments thereof through text, tables, graphs, or visualizations.
[0697] According to one aspect, a computer processing method 600 is provided for mapping proteins, protein-like molecules, or fragments thereof onto a visual interface and generating digital maps for display and interaction. The method 600 includes: receiving one or more sets of one or more input data 102 of proteins, protein-like molecules, or fragments thereof (box 602), wherein the data 102 includes features of one or more proteins, protein-like molecules, or fragments thereof; generating at least one dataset 106 by processing one or more sets of one or more input data of proteins, protein-like molecules, or fragments thereof (box 604); transforming the data and one or more of the at least one dataset 106 by feature extraction or feature selection to generate a map and additional features, thereby reducing a higher-dimensional data representation to a lower-dimensional data representation for visualization via a visual interface (box 606), wherein the lower-dimensional data representation captures valuable information from the data and one or more of the at least one dataset 106, and the lower-dimensional data representation includes one or more clusters of proteins, protein-like molecules, or fragments thereof, and the proteins, protein-like molecules, or fragments thereof. This includes one or more input proteins, protein-like molecules or fragments thereof, or generated proteins, protein-like molecules or fragments thereof; generating a visual map interface, which includes a map as a visual representation of the lower-dimensional data of the dataset, the visual representation including a visualization (box 608) representing the dataset 106 as one or more layers of proteins, protein-like molecules or fragments thereof, each layer including one or more clusters of proteins, protein-like molecules or fragments thereof; providing the visual map interface with tools for interacting with the map, wherein the interaction with the map includes one or more of the following: inspection, searching, sampling, clustering and analysis of one or more proteins, protein-like molecules or fragments thereof or newly generated proteins, protein-like molecules or fragments thereof; receiving commands or detecting interactions with the map through tools at the visual map interface; and triggering updates to the visual map interface and the map based on commands or interactions.
[0698] According to one aspect, a computer-readable medium encoded with instructions 600 is provided, which, when executed by a processor, cause the processor to map proteins, protein-like molecules, or fragments thereof onto a visual interface and generate a digital map 200 for display and interaction. Instructions 600 include instructions for performing the following: receiving one or more sets of one or more sets of input data 102 of proteins, protein-like molecules, or fragments thereof (box 602), wherein data 102 includes features of one or more proteins, protein-like molecules, or fragments thereof; generating at least one dataset 106 by processing one or more sets of one or more sets of input data of proteins, protein-like molecules, or fragments thereof (box 604); transforming one or more of data 102 and at least one dataset 106 by feature extraction or feature selection to generate map 200 and additional features, thereby reducing a higher-dimensional data representation to a lower-dimensional data representation for visualization via a visual interface (box 606), wherein the lower-dimensional data representation captures valuable information from the data and one or more of at least one dataset, and the lower-dimensional data representation includes one or more clusters of proteins, protein-like molecules, or fragments thereof, wherein the proteins, protein-like molecules, or fragments thereof include a The system can take one or more input proteins, protein-like molecules or fragments thereof, or generated proteins, protein-like molecules or fragments thereof; generate visual map interfaces 300 and 400, which include a map 200 (box 608) as a visual representation of the lower-dimensional data of dataset 106. The visual representation includes a visualization of dataset 106 as one or more layers of proteins, protein-like molecules or fragments thereof, each layer including one or more clusters of proteins, protein-like molecules or fragments thereof; provide tools to the visual map interfaces 300 and 400 for interacting with the map, wherein interaction with the map 200 includes one or more of the following: inspection, search, sampling, clustering and analysis of one or more proteins, protein-like molecules or fragments thereof or newly generated proteins, protein-like molecules or fragments thereof; receive commands or detect interactions with the map through tools at the visual map interfaces; and trigger updates to the visual map interfaces and the map based on commands or interactions.
[0699] According to one aspect, a computer-implemented system 100 is provided for mapping molecules onto an interface and generating a map 200 for interfaces 300, 400. System 100 includes a processing subsystem comprising one or more processors and one or more memories coupled to the processors. The processing subsystem is configured such that system 100: receives data 102 of one or more sets of input molecules, wherein the input molecules are one or more sets of one or more molecules, and wherein the data 102 includes features of the molecules; generates at least one dataset 106 by processing the data 102 of one or more sets of input molecules, wherein the dataset 106 includes an encoded sequence of information, coordinates for each molecule to define its molecular structure, and features; and transforms the data 102 and at least [dataset 106] by feature extraction or feature selection. One or more of a dataset 106 are used to generate maps and additional features, thereby reducing a higher-dimensional data representation to a lower-dimensional data representation that can be indicated, depicted, or visualized by an interface, while capturing valuable information in the lower-dimensional data representation, which includes one or more clusters of input molecules or newly generated molecules; generating map user interfaces 300 and 400, the map user interface including a map 200 as a representation of the lower-dimensional data representation of dataset 106, representing at least one dataset 106 as one or more molecular layers, each layer including one or more of one or more clusters of molecules; and providing a map user interface.
[0700] In some embodiments, the processing subsystem enables system 100 to: provide map user interfaces 300, 400 with tools for interacting with the map and for inspecting, searching, sampling, clustering, and analyzing one or more input molecules or newly generated molecules; receive commands or detect interactions with map 200 via tools at the visual map interfaces 300, 400; update map 200 based on commands or interactions; and trigger updates to map interfaces 300, 400 using the updated map 200.
[0701] In some embodiments, the molecule is a protein, a protein-like molecule, a fragment thereof, a small molecule drug, or a nucleic acid molecule.
[0702] In some embodiments, the input protein, protein-like molecule or fragment thereof includes an antibody or fragment thereof and / or an antigen or fragment thereof, and wherein the map is a complementary site map or an epitope map or a map including a protein or protein-like molecule or fragment thereof.
[0703] In some embodiments, proteins, protein-like molecules or fragments thereof are selected from the group consisting of antibodies, antigen-binding fragments, drug candidates, compounds, candidate conjugates and binding agents.
[0704] In some embodiments, map user interfaces 300, 400 include a visual interface, and lower-dimensional data representations can be visualized in the visual interface.
[0705] In some embodiments, the input molecule is a protein or protein-like molecule comprising an antibody or its antigen-binding fragment and / or an antigen.
[0706] In some embodiments, the antibody or its antigen-binding fragment includes antibodies derived from a species, including but not limited to mice, cattle, rabbits, camels, llamas, humans, alpacas, and standard species.
[0707] In some embodiments, the input molecule is an antibody or its antigen-binding fragment and / or antigen, and wherein the molecular map 200 is a complementary site map.
[0708] In some embodiments, the input molecule is an antigen, and the molecular map 200 is an epitope map.
[0709] In some embodiments, the features include fingerprints, wherein the processing subsystem uses machine learning or statistical methods involving one or more of dimensionality reduction, reconstructive autoencoders, variational autoencoders, adversarial autoencoders, neural networks, graph neural networks, attention networks, and recurrent networks to extract features and generate fingerprints.
[0710] In some embodiments, system 100 further includes a data storage device 118 for a database of molecules, wherein each molecule is assigned a unique index.
[0711] In some embodiments, system 100 compares a molecule with a database of molecules and assigns the molecule an index to the closest molecule in the database of molecules, wherein molecules assigned to the same index have similar sequences, structures, or properties.
[0712] In some embodiments, feature extraction 202 includes layer-by-layer embedding, wherein features of each dataset in the plurality of datasets or features of each part of the plurality of datasets are extracted individually for a plurality of datasets or a plurality of parts of a dataset.
[0713] In some embodiments, features can be drawn as overlapping visual layers.
[0714] In some embodiments, the visualization map user interface 300, 400 including dataset 106 has control inputs to enable viewing of layers individually or in relation to other layers.
[0715] In some embodiments, one or more layers are used for training feature extraction, and one or more other layers are used for testing feature extraction 202.
[0716] In some embodiments, feature extraction 202 includes arranging embeddings in clusters.
[0717] In some embodiments, feature extraction 202 includes embedding individual molecules.
[0718] In some embodiments, feature extraction 202 includes generating clusters around molecules of interest.
[0719] In some embodiments, feature extraction 202 includes sampling from molecular map 200.
[0720] In some embodiments, feature extraction 202 includes encoding fractions in the molecular map 200.
[0721] In some embodiments, the map user interface 300, 400 includes visualization of extracted features of the dataset 106.
[0722] In some embodiments, a user device 114 for displaying a map user interface 300, 400 is also included.
[0723] In some embodiments, the map user interface 300, 400 is one-dimensional, two-dimensional, three-dimensional, four-dimensional, or higher-dimensional.
[0724] In some embodiments, the layer-by-layer embedding is a time series that varies over time, and the map user interface includes a two-dimensional or three-dimensional embedding that varies over time as a time series representing a four-dimensional embedding.
[0725] In some embodiments, the map user interface 300, 400 includes visualization of clusters around molecules of interest.
[0726] In some embodiments, the map user interface 300, 400 includes one or more clusters of molecules.
[0727] In some embodiments, the processing subsystem performs feature extraction 202 on each molecule of the dataset to obtain each molecule embedding.
[0728] In some embodiments, layer-by-layer embedding includes individual molecular embedding.
[0729] In some embodiments, user interfaces 300, 400 use extracted features to characterize and / or obtain information about input molecules, wherein the input molecules may optionally come from clusters of interest.
[0730] In some embodiments, the information includes the extent, nature, and / or robustness of the immune response (against an antigen such as an immunogen or vaccine).
[0731] In some embodiments, the input molecule includes an antibody or an antigen-binding fragment thereof, and the information includes an amino acid sequence of one or more of the antibody or antigen-binding fragment.
[0732] In some embodiments, the input molecule includes an antigen, and the information includes the amino acid sequence of one or more of the antigens.
[0733] In some embodiments, system 100 allows a user to modify the amino acid sequence, and the system predicts the effect of the modification on the characteristics of the molecule (binding, function, stability, expressibility, affinity, immunogenicity).
[0734] In some embodiments, modifications are made to amino acid substitutions, deletions, and / or additions in one or more CDRs, variable regions, frame regions, and / or constant regions of the antibody or its antigen-binding fragment.
[0735] In some embodiments, the modification is the humanization, deimmunization, glycosylation, or deglycosylation of the antibody or its antigen-binding fragment.
[0736] In some embodiments, system 100 allows users to import other molecules and determine their similarity to the input molecule.
[0737] In some embodiments, the other molecules are other antibodies or their antigen-binding fragments, and the similarity is complementary site similarity.
[0738] In some embodiments, the output includes antibodies or antigen-binding fragments thereof selected from the map or variants thereof.
[0739] In some embodiments, the user synthesizes an input molecule or an output molecule or a variant thereof, or causes an input molecule or an output molecule or a variant thereof to be synthesized.
[0740] In some embodiments, information provided by the system is used to manufacture molecules.
[0741] In some embodiments, the input molecule includes a single-domain antibody or an antigen-binding fragment thereof, and the output molecule includes an antibody or an antigen-binding fragment thereof selected from conventional antibodies, single-domain antibodies, single-chain variable fragments, humanized antibodies, or chimeric antibodies.
[0742] In some embodiments, the processing subsystem connects the features of the parts of the molecule together to have the overall features of the molecule.
[0743] In some embodiments, the encoded information sequence refers to the encoded amino acid sequence.
[0744] In some embodiments, the processing subsystem outputs the selected molecule.
[0745] According to one aspect, a system 100 implemented by a computer as described herein provides an article obtained from the selected molecular output.
[0746] According to one aspect, a product obtained by a system 100 implemented by a computer as described herein is provided.
[0747] According to one aspect, a computer-implemented method 600 is provided for mapping molecules onto an interface and generating a map 200 for interfaces 300, 400. Method 600 includes: receiving data 102 (box 602) of one or more sets of one or more input molecules, wherein the input molecules are one or more sets of one or more molecules, and wherein the data 102 includes features of the molecules; generating at least one dataset 106 (box 604) by processing the data of one or more sets of one or more input molecules, wherein the dataset 106 includes an encoded sequence of information, coordinates for each molecule to define its molecular structure, features, and fingerprints; and transforming one or more of the data 102 and the at least one dataset 106 by feature extraction 202, feature selection, or fingerprint generation to generate a map 200 and additional features. This reduces the higher-dimensional data representation to a lower-dimensional data representation that can be indicated, depicted, or visualized by the interface (box 606), while capturing valuable information in the lower-dimensional data representation, which includes one or more clusters of input molecules or newly generated molecules; generates map user interfaces 300 and 400, which include a map 200 (box 608) representing the lower-dimensional data representation of the dataset, indicating that at least one dataset 106 is represented as one or more molecular layers, each layer including one or more of one or more clusters of molecules; and provides map user interfaces 300 and 400.
[0748] In some embodiments, map user interfaces 300, 400 include a visual interface, and lower-dimensional data representations can be visualized in the visual interface.
[0749] In some embodiments, the input molecule is a protein or protein-like molecule comprising an antibody or its antigen-binding fragment and / or an antigen.
[0750] In some embodiments, the antibody or its antigen-binding fragment includes antibodies derived from a species, including but not limited to mice, cattle, rabbits, camels, llamas, humans, alpacas, and standard species.
[0751] In some embodiments, the input molecule is an antibody or its antigen-binding fragment and / or antigen, and wherein the molecular map 200 is a complementary site map.
[0752] In some embodiments, the input molecule is an antigen, and the molecular map 200 is an epitope map.
[0753] In some embodiments, the features include fingerprints, wherein the processing subsystem uses machine learning or statistical methods involving one or more of dimensionality reduction, reconstructive autoencoders, variational autoencoders, adversarial autoencoders, neural networks, graph neural networks, attention networks, and recurrent networks to extract features 202 and generate fingerprints.
[0754] In some embodiments, method 600 further includes storing a database of molecules in data storage device 118, wherein each molecule is assigned a unique index.
[0755] In some embodiments, method 600 further includes comparing the molecule with a database of molecules and assigning the molecule an index to the closest molecule in the database of molecules, wherein molecules assigned the same index have similar sequences, structures, or properties.
[0756] In some embodiments, method 600 further includes feature extraction 202 with layer-by-layer embedding, wherein features of each of the multiple datasets or features of each of the multiple parts of the dataset are extracted individually for the multiple datasets 106 or multiple parts of the dataset 106.
[0757] In some embodiments, features can be drawn as overlapping visual layers.
[0758] In some embodiments, method 600 further includes providing a map user interface 300, 400 that visualizes the dataset 106, having control inputs to enable viewing of layers individually or in relation to other layers 302.
[0759] In some embodiments, one or more layers are used for training feature extraction 202, and one or more other layers are used for testing feature extraction 202.
[0760] In some embodiments, feature extraction 202 includes embedding in clusters 206.
[0761] In some embodiments, feature extraction 202 includes embedding individual molecules 208.
[0762] In some embodiments, feature extraction 202 includes generating clusters 210 around molecules of interest.
[0763] In some embodiments, feature extraction 202 includes sampling 214 from a molecular map.
[0764] In some embodiments, feature extraction 202 includes encoding the fractions in the molecular map 216.
[0765] In some embodiments, the map user interface 300, 400 includes visualization of extracted features of the dataset 106.
[0766] In some embodiments, method 600 further includes using user device 118 to display map user interface 300, 400.
[0767] In some embodiments, the map user interface 300, 400 is one-dimensional, two-dimensional, three-dimensional, four-dimensional, or higher-dimensional.
[0768] In some embodiments, the layer-by-layer embedding is a time series that varies over time, and the map user interface includes a two-dimensional or three-dimensional embedding that varies over time as a time series representing a four-dimensional embedding.
[0769] In some embodiments, method 600 further includes using map user interfaces 300, 400 to provide visualization of clusters around molecules of interest.
[0770] In some embodiments, the map user interface 300, 400 includes one or more clusters of molecules.
[0771] In some embodiments, method 600 further includes performing feature extraction 202 on each molecule 208 of dataset 106 to obtain each molecule embedding.
[0772] In some embodiments, layer-by-layer embedding includes individual molecular embedding.
[0773] In some embodiments, method 600 further includes using extracted features to characterize and / or obtain information about input molecules, wherein the input molecules may optionally come from clusters of interest.
[0774] In some embodiments, the information includes the extent, nature, and / or robustness of the immune response (against an antigen such as an immunogen or vaccine).
[0775] In some embodiments, the input molecule includes an antibody or an antigen-binding fragment thereof, and the information includes an amino acid sequence of one or more of the antibody or antigen-binding fragment.
[0776] In some embodiments, the input molecule includes an antigen, and the information includes the amino acid sequence of one or more of the antigens.
[0777] In some embodiments, method 600 further includes modifying the amino acid sequence, and wherein the system predicts the effect of the modification on the characteristics of the molecule (binding, function, stability, expressibility, affinity, immunogenicity, etc.).
[0778] In some embodiments, modifications are made to amino acid substitutions, deletions, and / or additions in one or more CDRs, variable regions, frame regions, and / or constant regions of the antibody or its antigen-binding fragment.
[0779] In some embodiments, the modification is the humanization, deimmunization, glycosylation, or deglycosylation of the antibody or its antigen-binding fragment.
[0780] In some embodiments, system 100 allows users to import other molecules and determine their similarity to the input molecule.
[0781] In some embodiments, the other molecules are other antibodies or their antigen-binding fragments, and the similarity is complementary site similarity.
[0782] In some embodiments, the output includes antibodies or antigen-binding fragments thereof selected from the map or variants thereof.
[0783] In some embodiments, the user synthesizes an input molecule or an output molecule or a variant thereof, or causes the input molecule or an output molecule or a variant thereof to be synthesized.
[0784] In some embodiments, method 600 further includes using information provided by the system to manufacture molecules.
[0785] In some embodiments, the input molecule includes a single-domain antibody or an antigen-binding fragment thereof, and the output molecule includes an antibody or an antigen-binding fragment thereof selected from conventional antibodies, single-domain antibodies, single-chain variable fragments, humanized antibodies, chimeric antibodies.
[0786] In some embodiments, method 600 further includes linking features of the portions of the molecule together to have the overall features of the molecule.
[0787] In some embodiments, the encoded information sequence refers to the encoded amino acid sequence.
[0788] In some embodiments, method 600 further includes outputting the selected molecule.
[0789] According to one aspect, a method 600 implemented by a computer as described herein provides an article obtained from the output of a selected molecule.
[0790] According to one aspect, a product obtained by a computer-implemented method 600 described herein is provided.
[0791] In some embodiments, method 600 includes the step of generating molecules identified or selected from map user interfaces 300, 400.
[0792] According to one aspect, a product obtained by a computer-implemented method 600 described herein is provided.
[0793] In some embodiments, the product is an antibody or an antigen-binding fragment thereof.
[0794] According to one aspect, one or more non-transitory computer-readable media having machine-interpretable instructions stored thereon are provided, which, when executed by a processing subsystem, cause the processing subsystem to perform a method 600 for mapping molecules onto a visual interface and generating a map 200 for the interface. Method 600 includes receiving data 102 (box 602) of one or more sets of one or more input molecules, wherein the input molecules are one or more sets of one or more molecules, and wherein the data 102 includes features of the molecules; generating at least one dataset 106 (box 604) by processing the data 102 of one or more sets of one or more input molecules, wherein the dataset 106 includes an encoded sequence of information, coordinates for each molecule to define its molecular structure, features, and fingerprints; and transforming one or more of the data 102 and the at least one dataset 106 by feature extraction 202, feature selection, or fingerprint generation to generate map 200. The system includes additional features to reduce the higher-dimensional data representation to a lower-dimensional data representation that can be indicated by the interface, while capturing valuable information in the lower-dimensional data representation (box 606), which includes one or more clusters of input molecules or newly generated molecules; generates map user interfaces 300 and 400, which include a map 200 representing the lower-dimensional data representation of dataset 106, representing at least one dataset 106 as one or more molecular layers, each layer including one or more of one or more clusters of molecules; and provides map user interfaces 300 and 400.
[0795] According to one aspect, a computer-implemented system 100 for a molecule-related interface is provided. System 100 includes: a processing subsystem comprising one or more processors and one or more memories coupled to the processors, the processing subsystem being configured such that system 100: receives input molecules (e.g., as data 102), wherein the input molecules are a set of one or more sets of molecules, each molecule being defined as an information sequence; encodes the information sequence; generates a dataset 106 by processing the input molecules, wherein dataset 106 includes the encoded information sequence, three-dimensional coordinates of the information sequence for each molecule defining the molecular structure, features, and fingerprints; generates a transformed dataset 106 by feature extraction 202 and fingerprints, thereby reducing a higher number of dimensions to a lower-dimensional data representation that can be indicated by the interface, while capturing valuable data in the lower-dimensional data representation; generates one or more metrics from the transformed dataset 106, wherein the one or more metrics include the lower-dimensional data representation of dataset 106 and summarize the characteristics of the input molecules; and provides one or more metrics to the interface.
[0796] According to one aspect, a computer-implemented system 100 for a visual interface for mapping molecules is provided. System 100 includes a processing subsystem comprising one or more processors and one or more memories coupled to the processors. The processing subsystem provides a map user interface 300, 400, wherein the map user interface 300, 400: receives input molecules (e.g., as data 102), wherein the input molecules are one or more groups of one or more molecules, wherein each molecule is defined as an information sequence; and provides the map interface 300, 400, which includes a representation of a molecular map as a lower-dimensional data representation of the dataset of input molecules, wherein the dataset includes the information sequence of each molecule, three-dimensional coordinates of the information sequence for defining the molecular structure of each molecule, and features; wherein the molecular map 200 includes a transformed dataset 106 generated by feature extraction 202 and fingerprint generation, thereby reducing a higher number of dimensions to a lower-dimensional data representation that can be indicated by the interface, while capturing valuable data in the lower-dimensional data representation.
[0797] In some embodiments, the input molecule is a protein or protein-like molecule that includes antibodies and antigens.
[0798] In some embodiments, antibodies include species-specific antibodies, including but not limited to mice, cattle, rabbits, camels, llamas, humans, alpacas, and standard species.
[0799] In some embodiments, the input molecules are antibodies and antigens, and the molecular map is a complementary site map.
[0800] In some embodiments, the input molecule is an antigen, and the molecular map is an epitope map.
[0801] In some embodiments, the map user interface 300, 400 includes layer-by-layer embedding to provide layers for map visualization.
[0802] In some embodiments, the map user interface 300, 400 renders features as overlaid map visualization layers.
[0803] In some embodiments, the map user interface 300, 400 has control inputs that enable viewing layers individually or in relation to other layers.
[0804] In some embodiments, map user interfaces 300, 400 have control inputs to add or remove layers in the layers used for map visualization 302.
[0805] In some embodiments, the map user interface 300, 400 includes visualization of extracted features of the dataset 106.
[0806] In some embodiments, map user interfaces 300, 400 receive one or more reference molecules or target molecules, wherein the transformation of dataset 106 is based on one or more reference molecules or target molecules.
[0807] In some embodiments, map user interfaces 300, 400 receive one or more scores of molecules, wherein the scores include expressibility scores and fuzzy filtering scores.
[0808] In some embodiments, map visualization includes one or more clusters corresponding to molecules, wherein map user interfaces 300, 400 receive clustering control commands to update the map visualization using hyperparameters of one or more clusters.
[0809] In some embodiments, the map visualization displays one or more scores associated with the molecule, including expressibility scores or fuzzy screening scores.
[0810] In some embodiments, map user interfaces 300, 400 receive control commands for sampling from at least a portion of a map visualization.
[0811] In some embodiments, map user interfaces 300, 400 receive control commands for editing samples extracted from at least a portion of a map visualization.
[0812] In some embodiments, the map user interface 300, 400 receives drawing settings corresponding to the visualization characteristics of the map visualization.
[0813] In some embodiments, system 100 further includes user equipment 114 for displaying map user interfaces 300, 400.
[0814] In some embodiments, the map user interface 300, 400 is one-dimensional, two-dimensional, three-dimensional, four-dimensional, or higher-dimensional.
[0815] In some embodiments, the layer-by-layer embedding is a time series that varies over time, and the map user interface includes a two-dimensional or three-dimensional embedding that varies over time as a time series representing a four-dimensional embedding.
[0816] In some embodiments, the map user interface 300, 400 includes visualization of clusters around molecules of interest.
[0817] In some embodiments, the map user interface 300, 400 includes one or more clusters of molecules.
[0818] In some embodiments, the map user interface 300, 400 includes various molecular embeddings.
[0819] In some embodiments, layer-by-layer embedding includes individual molecular embedding.
[0820] In some embodiments, the map user interface 300, 400 receives scores.
[0821] In some embodiments, the map user interface 300, 400 exports files.
[0822] In some embodiments, the map user interface 300, 400 includes one or more buttons for adding or removing layers 302, one or more buttons for receiving input molecules 304, one or more buttons for adding or removing individual molecules 306, and one or more buttons for importing fractions 308.
[0823] In some embodiments, the map user interface 300, 400 includes multiple settings selected from the group consisting of: navigation settings for map visualization, drawing settings, settings for encoding scores in map visualization, clustering settings, settings for editing samples, sample settings, reporting settings, map analysis settings, and export settings.
[0824] In some embodiments, the processing subsystem causes system 100 to selectively subtract a non-immunized molecular library from an immunized molecular library to filter out non-specific molecules and reduce the search space for sampling and searching for specific candidate drug molecules for one or more targets; if the immunized library contains multiple layers or dataset 106, the subsystem causes system 100 to selectively intersect the layers after the subtraction to further reduce the search space.
[0825] In some embodiments, the processing subsystem causes system 100 to selectively subtract a molecular library immunized against one or more targets from a molecular library immunized against a target of interest to filter out molecules that do not bind to the target of interest and reduce the search space used for sampling and searching for specific candidate drug molecules against the target of interest; if the molecular library immunized against the target of interest contains multiple layers or datasets 106, then the subsystem causes system 100 to selectively intersect the layers after the subtraction to further reduce the search space.
[0826] Implementation details.
[0827] The foregoing discussion provides numerous example embodiments of the inventive subject matter. While each embodiment represents a single combination of inventive elements, the inventive subject matter is considered to include all possible combinations of the disclosed elements. Thus, if one embodiment includes elements A, B, and C, and a second embodiment includes elements B and D, then the inventive subject matter is also considered to include other remaining combinations of A, B, C, or D, even if not explicitly disclosed.
[0828] Embodiments of the devices, systems, and methods described herein can be implemented in a combination of hardware and software. These embodiments can be implemented on a programmable computer, each computer including at least one processor, a data storage system (including volatile or non-volatile memory or other data storage elements or combinations thereof), and at least one communication interface.
[0829] Program code is applied to input data to perform the functions described herein and generate output information. The output information is applied to one or more output devices. In some embodiments, the communication interface may be a network communication interface. In embodiments where elements can be combined, the communication interface may be a software communication interface, such as a software communication interface for inter-process communication. In other embodiments, there may be a combination of communication interfaces implemented as hardware, software, or a combination thereof.
[0830] Throughout the foregoing discussion, references will be made to servers, services, interfaces, portals, platforms, or other systems formed by computing devices. It should be recognized that the use of such terms is considered to refer to one or more computing devices having at least one processor configured to execute software instructions stored on a computer-readable tangible, non-transitory medium. For example, a server may include one or more computers operating as a web server, database server, or other type of computer server in a manner that fulfills the described roles, responsibilities, or functions.
[0831] The technical solutions of the embodiments can be in the form of software products. Software products can be stored on non-volatile or non-transitory storage media, such as compact disk read-only memory (CD-ROM), USB flash drives, or removable hard drives. The software products include numerous instructions that enable a computer device (personal computer, server, or network device) to perform the methods provided by the embodiments.
[0832] The embodiments described herein are implemented using physical computer hardware, including computing devices, servers, receivers, transmitters, processors, memory, displays, and networks. The embodiments described herein provide useful physical machines and specially configured computer hardware arrangements.
[0833] The embodiments and examples described herein are illustrative and not restrictive. Actual implementations of the features may be incorporated into some or all of the aspects, and the features described herein should not be considered as indications of future or existing product plans. The applicant is involved in basic and applied research, and in some cases, the described features have been developed on an exploratory basis.
[0834] Of course, the above embodiments are intended to be illustrative only and are by no means limiting. The described embodiments are readily adaptable to numerous modifications in terms of form, arrangement of components, details, and order of operation. This disclosure is intended to cover all such modifications within the scope defined by the claims.
Claims
1. A computer-implemented system for mapping proteins, protein-like molecules, or fragments thereof onto a visual interface and generating maps for display and interaction, the system comprising: A processing subsystem, comprising one or more processors and one or more memories coupled to the one or more processors, the processing subsystem being configured such that the system: Receive one or more sets of data for one or more input proteins, protein-like molecules or fragments thereof, wherein the data includes features of the one or more proteins, protein-like molecules or fragments thereof; At least one dataset is generated by processing one or more sets of data of one or more input proteins, protein-like molecules or fragments thereof; The data and one or more of the at least one dataset are transformed by feature extraction or feature selection to generate maps and additional features, thereby reducing a higher-dimensional data representation to a lower-dimensional data representation for visualization through the visual interface, wherein the lower-dimensional data representation captures valuable information from the data and one or more of the at least one dataset, and the lower-dimensional data representation includes one or more clusters of proteins, protein-like molecules or fragments thereof, including the one or more input proteins, protein-like molecules or fragments thereof, or generated proteins, protein-like molecules or fragments thereof; Generate a visual map interface, the visual map interface comprising the map as a visual representation of the lower-dimensional data representation of the dataset, the visual representation comprising a visualization of the dataset as one or more layers of proteins, protein-like molecules or fragments thereof, each layer comprising one or more of the one or more clusters of proteins, protein-like molecules or fragments thereof; as well as The visual map interface is provided with tools for interacting with the map, wherein the interaction with the map includes one or more of the following: inspection, searching, sampling, clustering, and analysis of the one or more proteins, protein-like molecules or fragments thereof, or newly generated proteins, protein-like molecules or fragments thereof; The tool is used to receive commands or detect interactions with the map at the visual map interface. Update the map based on the command or interaction; as well as Trigger an update to the visual map interface using the updated map.
2. The computer-implemented system of claim 1, wherein the protein, protein-like molecule or fragment thereof is selected from: antibodies, antigens, lectins, receptors, ligands, enzymes or fragments thereof.
3. The computer-implemented system of claim 2, wherein the protein, protein-like molecule or fragment thereof comprises an antibody or fragment thereof and / or an antigen or fragment thereof, and wherein the map is a complementary site map or an epitope map or a map comprising a protein or protein-like molecule or fragment thereof.
4. The computer-implemented system of claim 1, wherein the protein, protein-like molecule or fragment thereof includes an antibody or antibody fragment thereof, and wherein the data includes the structure of one or more antibodies or antibody fragments thereof, the amino acid sequence of one or more of the antibodies or antibody fragments thereof, the amino acid atomic or molecular coordinates, and one or more of the biophysical properties of one or more antibodies or antibody fragments thereof.
5. The computer-implemented system of claim 1, wherein the protein, protein-like molecule or fragment thereof is selected from conventional antibodies, antibody-like molecules, artificial antibodies, antibody mimics, single-domain antibodies, single-chain antibodies, humanized antibodies, chimeric antibodies or fragments thereof.
6. The computer-implemented system of claim 5, wherein the fragment comprises an antigen-binding fragment or an antigen-binding domain.
7. The computer-implemented system of claim 6, wherein the antigen-binding fragment or the antigen-binding domain is selected from one or more complementarity-determining regions and / or one or more frame regions, one or more variable domains, or complementary sites.
8. The computer-implemented system of claim 1, wherein the visual representation of the lower-dimensional data includes different colors and / or marker shapes and / or marker sizes and / or color transparency and / or color gradients to indicate the one or more layers and the one or more clusters of proteins, protein-like molecules or fragments thereof.
9. The computer-implemented system of claim 1, wherein the lower-dimensional data representation is a one-dimensional, two-dimensional, three-dimensional, or four-dimensional data representation.
10. The computer-implemented system of claim 1, wherein the feature includes a fingerprint, and wherein the processing subsystem uses a machine learning or statistical method involving one or more of dimensionality reduction, reconstructive autoencoders, variational autoencoders, adversarial autoencoders, neural networks, graph neural networks, attention networks, recurrent networks, sequence processing algorithms, image processing algorithms, computer vision algorithms, and identity transformations to extract features and generate fingerprints.
11. The computer-implemented system of claim 1, wherein the processing subsystem causes the system to cluster the data and one or more of the dataset to generate said one or more clusters of proteins, protein-like molecules or fragments thereof.
12. The computer-implemented system of claim 1, wherein the processing subsystem causes the system to encode the raw data and generate additional features from the encoded data.
13. The computer-implemented system of claim 1, wherein the visual representation overlays one or more layers of proteins, protein-like molecules or fragments thereof as overlays as part of a visualization representing the dataset, wherein the tool triggers the one or more layers to move to different positions or levels, or removes them from the map, or changes the order in which the layers are displayed, or zooms in or out of one or more layers, or moves across layers in the map.
14. The computer-implemented system of claim 1, wherein the processing subsystem enables the system to perform map analysis, wherein the map analysis includes one or more of the following: generating clusters around proteins, protein-like molecules or fragments of interest, arranging embeddings in the clusters, embedding layer by layer, embedding individual proteins, protein-like molecules or fragments of them, sampling from the map, encoding scores in the map, wherein the map contains the one or more clusters and visualizes the one or more clusters.
15. The computer-implemented system of claim 1, wherein the processing subsystem causes the system to generate or compute one or more clusters of proteins, protein-like molecules or fragments thereof, and wherein the map user interface includes visualization of the one or more clusters of proteins, protein-like molecules or fragments thereof.
16. The computer-implemented system of claim 1, wherein feature extraction This includes extracting useful information from the dataset, and feature selection includes selecting a subset of the dataset of proteins, protein-like molecules, or fragments thereof.
17. The computer-implemented system of claim 1, wherein the processing subsystem causes the system to transform the data and one or more of the dataset by one or more of sequencing and clustering, sampling, intersection of data subsets, and subtraction of data subsets to generate the map.
18. The computer-implemented system of claim 1, wherein the processing subsystem causes the system to partition or segment a digital map into a plurality of map tiles, label each of the one or more clusters with a corresponding map tile in the plurality of map tiles, and display the one or more clusters within the plurality of map tiles using the label, wherein the visualization indicates the plurality of map tiles and the one or more clusters.
19. The computer-implemented system of claim 1, wherein the processing subsystem causes the system to: (i) intersect one or more layers of proteins, protein-like molecules or fragments thereof, or (ii) subtract one or more layers of proteins, protein-like molecules or fragments thereof, or (iii) add one or more layers of proteins, protein-like molecules or fragments thereof, to update the map based on the command or interaction.
20. The computer-implemented system of claim 1, wherein the tool at the visual map interface includes a sampling tool for sampling proteins, protein-like molecules, or fragments thereof from one or more clusters, wherein the processing subsystem, in response to activation of the sampling tool, updates the map by sampling proteins, protein-like molecules, or fragments thereof, and triggers an update of the visual map interface using the updated map to visualize the sampling.
21. The computer-implemented system of claim 1, wherein the processing subsystem causes the system to subtract a library of non-immunized proteins, protein-like molecules, or fragments thereof from a library of immunized proteins, protein-like molecules, or fragments thereof to filter out non-specific proteins, protein-like molecules, or fragments thereof, and to reduce the search space used for sampling and searching for specific candidate molecules against one or more targets, wherein if multiple layers or datasets exist for the immunized library, the subsystem causes the system to intersect the layers or datasets after the subtraction to further reduce the search space.
22. The computer-implemented system of claim 1, wherein the processing subsystem causes the system to subtract from a library of proteins, protein-like molecules, or fragments thereof immunized against one or more targets to filter out non-binding portions of the target of interest and reduce the search space for sampling and searching for specific candidate molecules against the target of interest, wherein if multiple layers or datasets exist for the immune library against the target of interest, the subsystem causes the system to intersect the layers or datasets after subtraction to further reduce the search space.
23. The computer-implemented system of claim 1, wherein the processing subsystem enables the system to export or report inspection, search, sampling, clustering, and analysis of the protein, protein-like molecule, or fragment thereof through text, tables, graphs, or visualizations.
24. A computer processing method for mapping proteins, protein-like molecules, or fragments thereof onto a visual interface and generating digital maps for display and interaction, the method comprising: Receive one or more sets of data for one or more input proteins, protein-like molecules or fragments thereof, wherein the data includes features of the one or more proteins, protein-like molecules or fragments thereof; At least one dataset is generated by processing the data of one or more sets of input proteins, protein-like molecules or fragments thereof; The data and one or more of the at least one dataset are transformed by feature extraction or feature selection to generate maps and additional features, thereby reducing a higher-dimensional data representation to a lower-dimensional data representation for visualization through the visual interface, wherein the lower-dimensional data representation captures valuable information from the data and one or more of the at least one dataset, and the lower-dimensional data representation includes one or more clusters of proteins, protein-like molecules or fragments thereof, including the one or more input proteins, protein-like molecules or fragments thereof, or generated proteins, protein-like molecules or fragments thereof; Generate a visual map interface, the visual map interface including a map as a visual representation of the lower-dimensional data representation of the dataset, the visual representation including a visualization of the dataset as one or more layers of proteins, protein-like molecules or fragments thereof, each layer including one or more of the one or more clusters of proteins, protein-like molecules or fragments thereof; as well as The visual map interface is provided with tools for interacting with the map, wherein the interaction with the map includes one or more of the following: inspection, searching, sampling, clustering, and analysis of the one or more proteins, protein-like molecules or fragments thereof, or newly generated proteins, protein-like molecules or fragments thereof; The tool is used to receive commands or detect interactions with the map at the visual map interface. as well as The visual map interface and the map are updated based on the command or interaction.
25. A computer-readable medium encoded with instructions, which, when executed by a processor, cause the processor to map proteins, protein-like molecules, or fragments thereof onto a visual interface and generate a digital map for display and interaction, the instructions including instructions for performing the following: Receive one or more sets of data for one or more input proteins, protein-like molecules or fragments thereof, wherein the data includes features of the one or more proteins, protein-like molecules or fragments thereof; At least one dataset is generated by processing the data of one or more sets of input proteins, protein-like molecules or fragments thereof; The data and one or more of the at least one dataset are transformed by feature extraction or feature selection to generate maps and additional features, thereby reducing a higher-dimensional data representation to a lower-dimensional data representation for visualization through the visual interface, wherein the lower-dimensional data representation captures valuable information from the data and one or more of the at least one dataset, and the lower-dimensional data representation includes one or more clusters of proteins, protein-like molecules or fragments thereof, including the one or more input proteins, protein-like molecules or fragments thereof, or generated proteins, protein-like molecules or fragments thereof; Generate a visual map interface, the visual map interface including a map as a visual representation of the lower-dimensional data representation of the dataset, the visual representation including a visualization of the dataset as one or more layers of proteins, protein-like molecules or fragments thereof, each layer including one or more of the one or more clusters of proteins, protein-like molecules or fragments thereof; as well as The visual map interface is provided with tools for interacting with the map, wherein the interaction with the map includes one or more of the following: inspection, searching, sampling, clustering, and analysis of the one or more proteins, protein-like molecules or fragments thereof, or newly generated proteins, protein-like molecules or fragments thereof; The tool is used to receive commands or detect interactions with the map at the visual map interface. as well as The visual map interface and the map are updated based on the command or interaction.
26. A computer-implemented system for mapping molecules onto an interface and generating a map for the interface, the system comprising: A processing subsystem, comprising one or more processors and one or more memories coupled to the one or more processors, the processing subsystem being configured such that the system: Receive data of one or more sets of one or more input molecules, wherein the input molecules are one or more sets of one or more molecules, and wherein the data includes molecular characteristics; At least one dataset is generated by processing data from one or more sets of the input molecules, wherein the dataset includes an encoded sequence of information, coordinates for each molecule to define its molecular structure, and features; Transform the data and one or more of the at least one dataset by feature extraction or feature selection to generate maps and additional features, thereby reducing the higher-dimensional data representation to a lower-dimensional data representation that can be indicated, depicted or visualized by the interface, while capturing valuable information in the lower-dimensional data representation, which includes one or more clusters of the input molecule or newly generated molecule. Generate a map user interface, the map user interface comprising a map as a representation of the lower-dimensional data of the dataset, the representation representing the at least one dataset as one or more molecular layers, each layer including one or more of the one or more clusters of the molecules; and Provide the map user interface.
27. The computer-implemented system of claim 26, wherein the processing subsystem causes the system to: Provide the map user interface with tools for interacting with the map and for inspecting, searching, sampling, clustering, and analyzing the one or more input molecules or newly generated molecules; The tool is used to receive commands or detect interactions with the map at the visual map interface. Update the map based on the command or interaction; as well as Trigger an update to the map interface using the updated map.
28. The computer-implemented system of claim 26, wherein the molecule is a protein, a protein-like molecule, a fragment thereof, a small molecule drug, or a nucleic acid molecule.
29. The computer-implemented system of claim 28, wherein the input protein, protein-like molecule or fragment thereof comprises an antibody or fragment thereof and / or an antigen or fragment thereof, and wherein the map is a complementary site map or an epitope map or a map comprising a protein or protein-like molecule or fragment thereof.
30. The computer-implemented system of claim 28, wherein the protein, protein-like molecule or fragment thereof is selected from the group consisting of: antibodies, antigen-binding fragments, drug candidates, compounds, candidate conjugates and binding agents.
31. The computer-implemented system of claim 26, wherein the map user interface comprises a visual interface, and wherein the lower-dimensional data representation can be visualized in the visual interface.
32. The computer-implemented system of claim 26, wherein the input molecule is a protein or protein-like molecule comprising an antibody or an antigen-binding fragment thereof and / or an antigen.
33. The computer-implemented system of claim 32, wherein the antibody or antigen-binding fragment thereof comprises antibodies from, but not limited to, the following species: mice, cattle, rabbits, camels, llamas, humans, alpacas, and standard species.
34. The computer-implemented system of claim 26, wherein the input molecule is an antibody or its antigen-binding fragment and / or antigen, and wherein the molecular map is a complementary site map.
35. The computer-implemented system of claim 26, wherein the input molecule is an antigen, and wherein the molecular map is an epitope map.
36. The computer-implemented system of claim 26, wherein the feature includes a fingerprint, and wherein the processing subsystem uses one or more machine learning or statistical methods involving dimensionality reduction, reconstructive autoencoders, variational autoencoders, adversarial autoencoders, neural networks, graph neural networks, attention networks, and recurrent networks to extract the feature and generate the fingerprint.
37. The computer-implemented system of claim 26 further includes a data storage device for a molecular database, wherein each molecule is assigned a unique index.
38. The computer-implemented system of claim 37, wherein the system compares a molecule with a database of molecules and assigns the molecule an index to the closest molecule in the database of molecules, wherein molecules assigned the same index have similar sequences, structures, or properties.
39. The computer-implemented system of claim 26, wherein the feature extraction comprises layer-by-layer embedding, wherein, For multiple datasets or multiple parts of a dataset, extract features from each of the multiple datasets or features from each of the multiple parts of the dataset individually.
40. The computer-implemented system of claim 39, wherein the features can be drawn as superimposed visualization layers.
41. The computer-implemented system of claim 39, wherein the map user interface for visualizing the dataset has control inputs that enable viewing the layers individually or in relation to other layers.
42. The computer-implemented system of claim 39, wherein one or more layers are used for training the feature extraction, and one or more other layers are used for testing the feature extraction.
43. The computer-implemented system of claim 26, wherein the feature extraction includes arranging embeddings in clusters.
44. The computer-implemented system of claim 26, wherein the feature extraction includes embedding individual molecules.
45. The computer-implemented system of claim 26, wherein the feature extraction includes generating clusters around the molecules of interest.
46. The computer-implemented system of claim 26, wherein the feature extraction includes sampling from a molecular map.
47. The computer-implemented system of claim 26, wherein the feature extraction includes encoding fractions in a molecular map.
48. The computer-implemented system of claim 26, wherein the map user interface includes a visualization of the extracted features of the dataset.
49. The computer-implemented system of claim 26 further includes a user device for displaying the map user interface.
50. The computer-implemented system of claim 26, wherein the map user interface is one-dimensional, two-dimensional, three-dimensional, four-dimensional, or higher-dimensional.
51. The computer-implemented system of claim 39, wherein the layer-by-layer embedding is a time series that varies over time, and wherein the map user interface includes a two-dimensional or three-dimensional embedding that varies over time as a time series representing a four-dimensional embedding.
52. The computer-implemented system of claim 26, wherein the map user interface includes visualization of clusters around molecules of interest.
53. The computer-implemented system of claim 26, wherein the map user interface comprises one or more clusters of molecules.
54. The computer-implemented system of claim 26, wherein the processing subsystem performs feature extraction on each molecule of the dataset to obtain each molecule embedding.
55. The computer-implemented system of claim 39, wherein the layer-by-layer embedding comprises individual molecular embeddings.
56. A computer-implemented system as claimed in any of the preceding claims, wherein the user interface uses extracted features to characterize and / or obtain information about the input molecule, wherein the input molecule may optionally be derived from a cluster of interest.
57. The computer-implemented system of claim 56, wherein the information includes the extent, nature, and / or robustness of an immune response (against an antigen such as an immunogen or vaccine).
58. The computer-implemented system of claim 56, wherein the input molecule comprises an antibody or an antigen-binding fragment thereof, and wherein the information comprises an amino acid sequence of one or more of the antibody or antigen-binding fragment.
59. The computer-implemented system of claim 56, wherein the input molecule comprises an antigen, and wherein the information comprises an amino acid sequence of one or more of the antigens.
60. The computer-implemented system of claim 58 or 59, wherein the system allows a user to modify an amino acid sequence, and wherein the system predicts the effect of the modification on the characteristics of the molecule (binding, function, stability, expressibility, affinity, immunogenicity).
61. The computer-implemented system of claim 60, wherein the modification comprises amino acid substitution, deletion, and / or addition in one or more CDRs, variable regions, frame regions, and / or constant regions of the antibody or its antigen-binding fragment.
62. The computer-implemented system of claim 61, wherein the modification is the humanization, deimmunization, glycosylation, or deglycosylation of the antibody or its antigen-binding fragment.
63. A computer-implemented system as described in any of the preceding claims, wherein the system allows a user to import other molecules and determine their similarity to the input molecule.
64. The computer-implemented system of claim 63, wherein the other molecule is another antibody or an antigen-binding fragment thereof, and the similarity is complementary site similarity.
65. A computer-implemented system as described in any of the preceding claims, wherein the output includes an antibody or an antigen-binding fragment thereof selected from the map or a variant thereof.
66. A computer-implemented system as described in any of the preceding claims, wherein a user synthesizes an input molecule or an output molecule or a variant thereof, or causes the input molecule or the output molecule or a variant thereof to be synthesized.
67. A computer-implemented system as described in any of the preceding claims, using information provided by said system to manufacture molecules.
68. The computer-implemented system of any of the preceding claims, wherein the input molecule comprises a single-domain antibody or an antigen-binding fragment thereof, and wherein the output molecule comprises an antibody or an antigen-binding fragment thereof selected from conventional antibodies, single-domain antibodies, single-chain variable fragments, humanized antibodies, chimeric antibodies.
69. A computer-implemented system as described in any of the preceding claims, wherein the processing subsystem connects features of portions of a molecule together to have the overall features of the molecule.
70. A computer-implemented system as described in any of the preceding claims, wherein the encoded information sequence refers to an encoded amino acid sequence.
71. A computer-implemented system as described in any of the preceding claims, wherein the processing subsystem outputs the selected molecule.
72. An article of manufacture obtained from a selected molecular output by a computer-implemented system as described in any of the preceding claims.
73. A product obtained by a system implemented by a computer as described in any one of the preceding claims.
74. A computer-implemented method for mapping molecules onto an interface and generating a map for the interface, the method comprising: Receive data of one or more sets of one or more input molecules, wherein the input molecules are one or more sets of one or more molecules, and wherein the data includes molecular characteristics; At least one dataset is generated by processing data from one or more sets of one or more input molecules, wherein the dataset includes an encoded sequence of information, coordinates for each molecule to define the molecular structure, features, and fingerprints; Transform the data and one or more of the at least one dataset by feature extraction, feature selection or fingerprint generation to generate maps and additional features, thereby reducing the higher-dimensional data representation to a lower-dimensional data representation that can be indicated, depicted or visualized by the interface, while capturing valuable information in the lower-dimensional data representation, which includes one or more clusters of the input molecule or newly generated molecule. Generate a map user interface, the map user interface including the map as a representation of the lower-dimensional data of the dataset, the representation representing the at least one dataset as one or more molecular layers, each layer including one or more of the one or more clusters of the molecules; and Provide the map user interface.
75. The computer-implemented method for mapping molecules as described in claim 74, wherein the map user interface includes a visual interface, and wherein the lower-dimensional data representation can be visualized in the visual interface.
76. The computer-implemented method of claim 74, wherein the input molecule is a protein or protein-like molecule comprising an antibody or an antigen-binding fragment thereof and / or an antigen.
77. The computer-implemented method of claim 76, wherein the antibody or antigen-binding fragment thereof comprises antibodies from, but not limited to, the following species: mice, cattle, rabbits, camels, llamas, humans, alpacas, and standard species.
78. The computer-implemented method of claim 74, wherein the input molecule is an antibody or its antigen-binding fragment and / or antigen, and wherein the molecular map is a complementary site map.
79. The computer-implemented method of claim 74, wherein the input molecule is an antigen, and wherein the molecular map is an epitope map.
80. The computer-implemented method of claim 74, wherein the feature includes a fingerprint, and wherein the processing subsystem uses one or more machine learning or statistical methods involving dimensionality reduction, reconstructive autoencoders, variational autoencoders, adversarial autoencoders, neural networks, graph neural networks, attention networks, and recurrent networks to extract features and generate fingerprints.
81. The computer-implemented method of claim 74 further includes storing a database of molecules in a data storage device, wherein each molecule is assigned a unique index.
82. The computer-implemented method of claim 81 further includes comparing the molecule with a database of molecules and assigning the molecule an index to the closest molecule in the database of molecules, wherein molecules assigned the same index have similar sequences, structures, or properties.
83. The computer-implemented method of claim 81, further comprising feature extraction with layer-by-layer embedding, wherein, For multiple datasets or multiple parts of a dataset, extract features from each of the multiple datasets or features from each of the multiple parts of the dataset individually.
84. The computer-implemented method of claim 74, wherein the features can be drawn as superimposed visualization layers.
85. The computer-implemented method of claim 74, further comprising providing the map user interface that includes a visualization of the dataset, the map user interface having control inputs that enable viewing the layers individually or in relation to other layers.
86. The computer-implemented method of claim 74, wherein one or more layers are used for training the feature extraction, and one or more other layers are used for testing the feature extraction.
87. The computer-implemented method of claim 74, wherein the feature extraction includes arranging embeddings in clusters.
88. The computer-implemented method of claim 74, wherein the feature extraction includes embedding individual molecules.
89. The computer-implemented method of claim 74, wherein the feature extraction includes generating clusters around the molecules of interest.
90. The computer-implemented method of claim 74, wherein the feature extraction includes sampling from a molecular map.
91. The computer-implemented method of claim 74, wherein the feature extraction includes encoding fractions in a molecular map.
92. The computer-implemented method of claim 74, wherein the map user interface includes a visualization of the extracted features of the dataset.
93. The computer-implemented method of claim 74 further includes using a user device to display the map user interface.
94. The computer-implemented method of claim 74, wherein the map user interface is one-dimensional, two-dimensional, three-dimensional, four-dimensional, or higher-dimensional.
95. The computer-implemented method of claim 74, wherein the layer-by-layer embedding is a time series that varies over time, and wherein the map user interface includes a two-dimensional or three-dimensional embedding that varies over time as a time series representing a four-dimensional embedding.
96. The computer-implemented method of claim 74, further comprising using the map user interface to provide visualization of clusters around molecules of interest.
97. The computer-implemented method of claim 74, wherein the map user interface comprises one or more clusters of molecules.
98. The computer-implemented method of claim 74 further comprises performing feature extraction on each molecule of the dataset to obtain each molecule embedding.
99. The computer-implemented method of claim 74, wherein the layer-by-layer embedding comprises individual molecular embeddings.
100. The computer-implemented method as described in any of the preceding claims, further comprising using the extracted features to characterize and / or obtain information about the input molecule, wherein the input molecule may optionally be derived from a cluster of interest.
101. The computer-implemented method of claim 100, wherein the information includes the extent, nature, and / or robustness of an immune response (against an antigen such as an immunogen or vaccine).
102. The computer-implemented method of claim 100, wherein the input molecule comprises an antibody or an antigen-binding fragment thereof, and wherein the information comprises an amino acid sequence of one or more of the antibody or antigen-binding fragment.
103. The computer-implemented method of claim 100, wherein the input molecule comprises an antigen, and wherein the information comprises an amino acid sequence of one or more of the antigens.
104. The computer-implemented method of claim 102 or 103 further includes modifying the amino acid sequence, wherein the system predicts the effect of the modification on the characteristics of the molecule (binding, function, stability, expressibility, affinity, immunogenicity, etc.).
105. The computer-implemented method of claim 104, wherein the modification comprises amino acid substitution, deletion, and / or addition in one or more CDRs, variable regions, frame regions, and / or constant regions of the antibody or its antigen-binding fragment.
106. The computer-implemented method of claim 105, wherein the modification is humanization, deimmunization, glycosylation, or deglycosylation of the antibody or its antigen-binding fragment.
107. A computer-implemented method as described in any of the preceding claims, wherein the system allows a user to import other molecules and determine their similarity to the input molecule.
108. The computer-implemented method of claim 107, wherein the other molecule is another antibody or an antigen-binding fragment thereof, and the similarity is complementary site similarity.
109. A computer-implemented method as described in any of the preceding claims, wherein the output comprises an antibody or an antigen-binding fragment thereof selected from the map or a variant thereof.
110. A computer-implemented method as described in any of the preceding claims, wherein a user synthesizes an input molecule or an output molecule or a variant thereof, or causes the input molecule or the output molecule or a variant thereof to be synthesized.
111. The computer-implemented method as described in any of the preceding claims, further comprising using information provided by the system to manufacture molecules.
112. The computer-implemented method as described in any of the preceding claims, wherein the input molecule comprises a single-domain antibody or an antigen-binding fragment thereof, and wherein the output molecule comprises an antibody or an antigen-binding fragment thereof selected from conventional antibodies, single-domain antibodies, single-chain variable fragments, humanized antibodies, chimeric antibodies.
113. The computer-implemented method as described in any of the preceding claims further includes linking features of portions of the molecule together to have the overall features of the molecule.
114. The computer-implemented method as described in any of the preceding claims, wherein the encoded information sequence refers to an encoded amino acid sequence.
115. The computer-implemented method as described in any of the preceding claims further includes outputting the selected molecule.
116. An article of manufacture obtained by a selected molecular output of a computer-implemented method as described in any of the preceding claims.
117. A product obtained by a computer-implemented method as described in any one of the preceding claims.
118. A computer-implemented method as described in any of the preceding claims, wherein the method includes the step of generating molecules identified or selected from the map user interface.
119. A product obtained by a computer-implemented method according to any one of the preceding claims.
120. The product of claim 119, wherein the product is an antibody or an antigen-binding fragment thereof.
121. A non-transitory computer-readable medium having stored thereon machine-interpretable instructions, which, when executed by a processing subsystem, cause the processing subsystem to perform a method for mapping molecules onto a visual interface and generating a map for the interface, the method comprising: Receive data of one or more sets of one or more input molecules, wherein the input molecules are one or more sets of one or more molecules, and wherein the data includes molecular characteristics; At least one dataset is generated by processing data about one or more sets of one or more input molecules, wherein the dataset includes an encoded sequence of information, coordinates for each molecule to define the molecular structure, features, and fingerprints; Transform the data and one or more of the at least one dataset by feature extraction, feature selection or fingerprint generation to generate maps and additional features, thereby reducing the higher-dimensional data representation to a lower-dimensional data representation that can be indicated by the interface, while capturing valuable information in the lower-dimensional data representation, which includes one or more clusters of the input molecule or newly generated molecule. Generate a map user interface, the map user interface comprising a map as a representation of the lower-dimensional data representation of the dataset, the representation representing the at least one dataset as one or more molecular layers, each layer including one or more clusters of molecules; and Provide the map user interface.
122. A computer-implemented system for a molecularly related interface, the system comprising: A processing subsystem, comprising one or more processors and one or more memories coupled to the one or more processors, the processing subsystem being configured such that the system: Receive input molecules, wherein the input molecules are one or more groups of one or more molecules, and each molecule is defined as an information sequence; Encode the information sequence; A dataset is generated by processing the input molecules, wherein the dataset includes an encoded information sequence, three-dimensional coordinates of the information sequence for each molecule to define the molecular structure, features, and fingerprint; The dataset is transformed by feature extraction and fingerprint generation, thereby reducing the high number of dimensions to a lower-dimensional data representation that can be indicated by the interface, while capturing valuable data in the lower-dimensional data representation. One or more metrics are generated from the transformed dataset, wherein the one or more metrics include a lower-dimensional data representation of the dataset and summarize the characteristics of the input molecule; as well as Provide one or more of the aforementioned metrics to the interface.
123. A computer-implemented system for a visual interface for mapping molecules, the system comprising: A processing subsystem, comprising one or more processors and one or more memories coupled to the processors, provides a map user interface, wherein the map user interface: Receive input molecules, wherein the input molecules are one or more groups of one or more molecules, wherein each molecule is defined as an information sequence; and A map interface is provided, which includes a molecular map as a low-dimensional data representation of the dataset of the input molecules, wherein the dataset includes an information sequence for each molecule, three-dimensional coordinates of an information sequence for each molecule to define the molecular structure, and features; The molecular map described therein includes transforming the dataset through feature extraction and fingerprint generation, thereby reducing a high number of dimensions to a lower-dimensional data representation that can be indicated by the interface, while capturing valuable data in the lower-dimensional data representation.
124. The computer-implemented system of claim 123, wherein the input molecule is a protein or protein-like molecule comprising antibodies and antigens.
125. The computer-implemented system of claim 124, wherein the antibody comprises a species antibody, the species including but not limited to mice, cattle, rabbits, camels, llamas, humans, alpacas, and standard species.
126. The computer-implemented system of claim 123, wherein the input molecules are antibodies and antigens, and wherein the molecular map is a complementary bit map.
127. The computer-implemented system of claim 123, wherein the input molecule is an antigen, and wherein the molecular map is an epitope map.
128. The computer-implemented system of claim 123, wherein the map user interface includes layer-by-layer embedding that provides layers for map visualization.
129. The computer-implemented system of claim 128, wherein the map user interface renders features as overlaid map visualization layers.
130. The computer-implemented system of claim 128, wherein the map user interface has control inputs that enable viewing the layers individually or in relation to other layers.
131. The computer-implemented system of claim 128, wherein the map user interface has control inputs to add or remove layers in the layers used for map visualization.
132. The computer-implemented system of claim 123, wherein the map user interface includes visualization of extracted features of the dataset.
133. The computer-implemented system of claim 123, wherein the map user interface receives one or more reference molecules or target molecules, wherein the transformation of the dataset is based on the one or more reference molecules or target molecules.
134. The computer-implemented system of claim 123, wherein the map user interface receives one or more scores of a molecule, wherein the scores include expressibility scores and fuzzy filtering scores.
135. The computer-implemented system of claim 123, wherein the map visualization includes one or more clusters corresponding to molecules, wherein the map user interface receives clustering control commands to update the map visualization using hyperparameters of the one or more clusters.
136. The computer-implemented system of claim 123, wherein the map visualization displays one or more scores related to the molecule, the scores including expressibility scores or fuzzy screening scores.
137. The computer-implemented system of claim 123, wherein the map user interface receives control commands for sampling from at least a portion of the map visualization.
138. The computer-implemented system of claim 123, wherein the map user interface receives control commands for editing samples extracted from at least a portion of the map visualization.
139. The computer-implemented system of claim 123, wherein the map user interface receives drawing settings corresponding to the visualization characteristics of the map visualization.
140. The computer-implemented system of claim 123 further includes a user device for displaying the map user interface.
141. The computer-implemented system of claim 123, wherein the map user interface is one-dimensional, two-dimensional, three-dimensional, four-dimensional, or higher-dimensional.
142. The computer-implemented system of claim 128, wherein the layer-by-layer embedding is a time series that varies over time, and wherein the map user interface includes a two-dimensional or three-dimensional embedding that varies over time as a time series representing a four-dimensional embedding.
143. The computer-implemented system of claim 123, wherein the map user interface includes visualization of clusters around molecules of interest.
144. The computer-implemented system of claim 123, wherein the map user interface comprises one or more clusters of molecules.
145. The computer-implemented system of claim 123, wherein the map user interface includes individual molecular embeddings.
146. The computer-implemented system of claim 128, wherein the layer-by-layer embedding comprises individual molecular embeddings.
147. The computer-implemented system of claim 123, wherein the map user interface receives scores.
148. The computer-implemented system of claim 123, wherein the map user interface exports a file.
149. The computer-implemented system of claim 123, wherein the map user interface includes one or more buttons for adding or removing layers, one or more buttons for receiving input molecules, one or more buttons for adding or removing individual molecules, and one or more buttons for importing fractions.
150. The computer-implemented system of claim 123, wherein the map user interface includes a plurality of settings selected from the group consisting of: navigation settings for the map visualization, drawing settings, settings for encoding scores in the map visualization, clustering settings, settings for editing samples, sample settings, reporting settings, map analysis settings, and export settings.
151. The computer-implemented system of claim 26, wherein the processing subsystem causes the system to selectively subtract a non-immunized molecular library from an immunized molecular library to filter out non-specific molecules and reduce the search space for sampling and searching for specific candidate drug molecules for one or more targets; if the immunized library contains multiple layers or datasets, then the subsystem causes the system to selectively intersect the layers after subtraction to further reduce the search space.
152. The computer-implemented system of claim 26, wherein the processing subsystem causes the system to selectively subtract a molecular library immunized against one or more targets from a molecular library immunized against a target of interest to filter out molecules that do not bind to the target of interest and reduce the search space used for sampling and searching for specific candidate drug molecules against the target of interest; if the molecular library immunized against the target of interest contains multiple layers or datasets, then the subsystem causes the system to selectively intersect the layers after the subtraction to further reduce the search space.
Citation Information
Patent Citations
Transgenic animals expressing heavy chain antibodies
WO2022011457A1