Methods and systems for identifying peptide patterns for cancer vaccine
A heuristic search algorithm and machine learning models are used to iteratively identify patient-specific peptide patterns, addressing the inefficiencies in existing methods and improving the efficacy of personalized cancer vaccines by targeting precise antigens.
Patent Information
- Application Number
- PCT/US2025/037115
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-15
- Filing Date
- 2025-07-10
- Publication Date
- 2026-01-22
Smart Images

Figure US2025037115_22012026_PF_FP_ABST
Abstract
Description
METHODS AND SYSTEMS FOR IDENTIFYING PEPTIDE PATTERNS FORCANCER VACCINERELATED APPLICATION
[0001] This application claims priority to U.S. Provisional Patent Application Serial No. 63 / 671,652, filed July 15, 2024, which is incorporated herein by reference in its entirety for all purposes.TECHNICAL FIELD
[0002] The subject matter described herein relates to systems and methods for using Machine Learning (ML) technique to identify peptide patterns for developing cancer vaccine, more specifically, it involves the application of algorithms to determine relevant peptide sequences that could inform the design of personalized medical treatments.BACKGROUND
[0003] Peptides are short chains of amino acids that are linked by peptide bonds. They are a common biological entity and are involved in a wide range of biological processes. Peptides can be found in every cell and tissue and have a variety of functions, including acting as enzymes, hormones, or antibodies. In the field of bioinformatics, peptides are often represented as strings of letters, where each letter corresponds to a different amino acid. For example, the peptide "LEDVSKPPA" represents a chain of nine amino acids: Leucine (L), Glutamic Acid (E), Aspartic Acid (D), Valine (V). Serine (S), Lysine (K), Proline (P), Proline (P), and Alanine (A).
[0004] The identification of patterns in peptide sequences can provide valuable insights into the biological processes in which these peptides are involved. For instance, patterns in peptide sequences can indicate the presence of specific protein families, the occurrence of post- translational modifications, or the binding affinity to other molecules.
[0005] In the context of cancer research, the identification of peptide patterns can be particularly valuable. Cancer cells often present specific peptides on their surface, which can serve as markers for the identification of the cancer cells. By identifying these peptides and the patterns they form, it is possible to develop personalized therapies that specifically target these markers.SUMMARY
[0006] Methods, systems, and articles of manufacture, including computer program products, are provided for identifying peptide patterns. In one aspect, there is provided a method. The method includes receiving a data set representing peptides presented on one subject's cells; identify ing a plurality of root nodes in the data set, wherein each of the plurality of root nodes is a data string with a predefined level of specificity; iteratively identifying a plurality of patterns from each of the plurality of root nodes, wherein each of the plurality of patterns has a higher level of specificity than the predefined level of specificity of the root nodes it branches from; and outputting a visualization of the identified patterns upon a determination that a stopping criterion is reached.
[0007] In some variations, a level of specificity indicates a number of letters with known locations and identifications in the data string.
[0008] In some variations, the stopping criterion is based on a predetermined minimum number of matching peptides for a pattern.
[0009] In some variations, the stopping criterion for each pattern is based on a predetermined minimum number of matching peptides, wherein the predetermined minimum number depends on a predicted HLA type associated with the pattern being identified.
[0010] In some variations, the visualization of the identified patterns comprises a directed acyclic graph (DAG) where nodes represent patterns and edges represent parent-child relationships between patterns.
[0011] In some variations, the method further comprises employing a heuristic search algorithm that utilizes a probability distribution based on occurrences of each possible amino acid-position pair within the peptides to guide the identification of the plurality of patterns.
[0012] In some variations, the stopping criterion comprises stopping at one matching peptide for a pattern.
[0013] In some variations, the method further comprises calculating a cluster score for each identified pattern.
[0014] In some variations, the cluster score factors in a count score, and wherein the count score represents the count of peptides from the subject that belongs to the patern.
[0015] In some variations, the cluster score factors in the level of specificity associated with the identified patern, and wherein a higher level of specificity' contributes higher value to the cluster score.
[0016] In another aspect, there is provided a computer program product including a non- transitory computer readable medium storing instructions. The operations include receiving a data set representing peptides presented on one subject’s cells; identifying a plurality of root nodes in the data set, wherein each of the plurality of root nodes is a data string with a predefined level of specificity; iteratively identifying a plurality of paterns from each of the plurality of root nodes, wherein each of the plurality of paterns has a higher level of specificity than the predefined level of specificity of the root nodes it branches from; and outputing a visualization of the identified paterns upon a determination that a stopping criterion is reached.
[0017] In some variations, a level of specificity indicates a number of letters with known locations and identifications in the data string.
[0018] In some variations, the stopping criterion is based on a predetermined minimum number of matching peptides for a pattern.
[0019] In some variations, the stopping criterion for each pattern is based on a predetermined minimum number of matching peptides, wherein the predetermined minimum number depends on a predicted HLA type associated with the pattern being identified.
[0020] In some variations, the visualization of the identified patterns comprises a directed acyclic graph (DAG) where nodes represent patterns and edges represent parent-child relationships between patterns.
[0021] In some variations, the operations further comprise employing a heuristic search algorithm that utilizes a probability distribution based on occurrences of each possible amino acid-position pair within the peptides to guide the identification of the plurality of patterns.
[0022] In some variations, the stopping criterion comprises stopping at one matching peptide for a pattern.
[0023] In some variations, the operations further comprise calculating a cluster score for each identified pattern.
[0024] In some variations, the cluster score factors in a count score, and wherein the count score represents the count of peptides from the subject that belongs to the pattern.
[0025] In some variations, the cluster score factors in the level of specificity associated with the identified pattern, and wherein a higher level of specificity contributes higher value to the cluster score.
[0026] In another aspect, there is provided a system. The system may include at least one processor and at least one memory. The at least one memory may store instructions that result in operations when executed by the at least one processor. The operations may include: receiving a data set representing peptides presented on one subject’s cells; identifying a plurality of root nodes in the data set, wherein each of the plurality of root nodes is a data string with a predefined level of specificity; iteratively identify ing a plurality of patterns from each of the plurality of root nodes, wherein each of the plurality of patterns has a higher level of specificity than the predefined level of specificity of the root nodes it branches from; and outputting a visualization of the identified patterns upon a determination that a stopping criterion is reached.
[0027] In some variations, a level of specificity indicates a number of letters with known locations and identifications in the data string.
[0028] In some variations, the stopping criterion is based on a predetermined minimum number of matching peptides for a pattern.
[0029] In some variations, the stopping criterion for each pattern is based on a predetermined minimum number of matching peptides, wherein the predetermined minimum number depends on a predicted HLA type associated with the pattern being identified.
[0030] In some variations, the visualization of the identified patterns comprises a directed acyclic graph (DAG) where nodes represent patterns and edges represent parent-child relationships between patterns.
[0031] In some variations, the operations further comprise employing a heuristic search algorithm that utilizes a probability distribution based on occurrences of each possible amino acid-position pair within the peptides to guide the identification of the plurality of patterns.
[0032] In some variations, the stopping criterion comprises stopping at one matching peptide for a pattern.
[0033] In some variations, the operations further comprise calculating a cluster score for each identified pattern.
[0034] In some variations, the cluster score factors in a count score, and wherein the count score represents the count of peptides from the subject that belongs to the pattern.
[0035] In some variations, the cluster score factors in the level of specificity associated with the identified pattern, and wherein a higher level of specificity contributes higher value to the cluster score.
[0036] Implementations of the current subject matter can include, but are not limited to, methods consistent with the descriptions provided herein as well as articles that include a tangibly embodied machine-readable medium operable to cause one or more machines (e.g., computers, etc.) to result in operations implementing one or more of the described features. Similarly, computer systems are also described that may include one or more processors and one or more memories coupled to the one or more processors. A memory, which can include a computer- readable storage medium, may include, encode, store, or the like one or more programs that cause one or more processors to perform one or more of the operations described herein. Computer implemented methods consistent with one or more implementations of the current subject matter can be implemented by one or more data processors residing in a single computing system or multiple computing systems. Such multiple computing systems can be connected and can exchange data and / or commands or other instructions or the like via one or more connections, including but not limited to a connection over a network (e.g. the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, or the like), via a direct connection between one or more of the multiple computing systems, etc.
[0037] The details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features and advantages of thesubj ect matter described herein will be apparent from the description and drawings, and from the claims. The claims that follow this disclosure are intended to define the scope of the protected subj ect matter.DESCRIPTION OF DRAWINGS
[0038] The accompanying drawings, which are incorporated in and constitute a part of this specification, show certain aspects of the subject matter disclosed herein and, together with the description, help explain some of the principles associated with the disclosed implementations. In the drawings,
[0039] FIG. 1 is a diagram illustrating an example of a search graph for identifying peptide patterns in a subject’s cells, in accordance with one or more embodiments of the current subject matter.
[0040] FIG. 2 is a diagram illustrating an example of a search graph for identifying peptide patterns in a subject’s cells, in accordance with one or more embodiments of the cunent subject matter.
[0041] FIG. 3 is a diagram illustrating a flow chart of a process for identifying peptide patterns for a given subject, in accordance with one or more embodiments of the current subject matter.
[0042] FIG. 4 is a diagram illustrating a flow chart of a process for training an Al model for identifying peptide patterns of a given subject, in accordance with one or more embodiments of the current subject matter.
[0043] FIG. 5 depicts a block diagram illustrating a computing system consistent with implementations of the current subject matter.
[0044] When practical, like labels are used to refer to same or similar items in the drawings.DETAILED DESCRIPTION
[0045] The details of one or more variations of the subject matter described herein are set forth in the accompanying drawings.
[0046] As discussed herein elsewhere, the identification of patterns in peptide sequences can provide valuable insights into the biological processes. For example, in the context of cancer research, the identification of peptide patterns can be particularly valuable. By employing Artificial Intelligence (Al) models to systematically analyze these sequences, researchers can uncover specific peptide sequences that may be indicative of cancerous activity. Such Al models pattern identification may be utilized in developing personalized medical interventions, including vaccines that target these peptides, thereby offering a tailored therapeutic approach to cancer treatment.
[0047] FIG. 1 is a diagram illustrating an example of a search graph 100 for identifying peptide patterns in a subject’s cells, in accordance with one or more embodiments of the current subject matter. As shown in FIG. 1, peptides may be represented by the amino acids’ letters of their sequences. Each amino acid in a peptide is denoted by a specific letter, corresponding to its place in the standard amino acid alphabet. This representation allows for using data strings to represent peptide sequences, where each character in the string corresponds to an amino acid in the peptide, as defined by the standard single-letter amino acid code. This representation also allows for the systematic analysis of peptide sequences using Al models, which can identify patterns and correlations within the data that may not be readily apparent through manual examination. As shown in FIG. 1 , a level of specificity 110 may indicate a number of letters with known locations and identifications in the data string. For example, element 120 represents a level two (2) specificity that may indicate that there are two amino acids at specific positions within the peptide sequence that are known and fixed. This could be represented as a pattern where two positions are occupied by specific amino acids, and the remaining positions are variable or undetermined.For example, node 121, i.e., “ D * > Y indicates a level 2 of specificity, because it is known that the fourth amino acid is a “D”, and the second to last amino acid is a “Y”, with representing a w ildcard. The wildcard may indicate a variable number of unspecified amino acids that can occupy the positions between the known amino acids. In the given pattern, node 121, i.e., “ D * > Y the asterisk (“*”) sen es as a placeholder for any sequence of amino acids of any length, including a sequence of zero length. This allow s for flexibility in the peptide sequence, accommodating various lengths and compositions of peptides that still conform to the known specificity at the defined positions. The underscoresrepresent single positions that can be occupied by any amino acid, further contributing to the variability of the pattern.
[0048] As shown in FIG. 1, element 140 represents a level four (4) specificity that may indicate that there are four amino acids at specific positions within the peptide sequence that are known and fixed. For example, element 141, i.e., “ S D K *_Y may represent a peptide pattern where the third position is occupied by the amino acid "S", the fourth position by "D", the seventh position by "K", and the second to last position by "Y". The underscoresdenote positions that can be filled by any amino acid, and the asterisk ("*") serves as a wildcard for a stretch of any number of amino acids, which can also be of zero length. This pattern, therefore, specifies a sequence with a level of specificity of four (4), as it contains four know n and fixed amino acids at particular locations within the peptide sequence, allowing for the identification of peptides that match this specific arrangement.
[0049] In some embodiments, an input data set may include a number of data strings that represent peptides from a subject’s cells. It should be noted that the peptides are collected from one patient’s cells to ensure the personalization of the treatment or analysis. This personalization allow s for the identification of patterns that are specific to the individual's cellular makeup, which can be particularly beneficial for tailoring medical treatments such as personalized cancervaccines or targeted therapies. By analyzing peptides that are uniquely presented on a single patient's cells, the method can uncover patient-specific antigens that may serve as precise targets for immunotherapy, thereby increasing the efficacy and reducing potential side effects associated with broader-spectrum treatments. As described herein elsewhere, data strings may be utilized to represent peptide sequences, where each character in the string corresponds to an amino acid in the peptide, as defined by the standard single-letter amino acid code. In some embodiments, a data set representing peptides presented on one subject’s cells may be received, and wherein the subj ect may suffer from one or more t pes of cancer.
[0050] In some embodiments, once the data set representing the peptides in the subject’s cells is received, the approach disclosed herein may further identify a plurality of root nodes in the data set. In some embodiments, root nodes are the data strings with a predefined level of specificity. For example, the approach may employ a computational algorithm to analyze the peptide sequences and determine initial patterns with a predefined level of specificity, which serve as the root nodes. In some embodiments, these root nodes may act as starting points for the iterative identification of more specific patterns, or child nodes, which are derived from the root nodes. The identification of root nodes may involve a foundational step in constructing a hierarchical framework of peptide patterns, which can be used to systematically explore the peptide space for potential biomarkers or targets for therapeutic intervention. In some embodiments, the method may include a brute-force identification of root nodes within the data set representing peptides presented on the subject's cells. This process involves exhaustively analyzing the data set to determine initial patterns that meet a predefined level of specificity', which are designated as root nodes. The brute-force approach may ensure that the foundational patterns are identified without overlooking any potential root nodes. Alternatively or additionally, the method may employ heuristic algorithms, machine learning techniques, or other advanced computational methods to identify root nodes within the data set representing peptidespresented on the subject's cells. These methods may enhance the efficiency of the root node identification process, reducing the computational resources and time compared to brute-force methods.
[0051] As shown in FIG. 1 , in some embodiments, the approach described herein may further identify patterns. In some embodiments, patterns may include data strings that have a higher level of specificity than the predefined level of specificity of the root nodes it branches from. For example, as shown in FIG. 1, element 130 may include nodes with a level of specificity of three (3). Node 131 and node 132 may be considered as patterns, and they have higher level of specificity than the root node they branch from, which is the node 121. As shown in FIG. 1. node 131 is expressed as “ S D * > Y which contains one more known amino acid position than the node 121. i.e., “ D * > Y The additional known amino acid position in node 131 is the 'S' at the third position, which is specified, as opposed to the corresponding position in node 121, which is not specified and thus can be any amino acid. In some embodiments, node 131 and node 132 may not themselves constitute identified patterns; rather, they may serve as intermediate nodes that act as parent nodes to subsequent, more specific patterns. As parent nodes, node 131 and node 132 facilitate the iterative identification of patterns with increasing levels of specificity within the data set. The specificities regarding what constitutes a pattern are described in further detail elsewhere herein.
[0052] In some embodiments, node 141 may be a pattern that branches from node 121 via intermediate nodes such as node 131. Node 141 represents a more specific pattern within the hierarchical search structure, having evolved from the more general pattern of node 121. The transition from node 121 to node 141 involves the iterative process of increasing specificity by defining additional amino acid positions in the peptide sequence. This process may include the identification and establishment of intermediate nodes like node 131 and node 132, which potentially serve as parent nodes to node 141. Each intermediate node contributes to therefinement of the pattern, adding specificity until node 141 is reached with its level four specificity7. In some embodiments, the process of iteratively identifying a plurality of patterns from each of the plurality7of root nodes involves a recursive or iterative computational method that builds upon the initial patterns (e.g., root nodes or parent nodes) to discover more detailed and specific patterns (e.g., child nodes). Each root node or parent node serves as a starting point for the generation of new patterns that share a common structure with the root node but include additional specified amino acids at particular positions, thereby increasing the specificity' of the pattern. The predefined level of specificity' for a root node may be a baseline level that is set to ensure that the initial patterns are broad enough to capture a wide range of potential matches within the peptide data set. This level of specificity7is characterized by a minimum number of known amino acids at specific positions within the peptide sequence. As the process iterates, each subsequent pattern derived from a root node will have a higher level of specificity7than the root node it branches from. This is achieved by introducing additional known amino acids at previously unspecified positions or by' further specifying the positions of amino acids that were previously represented by wildcards or placeholders in the root node pattern.
[0053] In some embodiments, the pattern identification process may be performed via a heuristic search algorithm. This algorithm may employ a variety of heuristic techniques to efficiently navigate the search space 101 of the search graph 100 of potential peptide patterns. The heuristic search algorithm may prioritize the exploration of pattern branches that are statistically more likely to yield informative and biologically relevant patterns, based on prior knowledge or probabilistic models. In some embodiments, for example, the heuristic search algorithm may utilize a probability distribution that is derived from the frequency of occurrence of each amino acid at each position within the peptide data set. This distribution can guide the algorithm in selecting which amino acid to introduce at a given position in the pattern, thereby increasing the likelihood of identifying a pattern that matches a substantial number of peptidespresented on the subject's cells. In some embodiments, the heuristic search algorithm may also incorporate rules or constraints that are based on biological relevance or empirical data, such as favoring the introduction of amino acids that are known to be commonly presented in the context of the subject's HLA type or avoiding patterns that are unlikely to be processed and presented by the immune system. In some embodiments, the heuristic search algorithm may implement a scoring system for evaluating the quality of identified patterns. This scoring system may take into account factors such as the number of matching peptides, the level of specificity7of the pattern, and the potential immunogenicity7of the pattern. Patterns with higher scores may be considered more promising and may be selected for further expansion and refinement.
[0054] In some embodiments, the pattern identification process may be performed by an Al model or a Machine Learning (ML) model. In some embodiments, the Al model may7be trained by providing it with a dataset comprising a large number of known peptide sequences and their associated properties, such as binding affinities, immunogenicity, and frequency of occurrence within a population or a specific patient's cells. The training process may include data preprocessing, where peptide sequences are encoded into a format suitable for the Al or ML model, which could involve one-hot encoding, sequence embedding, or other numerical representation techniques. Relevant features may be extracted or selected from the peptide sequences, which could include the position-specific frequency of amino acids, known epitopes, motifs, or other biologically relevant signals. The Al or ML model may be trained using the preprocessed data and selected features by adjusting the model's parameters to minimize a loss function, which measures the difference between the model's predictions and the actual data. The model may be validated using a separate set of data not included in the training set to ensure that the model generalizes well to new7, unseen data. Hyperparameter tuning may be performed to find the model configuration that yields the best performance on the validation set. The trained model may be evaluated using various metrics, such as accuracy, precision, recall, or area underthe receiver operating characteristic curve (AUC-ROC), to assess its ability to correctly identify patterns. Once trained, the Al or ML model may be used to predict new patterns in peptide sequences from the subject's cells. The model may generate a score or probability for each potential pattern, indicating the likelihood that the pattern is a meaningful biomarker or target for therapeutic intervention. In some embodiments, the Al or ML model may be a neural network, such as a convolutional neural network (CNN) for capturing spatial patterns within sequences, or a recurrent neural network (RNN) for capturing sequential dependencies. Alternatively, the model may be a support vector machine (SVM), a random forest, or another suitable ML algorithm. The choice of model may depend on the nature of the data, the complexify of the patterns being identified, and the computational resources available.
[0055] In some embodiments, upon meeting a stopping criterion related to the pattern identification process, the system may generate a visualization to represent the identified patterns. This visualization can illustrate the hierarchical structure of the patterns, their levels of specificity, and the relationships between them, providing a clear and interpretable overview of the results obtained from the pattern identification. The stopping criterion may be based on various factors, such as reaching a predetermined level of pattern specificity, identifying a maximum number of patterns, or achieving a minimum threshold of matching peptides for a given pattern. The visualization could take the form of a directed acyclic graph (DAG), a tree structure, or any other graphical representation that effectively conveys the hierarchical relationships between the identified patterns and their respective levels of specificity. The stopping criterion may be determined by the desired outcome of the pattern identification process. For instance, if the goal is to identify patterns with a high likelihood of representing biologically relevant targets, the stopping criterion may be set to stop the search when patterns no longer match a sufficient number of peptides, indicating a high level of specificity but potentially lower biological relevance. Conversely, if the goal is to explore the peptide spacemore broadly, the stopping criterion may allow for the inclusion of patterns with fewer peptide matches, thus capturing a wider array of potential targets. Once the stopping criterion is met, the system may automatically generate the visualization, which can then be used by researchers or clinicians to analyze the identified patterns, explore their potential as biomarkers or therapeutic targets, and make informed decisions about further research or treatment strategies. The visualization may also include additional information, such as cluster scores, frequency of occurrence, or other relevant data, to provide a comprehensive overview of the pattern identification results.
[0056] In some embodiments, the stopping criterion for the pattern identification process may comprise a set of predefined conditions that determine when the iterative search for patterns is to be concluded. These conditions may include reaching a specific level of pattern specificity, where the search halts if patterns achieve a level of detail that meets or exceeds the predefined specificity threshold. In some embodiments, another condition may be the identification of a maximum number of patterns, which serves to limit the exploration to a predetermined quantity of potential targets. Additionally, the stopping criterion may involve a minimum number of matching peptides for a pattern, ceasing the search when patterns fail to match a sufficient number of peptides, which could indicate a lower biological or clinical relevance. Computational constraints, such as time limits or resource usage, may also be incorporated into the stopping criterion to ensure the search remains within practical bounds. In some cases, the stopping criterion may be adaptive, allowing for adjustments based on real-time analysis of the patterns' quality, such as their immunogenic potential or their likelihood to serve as effective therapeutic targets.
[0057] In some embodiments, the stopping criterion for each pattern during the pattern identification process is based on a predetermined minimum number of matching peptides, wherein the predetermined minimum number is tailored to the predicted HLA type associatedwith the pattern being identified. The goal of this approach is to include as many different HLA ty pes that may be present in the patient. Therefore, the stopping criterion is adjusted to reflect the specific HLA ty pe's propensity to present peptides. For example, if a particular HLA ty pe is known to present a wide array of peptides, a higher minimum number of matching peptides may be set as the stopping criterion for patterns associated with that HLA ty pe. Conversely, for HLA ty pes that present fewer peptides, a lower minimum number of matching peptides may be sufficient to indicate a pattern of interest. This adaptive approach allows for a more nuanced and targeted search, potentially leading to the identification of patterns that are more likely to be therapeutically relevant for the subject based on their individual HLA type profile.
[0058] In some embodiments, the method includes employing a heuristic search algorithm that utilizes a probability distribution based on the occurrences of each possible amino acid-position pair within the peptides to guide the identification of the plurality of patterns. This approach may utilize statistical analysis to inform the search process, wherein the frequency of each amino acid occurring at each position within the dataset of peptides is calculated to create a probability distribution. The heuristic search algorithm may then use this distribution to prioritize the exploration of patterns that are statistically more likely to occur, effectively guiding the search towards areas of the search space with a higher density of matching peptides. In some embodiments, the Al models may employ machine learning techniques to refine the probability distribution iteratively, allowing the heuristic search algorithm to adapt to emerging patterns as the search progresses. That is to say, in some embodiments, the heuristic search algorithm is not static but can be dynamically refined using machine learning techniques. This probabilistic approach aims to optimize the search process by focusing computational resources on the identification of patterns that have a higher likelihood of being biologically relevant, based on the observed distribution of amino acid-position pairs. By incorporating the probability distribution into the heuristic search, the method can more efficiently navigate the vastcombinatorial space of potential peptide patterns, thereby enhancing the overall efficacy of the pattern identification process.
[0059] In some embodiments, the approach described herein may include calculating a cluster score for each identified pattern. In some embodiments, the cluster score may indicate a quantitative measure of the pattern's relevance or prevalence within the dataset. This cluster score may be derived from various factors, such as the number of peptides that match the pattern, the level of specificity of the pattern, and the frequency of occurrence of the pattern's amino acidposition pairs within the dataset, etc. The calculation of the cluster score may involve aggregating these factors into a composite score that reflects the pattern's potential biological or clinical relevance. For example, a higher cluster score may indicate a pattern that is more likely to represent a biologically meaningful signal, such as a peptide sequence that is commonly presented on the surface of cancer cells and thus may be a target for therapeutic intervention. In some embodiments, the cluster score may be used to prioritize patterns for further analysis or experimental validation. Patterns with higher cluster scores may be selected for synthesis and testing as potential vaccine candidates or diagnostic markers. The cluster score may also be used to filter out patterns with low scores that are less likely to be of interest, thereby streamlining the search process and focusing resources on the most promising patterns.
[0060] In some embodiments, the approach described herein may include calculating a cluster score for each identified pattern using various alternative methods, each providing a different perspective on the pattern's relevance. Cluster Score per HLA Type: In some embodiments, the cluster score may be calculated separately for each HLA type present in the subject. This approach recognizes that different HLA types may present peptides differently, and thus, the relevance of a pattern may vary depending on the HLA type. The method may involve generating separate scores for each HLA type, resulting in multiple sets of scores and corresponding lists of clusters. This allows for a more personalized assessment of patterns based on the subject's HLAprofile. Impact of Specific Amino Acids on Cluster Score: In some embodiments, the presence of specific amino acids within a pattern, such as cysteine (C), may reduce the cluster score. This approach takes into account the biochemical properties of amino acids that may affect peptide processing or presentation. For example, cysteine's ability to form disulfide bonds, or become oxidized, or otherwise modified, which might render a peptide less likely to be presented by HLA molecules, thereby reducing the pattern's cluster score. Logarithmic Scaling of Peptide Counts: In some embodiments, the cluster score may be calculated using a logarithmic scale of the number of peptides matching the pattern. This method ensures that patterns matching a small number of peptides still contribute meaningfully to the analysis, promoting the identification of antigens across all branches of the search tree. The logarithmic scale balances the influence of large and small clusters on the overall search results. Database Cross-Reference Enhancement: In some embodiments, the cluster score may be increased if the pattern is found in an empirical database such as the Immune Epitope Database (IEDB), which contains peptides identified in other patients independent of their HLA types. If a pattern matches peptides in the database, it may suggest a broader relevance, and thus, the cluster score for that pattern may be increased. Dual HLA Type Binding: In some embodiments, the cluster score may be increased if a pattern is predicted to bind to two different HLA types. This approach recognizes the potential for a peptide to elicit a broader immune response if it can be presented by multiple HLA molecules, increasing the pattern's cluster score to reflect its enhanced immunogenic potential. Uniqueness to HLA Type: In some embodiments, the cluster score may be increased if a pattern is the sole representative found to likely bind to a particular HLA type. This approach values the diversity of antigens presented to the immune system and ensures that antigens capable of binding to each HLA type are included, thereby avoiding the exclusion of potential therapeutic targets.
[0061] Each of these alternative methods for calculating cluster scores may be combined into a weighted sum to provide an overall cluster score for each pattern. The weighted sum approachallows for the integration of the various factors that contribute to a pattern's relevance, with each method assigned a specific weight based on its perceived impact on the pattern's utility. For instance, the cluster score per HL A type might be given a higher weight if the subject's HL A profile is known to have a strong influence on peptide presentation, while the impact of specific amino acids on the cluster score might be weighted less if the peptide processing implications are deemed less consequential. The logarithmic scaling of peptide counts may facilitate that patterns matching a small number of peptides are not overlooked, and this factor can be adjusted in the weighted sum to balance the influence of large and small clusters. Database cross-reference enhancement can be weighted to reflect the added confidence that comes from empirical validation, and dual HLA type binding can be given prominence to emphasize the potential for broader immune responses. The uniqueness of an HLA type can also be weighted to ensure that all HLA types have potential binding antigens represented. By combining these methods into a weighted sum, the overall cluster score encapsulates a comprehensive assessment of each pattern's potential to serve as a biomarker or therapeutic target, providing a robust tool for guiding further research and clinical decision-making.
[0062] FIG. 2 is a diagram illustrating an example of a search graph 200 for identifying peptide patterns in a subject's cells, in accordance with one or more embodiments of the current subject matter. As shown in FIG. 2, peptides may be represented by the amino acid’s letters of their sequences. Each amino acid in a peptide is denoted by a specific letter, corresponding to its place in the standard amino acid alphabet. This representation allows for using data strings to represent peptide sequences, where each character in the string corresponds to an amino acid in the peptide, as defined by the standard single-letter amino acid code. This representation also allows for the systematic analysis of peptide sequences using Al models, which can identify patterns and correlations within the data that may not be readily apparent through manual examination. As show n in FIG. 2, a level of specificity may indicate a number of letters with known locations andidentifications in the data string. For example, element 220 represents a level two (2) specificity that may indicate that there are two amino acids at specific positions within the peptide sequence that are known and fixed. This could be represented as a pattern where tw o positions are occupied by specific amino acids, and the remaining positions are variable or undetermined. For example, node 221, i.e., “ D * > Y”, indicates a level 2 of specificity, because it is known that the third amino acid is a “D”, and the last amino acid is a “Y”, with “*” representing a wildcard. The wildcard may indicate a variable number of unspecified amino acids that can occupy the positions between the known amino acids. In the given pattern, node 221, i.e., “ D * > Y”, the asterisk (“*”) serves as a placeholder for any sequence of amino acids of any length, including a sequence of zero length. This allows for flexibility in the peptide sequence, accommodating various lengths and compositions of peptides that still conform to the known specificity at the defined positions. The underscoresrepresent single positions that can be occupied by any amino acid, further contributing to the variability of the pattern.
[0063] As shown in FIG. 2, element 240 represents a level four (4) specificity that may indicate that there are four amino acids at specific positions within the peptide sequence that are known and fixed. For example, element 241, i.e.. S D K *_ Y” may represent a peptide pattern where the second position is occupied by the amino acid "S", the third position by "D", the sixth position by "K", and the last position by "Y". The underscoresdenote positions that can be filled by any amino acid, and the asterisk ("*") serves as a wildcard for a stretch of any number of amino acids, which can also be of zero length. This pattern, therefore, specifies a sequence with a level of specificity of four (4), as it contains four know n and fixed amino acids at particular locations within the peptide sequence, allowing for the identification of peptides that match this specific arrangement.
[0064] In some embodiments, an input data set may include a number of data strings that represent peptides from a subject’s cells. It should be noted that the peptides are collected fromone patient’s cells to ensure the personalization of the treatment or analysis. This personalization allows for the identification of patterns that are specific to the individual's cellular makeup, which can be particularly beneficial for tailoring medical treatments such as personalized cancer vaccines or targeted therapies. By analyzing peptides that are uniquely presented on a single patient's cells, the method can uncover patient-specific antigens that may serve as precise targets for immunotherapy, thereby increasing the efficacy and reducing potential side effects associated with broader-spectrum treatments. As described herein elsewhere, data strings may be utilized to represent peptide sequences, where each character in the string corresponds to an amino acid in the peptide, as defined by the standard single-letter amino acid code, in some embodiments, a data set representing peptides presented on one subject’s cells may be received, and wherein the subj ect may suffer from one or more types of cancer.
[0065] In some embodiments, once the data set representing the peptides in the subject’s cells are received, the approach disclosed herein may further identify a plurality of root nodes in the data set. In some embodiments, rood nodes are the data strings with a predefined level of specificity. For example, the approach may employ a computational algorithm to analyze the peptide sequences and determine initial patterns with a predefined level of specificity, which serve as the root nodes. In some embodiments, these root nodes may act as starting points for the iterative identification of more specific patterns, or child nodes, which are derived from the root nodes. The identification of root nodes may involve a foundational step in constructing a hierarchical framework of peptide patterns, which can be used to systematically explore the peptide space for potential biomarkers or targets for therapeutic intervention. In some embodiment, the method may include a brute-force identification of root nodes within the data set representing peptides presented on the subject's cells. This process involves exhaustively analyzing the data set to determine initial patterns that meet a predefined level of specificity, which are designated as root nodes. The brute-force approach may ensure that the foundationalpaterns are identified without overlooking any potential root nodes. Alternatively or additionally, the method may employ heuristic algorithms, machine learning techniques, or other advanced computational methods to identify root nodes within the data set representing peptides presented on the subject's cells. These methods may enhance the efficiency of the root node identification process, reducing the computational resources and time compared to brute-force methods.
[0066] As shown in FIG. 2, in some embodiments, the approach described herein may further identify’ patterns. In some embodiments, paterns may include data strings that have a higher level of specificity than the predefined level of specificity of the root nodes it branches from. For example, as shown in FIG. 2, element 230 may include nodes with a level of specificity of three (3). Node 231 and node 232 may be considered as paterns, and they have higher level of specificity than the root node they branch from, which is the node 221. As shown in FIG. 2. node 231 is expressed asL‘_ S D * _ Y which contains one more known amino acid position than the node 221, i.e.. D * > Y” The additional known amino acid position in node 231 is the 'S' at the second position, which is specified, as opposed to the corresponding position in node 221, which is not specified and thus can be any amino acid. In some embodiments, node 231 and node 232 may not themselves constitute patterns; rather, they may serve as intermediate nodes that act as parent nodes to subsequent, more specific paterns. As parent nodes, node 231 and node 232 facilitate the iterative identification of paterns with increasing levels of specificity within the data set. The specificities regarding what constitutes a patern are described in further detail elsew here herein.
[0067] In some embodiments, node 241 may be a patern that branches from node 121 via intermediate nodes such as node 231. Node 241 represents a more specific patern within the hierarchical search structure, having evolved from the more general patern of node 221. The transition from node 221 to node 241 involves the iterative process of increasing specificity bydefining additional amino acid positions in the peptide sequence. This process may include the identification and establishment of intermediate nodes like node 231 and node 232, which potentially sen e as parent nodes to node 241. Each intermediate node contributes to the refinement of the pattern, adding specificity until node 241 is reached with its level four specificity. In some embodiments, the process of iteratively identifying a plurality of patterns from each of the plurality of root nodes involves a recursive or iterative computational method that builds upon the initial patterns (e.g., root nodes or parent nodes) to discover more detailed and specific patterns (e.g., child nodes). Each root node or parent node serves as a starting point for the generation of new patterns that share a common structure with the root node but include additional specified amino acids at particular positions, thereby increasing the specificity of the pattern. The predefined level of specificity for a root node may be a baseline level that is set to ensure that the initial patterns are broad enough to capture a wide range of potential matches within the peptide data set. This level of specificity is characterized by a minimum number of known amino acids at specific positions within the peptide sequence. As the process iterates, each subsequent pattern derived from a root node will have a higher level of specificity than the root node it branches from. This is achieved by introducing additional known amino acids at previously unspecified positions or by further specifying the positions of amino acids that were previously represented by wildcards or placeholders in the root node pattern.
[0068] FIG. 3 is a diagram illustrating a flow chart of a process for identifying peptide patterns for a given subject, in accordance with one or more embodiments of the current subject matter. As shown in FIG. 3, the process 300 may begin with operation 302, wherein the system and / or platform may receive a data set representing peptides presented on one subject’s cells. The data set may include a variety of peptides, each represented by a data string that reflects its amino acid sequence. In some embodiments, the peptides may be derived from a single patient, ensuring the personalization of the subsequent analysis and potential therapeutic interventions. The data setmay be preprocessed to standardize the format of the peptide sequences, facilitating their analysis by the system.
[0069] In some embodiments, the process 300 may proceed to operation 304, wherein the system or platform may identify a plurality of root nodes within the data set. Each root node may be a data string with a predefined level of specificity, serving as a foundational pattern for the identification process. The predefined level of specificity may be determined based on the minimum desired specificity for the initial patterns, which may be set according to the objectives of the analysis or the characteristics of the peptides.
[0070] In some embodiments, the process 300 may continue to operation 306, wherein the system or platform may iteratively identify a plurality of patterns from each of the plurality of root nodes. Each identified pattern may have a higher level of specificity than the predefined level of specificity of the root nodes it branches from. This iterative process may involve the application of heuristic algorithms or other computational methods to systematically explore the peptide space, refining the patterns to increase their specificity and potential biological relevance.
[0071] In some embodiments, the process 300 may conclude with operation 308, wherein the system or platform may output a visualization of the identified patterns upon a determination that a stopping criterion is reached. The stopping criterion may be based on various factors, such as a predetermined minimum number of matching peptides for a pattern or the attainment of a maximum level of specificity. The visualization may be in the form of a directed acyclic graph (DAG), a tree structure, or any other suitable graphical representation that effectively conveys the hierarchical relationships between the patterns and their respective levels of specificity. This visualization may aid researchers or clinicians in analyzing the patterns, exploring their potential as biomarkers or therapeutic targets, and making informed decisions about further research or treatment strategies.
[0072] FIG. 4 is a diagram illustrating a flow chart of a process 400 for training an Al model for identifying peptide patterns of a given subject, in accordance with one or more embodiments of the current subj ect matter. As shown in FIG. 4, the process 400 may begin with operation 402, wherein the system and / or platform may receive training data including a set of known peptide sequences and their associated properties. In some embodiments, the variety' of properties of peptides, which are not limited to, but include binding affinities that describe how strongly each peptide binds to specific molecules like antibodies, receptors, or major histocompatibility7complex (MHC) molecules. It also covers the immunogenicity7, which provides data on the peptides' ability to elicit an immune response, and the frequency of occurrence, detailing how often each peptide appears within a population or the specific patient's cells from which the peptides are derived. Additionally, the data may capture biological relevance, which encompasses any supplementary7information that might be pertinent to the biological functions or roles of the peptides, such as their involvement in particular cellular pathways, disease states, or therapeutic responses. In some embodiments, a data preprocessing stage may be utilized to transform the peptide sequences into a numerical format suitable for the Al model, employing methods such as one-hot encoding or sequence embedding, and involves selecting features that are particularly informative for the pattern recognition task.
[0073] In some embodiments, the process 400 may proceeds to operation 402. wherein the system or platform may train the Al model by automatically deriving model features from the training data, wherein the model feature may include position-specific amino acid frequencies, known epitopes, and biologically relevant sequence motifs. In some embodiments, motifs may comprise conserved sequences w ithin the peptide data that are indicative of a particular function or structural characteristic. In some embodiments the training may involve adjusting the weights of the neural netw ork or parameters of the machine learning model to optimize the identification of peptide patterns that correlate with the desired biological outcomes. This optimization may beguided by a set of rules or algorithms designed to enhance the predictive accuracy of the model without directly involving the calculation of a loss function. The training may also include the use of cross-validation techniques to prevent overfitting and ensure that the model remains generalizable to new data. Next, the process may proceed to operation 306, wherein a loss function is selected to evaluate and adjust the performance of the Al model. For example, the loss function may be chosen based on its ability to quantity' the discrepancy between the predicted patterns and the actual peptide sequences within the training data. This function serves as a guide for the model's learning process, steering the adjustments to the model parameters in a way that minimizes prediction errors and enhances the model's predictive capabilities.
[0074] FIG. 5 depicts a block diagram illustrating a computing system 500 consistent with implementations of the current subject matter. As shown in FIG. 5, the computing system 500 can include a processor 510, a memory 520. a storage device 530. and input / output devices 540. The processor 510, the memory 520, the storage device 530, and the input / output devices 540 can be interconnected via a system bus 550. The computing system 500 may additionally or alternatively include a graphic processing unit (GPU), such as for image processing, and / or an associated memory for the GPU. The GPU and / or the associated memory for the GPU may be interconnected via the system bus 550 with the processor 510, the memory 520, the storage device 530, and the input / output devices 540. The memory associated with the GPU may store one or more images described herein, and the GPU may process one or more of the images described herein. The GPU may be coupled to and / or form a part of the processor 510. The processor 510 is capable of processing instructions for execution within the computing system500. In some implementations of the current subject matter, the processor 510 can be a singlethreaded processor. Alternately, the processor 510 can be a multi-threaded processor. The processor 510 is capable of processing instructions stored in the memory' 520 and / or on thestorage device 530 to display graphical information for a user interface provided via the input / output device 540.
[0075] The memory 520 is a computer readable medium, such as volatile or non-volatile that stores information within the computing system 500. The memory 520 can store data structures representing configuration object databases, for example. The storage device 530 is capable of providing persistent storage for the computing system 500. The storage device 530 can be a floppy disk device, a hard disk device, an optical disk device, or a tape device, or other suitable persistent storage means. The input / output device 540 provides input / output operations for the computing system 500. In some implementations of the current subject matter, the input / output device 540 includes a keyboard and / or pointing device. In various implementations, the input / output device 540 includes a display unit for displaying graphical user interfaces.
[0076] According to some implementations of the current subject matter, the input / output device 540 can provide input / output operations for a network device. For example, the input / output device 540 can include Ethernet ports or other networking ports to communicate with one or more wired and / or wireless networks (e.g., a local area network (LAN), a wide area network (WAN), the Internet).
[0077] In some implementations of the current subject matter, the computing system 500 can be used to execute various interactive computer software applications that can be used for organization, analysis and / or storage of data in various (e.g., tabular) format (e.g., Microsoft Excel®, and / or any other type of software). Alternatively, the computing system 500 can be used to execute any type of software applications. These applications can be used to perform various functionalities, e.g., planning functionalities (e.g.. generating, managing, editing of spreadsheet documents, word processing documents, and / or any other objects, etc.), computing functionalities, communications functionalities, etc. The applications can include various add-in functionalities or can be standalone computing products and / or functionalities. Upon activationwithin the applications, the functionalities can be used to generate the user interface provided via the input / output device 540. The user interface can be generated and presented to a user by the computing system 500 (e.g., on a computer screen monitor, etc.).
[0078] One or more aspects or features of the subject matter described herein can be realized in digital electronic circuitry, integrated circuitry, specially designed framework specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) computer hardware, firmware, software, and / or combinations thereof. These various aspects or features can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to. a storage system, at least one input device, and at least one output device. The programmable system or computing system may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0079] These computer programs, which can also be referred to as programs, software, software frameworks, frameworks, components, or code, include machine instructions for a programmable processor, and can be implemented in a high-level procedural language, an object- oriented programming language, a functional programming language, a logical programming language, and / or in assembly / machine language. As used herein, the term '‘machine-readable medium” refers to any computer program product, apparatus and / or device, such as for example magnetic discs, optical disks, memory, and Programmable Logic Devices (PLDs), used to provide machine instructions and / or data to a programmable processor, including a machine- readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or datato a programmable processor. The machine-readable medium can store such machine instructions non-transitorily, such as for example as would a non-transient solid-state memory' or a magnetic hard drive or any equivalent storage medium. The machine-readable medium can alternatively or additionally store such machine instructions in a transient manner, such as for example as would a processor cache or other random access memory associated with one or more physical processor cores.
[0080] To provide for interaction with a user, one or more aspects or features of the subject matter described herein can be implemented on a computer having a display device, such as for example a cathode ray tube (CRT) or a liquid crystal display (LCD) or a light emitting diode (LED) monitor for displaying information to the user and a keyboard and a pointing device, such as for example a mouse or a trackball, by which the user may provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well. For example, feedback provided to the user can be any form of sensory feedback, such as for example visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including, but not limited to, acoustic, speech, or tactile input. Other possible input devices include, but are not limited to, touch screens or other touch-sensitive devices such as single or multi-point resistive or capacitive trackpads, voice recognition hardware and software, optical scanners, optical pointers, digital image capture devices and associated interpretation software, and the like.
[0081] In the descriptions above and in the claims, phrases such as “at least one of’ or “one or more of’ may occur followed by a conjunctive list of elements or features. The term “and / or” may also occur in a list of two or more elements or features. Unless otherwise implicitly or explicitly contradicted by the context in which it used, such a phrase is intended to mean any of the listed elements or features individually or any of the recited elements or features in combination with any of the other recited elements or features. For example, the phrases “at leastone of A and B;” “one or more of A and B;” and “A and / or B” are each intended to mean “A alone, B alone, or A and B together."’ A similar interpretation is also intended for lists including three or more items. For example, the phrases “at least one of A, B, and C;” “one or more of A, B, and C;” and “A, B, and / or C” are each intended to mean “A alone, B alone, C alone, A and B together, A and C together, B and C together, or A and B and C together.” Use of the term “based on,” above and in the claims is intended to mean, “based at least in part on,” such that an unrecited feature or element is also permissible.
[0082] The subject matter described herein can be embodied in systems, apparatus, methods, and / or articles depending on the desired configuration. The implementations set forth in the foregoing description do not represent all implementations consistent with the subject matter described herein. Instead, they are merely some examples consistent with aspects related to the described subject matter. Although a few variations have been described in detail above, other modifications or additions are possible. In particular, further features and / or variations can be provided in addition to those set forth herein. For example, the implementations described above can be directed to various combinations and subcombinations of the disclosed features and / or combinations and subcombinations of several further features disclosed above. In addition, the logic flows depicted in the accompanying figures and / or described herein do not necessarily require the particular order shown, or sequential order, to achieve desirable results. Other implementations may be within the scope of the follow ing claims.
Claims
CLAIMSWHAT IS CLAIMED IS:
1. A method, comprising: receiving a data set representing peptides presented on one subject’s cells; identifying a plurality of root nodes in the data set, wherein each of the plurality of root nodes is a data string with a predefined level of specificity; iteratively identifying a plurality' of patterns from each of the plurality of root nodes, wherein each of the plurality of patterns has a higher level of specificity than the predefined level of specificity of the root nodes it branches from; and outputting a visualization of the identified patterns upon a determination that a stopping criterion is reached.
2. The method of claim 1, wherein a level of specificity indicates a number of letters with known locations and identifications in the data string.
3. The method of claim 1, wherein the stopping criterion is based on a predetermined minimum number of matching peptides for a pattern.
4. The method of claim 1, wherein the stopping criterion for each pattern is based on a predetermined minimum number of matching peptides, wherein the predetermined minimum number depends on a predicted HLA type associated with the pattern being identified.
5. The method of claim 1. wherein the visualization of the identified patterns comprises a directed acyclic graph (DAG) where nodes represent patterns and edges represent parent-child relationships between patterns.
6. The method of claim 1, further comprising employing a heuristic search algorithm that utilizes a probability distribution based on occurrences of each possible amino acid-position pair within the peptides to guide the identification of the plurality of patterns.
7. The method of claim 3, wherein the stopping criterion comprises stopping at one matching peptide for a pattern.
8. The method of claim 1, further comprising calculating a cluster score for each identified pattern.
9. The method of claim 8, wherein the cluster score factors in a count score, and wherein the count score represents the count of peptides from the subject that belongs to the pattern.
10. The method of claim 8, wherein the cluster score factors in the level of specificity associated with the identified pattern, and wherein a higher level of specificity contributes higher value to the cluster score.1 1. A computer program product comprising a non-transient machine-readable medium storing instructions that, when executed by at least one programmable processor, cause the at least one programmable processor to perform operations comprising: receiving a data set representing peptides presented on one subject’s cells; identifying a plurality of root nodes in the data set, wherein each of the plurality of root nodes is a data stnng with a predefined level of specificity; iteratively identifying a plurality of patterns from each of the plurality of root nodes, wherein each of the plurality of patterns has a higher level of specificity than the predefined level of specificity of the root nodes it branches from; and outputting a visualization of the identified patterns upon a determination that a stopping criterion is reached.
12. The computer program product of claim 11, wherein a level of specificity indicates a number of letters with known locations and identifications in the data string.
13. The computer program product of claim 11, wherein the stopping criterion is based on a predetermined minimum number of matching peptides for a pattern.
14. The computer program product of claim 11, wherein the stopping criterion for each pattern is based on a predetermined minimum number of matching peptides, wherein the predetermined minimum number depends on a predicted HLA type associated with the pattern being identified.
15. The computer program product of claim 11, wherein the visualization of the identified patterns comprises a directed acyclic graph (DAG) where nodes represent patterns and edges represent parent-child relationships between patterns.
16. The computer program product of claim 11, wherein the operations further comprise employing a heuristic search algorithm that utilizes a probability distribution based on occurrences of each possible amino acid-position pair within the peptides to guide the identification of the plurality of patterns.
17. The computer program product of claim 13, wherein the stopping criterion comprises stopping at one matching peptide for a pattern.
18. The computer program product of claim 11, wherein the operations further comprise calculating a cluster score for each identified pattern.
19. The computer program product of claim 18, wherein the cluster score factors in a count score, and wherein the count score represents the count of peptides from the subject that belongs to the pattern.
20. The computer program product of claim 18, wherein the cluster score factors in the level of specificity associated with the identified pattern, and wherein a higher level of specificity contributes higher value to the cluster score.
21. A system comprising: a programmable processor; and a non-transient machine-readable medium storing instructions that, when executed by the processor, cause the at least one programmable processor to perform operations comprising: receiving a data set representing peptides presented on one subject’s cells; identifying a plurality of root nodes in the data set, wherein each of the plurality of root nodes is a data string with a predefined level of specificity; iteratively identifying a plurality' of patterns from each of the plurality of root nodes, wherein each of the plurality of patterns has a higher level of specificity’ than the predefined level of specificity of the root nodes it branches from; and outputting a visualization of the identified patterns upon a determination that a stopping criterion is reached.
22. The system of claim 21, wherein a level of specificity indicates a number of letters with known locations and identifications in the data string.
23. The system of claim 21, wherein the stopping criterion is based on a predetermined minimum number of matching peptides for a pattern.
24. The system of claim 21 , wherein the stopping criterion for each pattern is based on a predetermined minimum number of matching peptides, wherein the predetermined minimum number depends on a predicted HLA type associated with the pattern being identified.
25. The system of claim 21, wherein the visualization of the identified patterns comprises a directed acyclic graph (DAG) where nodes represent patterns and edges represent parent-child relationships between patterns.
26. The system of claim 21, wherein the operations further comprise employing a heuristic search algorithm that utilizes a probability distribution based on occurrences of eachpossible amino acid-position pair within the peptides to guide the identification of the plurality of patterns.
27. The system of claim 23, wherein the stopping criterion comprises stopping at one matching peptide for a pattern.
28. The system of claim 21, wherein the operations further comprise calculating a cluster score for each identified pattern.
29. The system of claim 28, wherein the cluster score factors in a count score, and wherein the count score represents the count of peptides from the subject that belongs to the pattern.
30. The system of claim 28, wherein the cluster score factors in the level of specificity associated with the identified pattern, and wherein a higher level of specificity contributes higher value to the cluster score.
Citation Information
Patent Citations
Systems and methods for de novo peptide sequencing from data-independent acquisition using deep learning
US20190147983A1
Peptide search system for immunotherapy
US20240071570A1