Systems and methods for designing vaccines

The system uses machine learning to predict optimal molecular sequences for vaccines, addressing the limitations of traditional methods by enhancing vaccine effectiveness against multiple pathogen strains through advanced predictive modeling.

JP7801214B2Active Publication Date: 2026-01-16サノフィ ワクチンズ ユーエス インコーポレイテッド
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2022523426
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-10-21
Filing Date
2020-10-20
Publication Date
2026-01-16
Estimated Expiration
2040-10-20

AI Technical Summary

Technical Problem

Traditional methods for selecting candidate vaccines are inadequate in addressing the complexity of current disease seasons with thousands of pathogenic isolates, leading to suboptimal vaccine effectiveness, as they rely on assumptions that were valid when fewer isolates were present, such as dominant strain selection and ferret cross-reactivity as a predictor of human efficacy.

Method used

A system utilizing machine learning models, including recurrent neural networks, to predict molecular sequences that generate optimal biological responses across multiple translational axes, selecting driver models that maximize aggregate immune response or coverage against multiple pathogen strains, validated through real-world experiments.

Benefits of technology

The system designs vaccines that confer broader and more effective protection against future disease seasons by predicting responses to rarely observed strains, improving vaccine efficacy beyond conventional techniques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007801214000001
    Figure 0007801214000001
  • Figure 0007801214000002
    Figure 0007801214000002
  • Figure 0007801214000003
    Figure 0007801214000003
Patent Text Reader

Abstract

A system for designing a vaccine includes one or more processors and a computer storage device storing executable computer instructions that, when executed by the one or more processors, cause the one or more processors to perform one or more operations. The one or more operations include applying a plurality of driver models configured to generate output data representing one or more molecular sequences to a first time series data set. The one or more operations include training a driver model for each of the plurality of driver models. The one or more operations include selecting a set of trained driver models from the plurality of driver models based on one or more trained translational responses. The one or more operations include selecting a subset of trained driver models from the set of trained driver models based on second translational response data.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 62 / 924,096, filed October 21, 2019, the entire contents of which are incorporated herein by reference.

[0002] The present disclosure relates generally to systems and methods for producing vaccines. [Background technology]

[0003] The mammalian immune system uses two general mechanisms to defend the body against environmental pathogens: upon encountering a pathogen-derived molecule, an immune response is activated to ensure protection against that pathogen.

[0004] The first immune system mechanism is the nonspecific (or innate) inflammatory response. The innate immune system appears to recognize specific molecules that are present on pathogens but not on the body itself.

[0005] The second immune system mechanism is the specific or acquired (or adaptive) immune response. The innate response is essentially the same for each injury or infection. In contrast, the acquired response is generated specifically in response to molecules in or derived from pathogens. The immune system recognizes and responds to structural differences between self proteins and non-self (e.g., pathogen or pathogen-derived) proteins. Proteins that the immune system recognizes as non-self are called antigens. Pathogens typically express numerous, highly complex antigens. The acquired immune system serves two functions: first, it produces immunoglobulins (antibodies) in response to the many different molecules, called antigens, present on pathogens; and second, it recruits receptors that bind to processed forms of antigens displayed on the cell surface, allowing other cells to identify infected cells.

[0006] In summary, adaptive immunity is mediated by specialized immune cells called B and T lymphocytes (or simply, B and T cells). Adaptive immunity has a specific memory for antigenic structures. Repeated exposure to the same antigen can elevate the response, thereby increasing the level of induced defense against that particular pathogen. B cells generate and mediate their function through the action of antibodies. B cell-dependent immune responses are called "humoral immunity" because antibodies are found in bodily fluids. T cell-dependent immune responses are called "cell-mediated immunity" because effector activity is directly mediated by the local action of effector T cells. The local action of effector T cells is amplified by synergistic interactions between T cells and secondary effector cells, such as activated macrophages. As a result, pathogens are killed and prevented from causing disease.

[0007] Like pathogens, vaccines work by initiating an innate immune response at the site of vaccination and activating antigen-specific T and B cells that can give rise to long-term memory cells in secondary lymphoid tissues. The correct interaction of the vaccine with cells at the site of vaccination, as well as with T and B cells, is critical to the ultimate success of the vaccine.

[0008] To determine whether a candidate antigen can be a functional and effective vaccine, the candidate antigen usually needs to undergo rigorous testing and evaluation protocols. Traditionally, candidate antigens are preclinically tested by in vitro assays, ex vivo assays, and various animal models (e.g., mouse models, ferret models, etc.) to evaluate the candidate antigen.

[0009] One exemplary type of assay that can be used to measure biological responses is the hemagglutination inhibition assay (HAI). HAI employs a process known as hemagglutination, in which sialic acid receptors on the surface of red blood cells (RBCs) bind to the hemagglutinin glycoprotein found on the surface of influenza viruses (and some other viruses), creating a network, or lattice structure, of interconnected red blood cells and virus particles, called hemagglutination. This hemagglutination occurs in a concentration-dependent manner for virus particles. HAI is a physical measurement that serves as a proxy for the ability of a virus to bind to similar sialic acid receptors on pathogen target cells in the body. The introduction of antiviral antibodies, generated in a human or animal immune response to another virus (which may be genetically similar or different from the virus used to bind to RBCs in the assay), disrupts the interaction of the virus with RBCs and alters the concentration of the virus sufficiently to alter the concentration at which hemagglutination is observed in the assay. One goal of HAI may be to characterize the concentration of antibody in antiserum, or other antibody-containing samples, relative to the antibody's ability to induce hemagglutination in an assay. The highest dilution of antibody that prevents hemagglutination is called the HAI titer (i.e., the assessed response).

[0010] Another approach to measuring biological responses is to measure the larger set of possible antibodies elicited by a human or animal immune response. This set is not necessarily capable of affecting hemagglutination in an HAI assay. A common approach for this measurement utilizes enzyme-linked immunosorbent assay (ELISA) technology, in which a viral antigen (e.g., hemagglutinin) is immobilized on a solid surface, after which antibodies from antisera bind to the antigen. The readout measures the catalytic activity of an exogenous enzyme substrate conjugated to either the antibody from the antisera or to another antibody that itself binds to the antibody in the antisera. The catalytic activity of the substrate produces a readily detectable product. There are many variations of this type of in vitro assay. One such variation is called antibody forensics (AF); it is a multiplexed bead array technique that allows a single serum sample to be compared simultaneously with many antigens. These assays characterize the concentration and total antibody recognition compared to HAI titers, which are understood to be specifically related to interference with sialic acid binding by the hemagglutinin molecule. Thus, the antibodies in the antisera may, in some cases, be proportionally higher or lower in measurement relative to the hemagglutinin molecule of one virus than the corresponding HAI titer of the hemagglutinin molecule of another virus; in other words, these two measurements, AF and HAI, are generally not linearly related.

[0011] Currently, traditional candidate antigen testing is only performed with the conditional assumption of eliciting a predetermined "protective" immune response. That is, if an animal or assay fails to demonstrate an adequate response to a candidate antigen, the candidate antigen is typically "thinned out" (i.e., discarded as a productive candidate). For example, influenza antigens are often tested using sequential selection protocols, in which the antigen is first evaluated using in vitro assays to ensure that the antigen is amenable to large-scale production. Provided that the antigen meets these requirements, it is then evaluated, for example, by immunization of mice to measure the antigen's ability to elicit a protective immune response in the mice. This response is typically expected to be protective against the antigen itself and various other virus strains and / or virus strain components against which protection is desired. Ferrets are then similarly evaluated, subject to previous measurements in mice or other models that have previously demonstrated what is understood to be indicative of a protective response. Penultimately, ex vivo platforms such as human immune system replicas or non-human primates are evaluated; again, subject to success in the previous steps. Summary of the Invention [Means for solving the problem]

[0012] In one aspect, a system for designing a vaccine is provided. The system includes one or more processors. The system includes a computer storage device storing executable computer instructions that, when executed by the one or more processors, cause the one or more processors to perform one or more operations. The one or more operations include applying a plurality of driver models to a first time series dataset, the driver models being configured to generate output data representing one or more molecular sequences, the first time series dataset representing the one or more molecular sequences and, for each of the one or more molecular sequences, one or more circulating periods of pathogen strains that include the molecular sequence as a natural antigen. The one or more operations include, for each of the plurality of driver models, training the driver model by: i) receiving from the driver model output data representing one or more predicted molecular sequences based on the received first time-series data set; ii) applying a translational model configured to predict biological responses to the molecular sequences along a plurality of translational axes to the output data representing the predicted one or more molecular sequences to generate first translational response data representing one or more first translational responses corresponding to a particular translational axis of the plurality of translational axes based on the one or more predicted molecular sequences in the output data; iii) adjusting one or more parameters of the driver model based on the first translational response data; and iv) repeating steps i-iii a certain number of iterations to generate learned translational response data representing one or more learned translational responses corresponding to the particular translational axis. The one or more operations include selecting a set of learned driver models from the plurality of driver models based on the one or more learned translational responses.The one or more operations include, for each trained driver model in the set of trained driver models: applying the trained driver model to the second time series data set to generate trained output data representing one or more predicted molecular sequences for a particular season; applying a translational model to the final output data to generate second translational response data representing one or more second translational responses for each translational axis of the plurality of translational axes; and selecting a subset of trained driver models in the set of trained driver models based on the second translational response data.

[0013] At least one of the plurality of driver models may include a recurrent neural network. At least one of the plurality of driver models includes a long-short-term memory recurrent neural network.

[0014] The output data representing one or more predicted molecular sequences based on the received first time series dataset may include output data representing antigens for each of a plurality of disease seasons. The output data representing antigens for each of a plurality of disease seasons may include antigens determined by predicting molecular sequences that will generate a maximized aggregate biological response across all pathogen strains circulating in a particular season. The output data representing antigens for each of a plurality of disease seasons may include antigens determined by predicting molecular sequences that will generate a response that effectively immunizes against a maximum number of viruses circulating in a particular season.

[0015] The multiple translational axes may include at least one of: a ferret antibody forensics (AF) axis, a ferret hemagglutination inhibition (HAI) axis, a mouse AF axis, a mouse HAI axis, a human replica AF axis, a human AF axis, or a human HAI axis. The number of iterations is based on a predetermined number of iterations. The number of iterations is based on a predetermined error value. The one or more first translational reactions may include at least one of: a predicted ferret HAI titer, a predicted ferret AF titer, a predicted mouse AF titer, a predicted mouse HAI titer, a predicted human replica AF titer, a predicted human AF titer, or a predicted human HAI titer.

[0016] The act of selecting a set of learned driver models from the plurality of driver models may include assigning each driver model of the plurality of driver models to a driver model class, each class being associated with a particular translational axis of the plurality of translational axes used to train the driver model. The act of selecting a set of learned driver models from the plurality of driver models may include, for each driver model of the plurality of driver models, comparing one or more learned translational responses of the driver model to one or more learned translational responses of at least one other driver model assigned to the same class as the driver model.

[0017] The operations may further include, for each learned driver model of the subset of learned driver models: validating the learned driver model by comparing second translational response data corresponding to the learned driver model with the observed experimental response data; and in response to the operation of validating the learned driver model, producing a vaccine that includes one or more molecular sequences represented by the learned output data corresponding to the learned driver model.

[0018] In one aspect, a system is provided. The system includes a computer-readable memory containing computer-executable instructions. The system includes at least one processor configured to execute executable logic including at least one machine learning model trained to predict one or more molecular sequences, wherein the at least one processor is configured to perform one or more operations when the at least one processor is executing the computer-executable instructions. The one or more operations include receiving time-series data indicating one or more molecular sequences and, for each of the one or more molecular sequences, one or more circulating periods of pathogen strains containing the molecular sequence as a natural antigen. The one or more operations include processing the time-series data through one or more data structures that store one or more portions of the executable logic included in the machine learning model to predict the one or more molecular sequences based on the time-series data.

[0019] Predicting one or more molecular sequences based on the time series data may include predicting one or more immunological properties that the predicted one or more molecular sequences will confer for future use. Predicting one or more molecular sequences based on the time series data may include predicting one or more molecular sequences that will generate a maximized aggregate biological response across all pathogen strains in the time series data. Predicting one or more molecular sequences based on the time series data may include predicting one or more molecular sequences that will generate a biological response that effectively covers the greatest number of pathogen strains in the time series data. The predicted one or more molecular sequences are used to design vaccines against pathogen strains circulating during a period following one or more circulation periods of the time series data.

[0020] The machine learning model may include a recurrent neural network.

[0021] These and other aspects, configurations, and implementations may be expressed as methods, apparatus, systems, components, program products, ways of doing business, means or steps for performing a function, and in other ways, and will become apparent from the following description, including the claims.

[0022] Embodiments of the present disclosure may provide one or more of the following advantages: Compared to conventional techniques, the vaccine is designed to confer more protection for a future disease season in terms of the amount of biological response to at least one pathogen strain in that future disease season; Compared to conventional techniques, the vaccine is designed for a future disease season to confer more protection in terms of the breadth of effective coverage against multiple pathogen strains in that future disease season (i.e., to induce an effective immunological response against several pathogen strains in the future disease season); Unlike conventional techniques, rarely observed strains, which may confer "more protection" because they cross-react with more strains than frequently observed strains, are evaluated, and vaccination efficacy for those strains is predicted.

[0023] These and other aspects, configurations, and implementations may be expressed as methods, apparatus, systems, components, program products, means or steps for performing a function, or in other ways.

[0024] These and other aspects, configurations, and implementations will become apparent from the following description, including the claims. [Brief explanation of the drawings]

[0025] [Figure 1] FIG. 1 illustrates an example of a system for designing a vaccine. [Figure 2A] FIG. 1 is a flow diagram of a method for designing a vaccine design system. [Figure 2B] FIG. 1 is a flow diagram of a method for designing a vaccine design system. [Figure 3] 1 is a flow chart of a method for designing a vaccine. [Figure 4] 1 is a flowchart of a method for training one or more driver models to design a vaccine. [Figure 5] 1 is a diagram showing improvements by translational axis compared to conventional techniques for designing vaccines. [Figure 6] FIG. 1 illustrates an example of a system for predicting biological responses using machine learning techniques. [Figure 7] 1 is a flowchart illustrating an example of a method for predicting biological responses using machine learning techniques. [Figure 8] 1 is an example of data used to train a machine learning model to predict biological responses. [Figure 9] FIG. 1 is a flow diagram of an example of training a machine learning model to predict biological responses. DETAILED DESCRIPTION OF THE INVENTION

[0026] Traditional methods for selecting candidate vaccines (CVs) and / or their antigens expressed as recombinant proteins generally rely on several assumptions. As an illustrative example, in the case of influenza, traditional methods for selecting CVs can assume the following: (1) there is a "dominant strain" for a given outbreak season; (2) naive ferrets are an accurate model of influenza drift (i.e., cross-reactivity in ferrets demonstrates whether one CV as an antigen confers protection against other circulating influenza strains); and (3) the acquisition of ferret cross-reactivity can be a reliable predictor of the acquisition of human vaccine efficacy. Based on these assumptions, traditional methods for selecting CVs can have the following solutions: (1) select a CV that protects against the dominant strain; (2) establish a correlate of protection, for example, using ferret HAI; and (3) evaluate the cross-reactivity of clinical isolates in ferrets. Furthermore, traditional methods for selecting CVs typically involve selecting a CV that was circulating in the year prior to the vaccine recommendation year and evaluating the selected CV against other frequently observed pathogenic strains (usually in ferrets).

[0027] While these assumptions may have facilitated effective CVV selection over 50 years ago, when 1–10 pathogenic isolates were observed per year, these assumptions cannot facilitate effective CVV selection in current disease seasons, when thousands of pathogenic isolates are observed and reported. This is because it can be difficult to scale ferret evaluations to thousands of pathogen isolates. As a result, in some cases, for example, current selection of seasonal influenza vaccines typically achieves vaccine effectiveness (i.e., the percentage reduction in severe disease among case-detected individuals in vaccinated populations compared with unvaccinated populations) of less than 50%.

[0028] The systems and methods described herein can be used to alleviate one or more of the aforementioned drawbacks of conventional CV selection techniques. According to the systems and methods described herein, a subset of an initial plurality of machine learning models (also referred to herein as driver models) is used to select one or more molecular sequences (e.g., antigen sequences) that are predicted to excel in at least one translational axis. The translational axis can represent, for example, an evaluation criterion of the biological response of a human or non-human model to an antigen (e.g., the resulting HAI titer of a mouse exposed to a particular antigen, or the resulting HAI titer of collected human serum, etc.). The subset of driver models is selected for rational use by first assigning each driver model from the initial plurality of driver models to one class of translational axes, each class of translational axes corresponding to one translational axis of a plurality of translational axes (e.g., at least one of the following: ferret AF, ferret HAI, mouse AF, mouse HAI, human replica AF, human AF, or human HAI).

[0029] In some embodiments, each driver model is trained to predict molecular sequences that will generate a maximal (e.g., maximized) biological response (e.g., maximized mouse HAI titers) among all circulating pathogen strains in a particular disease season, or that will generate a response that effectively covers the maximum number of circulating pathogen strains in a particular disease season, based on time-series data representing multiple molecular sequences and, for each molecular sequence, the circulating duration of pathogen strains that contain that molecular sequence as a native antigen. In some embodiments, for each driver model, translational models configured to predict biological responses to molecular sequences across multiple translational axes are used to provide feedback in the form of translational response data representing one or more translational responses corresponding to the translational axis classes assigned to the driver model.

[0030] This process is carried out over a number of iterations, with the driver model updating one or more parameters (often referred to as weights and biases) based on feedback from the translational model. After the number of iterations, a set of trained driver models is selected. The selected set of trained driver models may include, for each class of translational axis, a trained driver model that predicted the molecular sequence that results in the desired (often: highest) aggregate (e.g., averaged) biological response (e.g., immune response) predicted by the translational models for that class of translational axis. For each trained driver model in the set of trained driver models, the antigen predicted by that trained driver model is then applied to a translational model that predicts the response to that antigen for each translational axis.

[0031] Next, a subset of trained driver models from the set of trained driver models is selected. Selecting the subset of trained driver models may include, for each translational axis, selecting a trained driver model from the set of trained driver models that predicted the antigen that elicits the highest aggregate biological response across all pathogen strains for a particular pathogenic season predicted by the translational model for that translational axis. Each trained driver model from the subset of trained driver models is validated using observational data from human or non-human experiments. If the trained driver model is validated, the trained driver model is used to design a vaccine based on the antigen predicted by the validated trained driver model.

[0032] In the drawings, a particular arrangement or order of schematic elements, such as those representing devices, modules, instruction blocks, and data elements, is shown for ease of explanation. However, those skilled in the art should understand that the particular order or arrangement of schematic elements in the drawings does not imply that a particular order or sequence of operations, or separation of operations, is required. Furthermore, the inclusion of a schematic element in a drawing does not imply that such element is required in all embodiments, or that the structure represented by such element is not included in or combined with other elements in some embodiments.

[0033] Furthermore, when a connecting element, such as a solid or dashed line or arrow, is used in the drawings to illustrate a connection, relationship, or association between two or more other schematic elements, the absence of any such connecting element does not imply that the connection, relationship, or association cannot exist. In other words, some connections, relationships, or associations between elements may not be shown in the drawings so as not to obscure the disclosure. Additionally, for ease of illustration, a single connecting element may be used to represent multiple connections, relationships, or associations between elements. For example, when a connecting element represents communication of signals, data, or instructions, it should be understood by those skilled in the art that such element represents one or more signal paths (e.g., buses) to affect the required communication.

[0034] Reference will now be made in detail to the embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the various embodiments described. However, it will be apparent to those skilled in the art that the various embodiments described may be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.

[0035] Below are described several configurations, each of which can be used independently of one another or with any combination of the other configurations. However, any individual configuration may not address all of the problems discussed above, or may only address one of the problems discussed above. Some of the problems discussed above may not be completely solved by any of the configurations described herein. Even if a heading is provided, data related to a particular heading may not be found in the section bearing that heading, but may be found elsewhere in this specification.

[0036] 1 shows an example of a system 100 for designing a vaccine. System 100 includes a computer processor 110. Computer processor 110 includes computer-readable memory 111 and computer-readable instructions 112. System 100 also includes a machine learning system 150. Machine learning system 150 includes a machine learning model 120. Machine learning system 150 may be separate from computer processor 110 or may be integrated with computer processor 110.

[0037] The computer readable memory 111 (or computer readable medium) may include any data storage technology type suitable for the local technology environment, including but not limited to semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, removable memory, disk memory, flash memory, dynamic random access memory (DRAM), static random access memory (SRAM), electronically erasable programmable read-only memory (EEPROM), etc. In one embodiment, the computer readable memory 111 includes code segments having executable instructions.

[0038] In some embodiments, computer processor 110 includes a general-purpose processor. In some embodiments, computer processor 110 includes a central processing unit (CPU). In some embodiments, computer processor 110 includes at least one application-specific integrated circuit (ASIC). Computer processor 110 may also include a general-purpose programmable microprocessor, a special-purpose programmable microprocessor, a digital signal processor (DSP), a programmable logic array (PLA), a field programmable gate array (FPGA), special-purpose electronic circuitry, or the like, or a combination thereof. Computer processor 110 is configured to execute program code means, such as computer-executable instructions 112. In some embodiments, computer processor 110 is configured to execute machine learning model 120.

[0039] The computer processor 110 is configured to receive a time series dataset 161. The time series dataset 161 may include data representing one or more molecular sequences and, for each of the one or more molecular sequences, one or more circulation periods of pathogen strains that include the molecular sequence as a natural antigen. As an illustrative example, the time series dataset 161 may represent molecular sequences and circulation periods (e.g., specific months, specific seasons of onset, etc.) for A / SINGAPORE / INFIMH160019 / 2016, A / MISSOURI / 37 / 2017, A / KENYA / 105 / 2017, A / MIYAZAKI / 89 / 2017, A / ETHIOPIA / 1877 / 201, A / OSORNO / 60580 / 2017, A / BRISBANE / 1059 / 2017, and A / VICTORIA / 11 / 2017. Although only eight pathogen strains have been described, the time series dataset 161 may contain molecular sequence information and circulation periods corresponding to billions of pathogen strains. The time series dataset 161 is obtained via one or more means, such as wired or wireless communication with a database (including cloud-based environments), fiber optic communication, universal serial bus (USB), read-only memory (CD-ROM), etc.

[0040] In machine learning system 150, machine learning techniques are applied to train machine learning models 120, which, when applied to input data, generate an indication of whether an input data item has a relevant property, such as the probability that the input data item has a particular Boolean property, an estimate of a scalar property, or an estimate of a vector (i.e., an ordered combination of multiple scalars).

[0041] As part of training the machine learning model 120, the machine learning system 150 can form a training set of input data by identifying a positive training set of input data items determined to have the property of interest, and in some embodiments, a negative training set of input data items that lack the property of interest.

[0042] The machine learning system 150 extracts configuration values ​​from the input data of a training set, where these configurations are variables that are deemed potentially relevant to whether an input data item has an associated property. An ordered list of configurations of the input data is referred to herein as a configuration vector of the input data. In some implementations, the machine learning system 150 applies dimensionality reduction (e.g., by linear discriminant analysis (LDA), principal component analysis (PCA), learned deep configurations from neural networks, etc.) to reduce the amount of data in the configuration vector of the input data to a smaller, more representative set of data.

[0043] In some embodiments, machine learning system 150 uses supervised machine learning to train machine learning model 120, with constituent vectors from a positive training set and a negative training set as input. Various machine learning techniques are used in some embodiments, such as linear support vector machines (linear SVMs), boosting other algorithms (e.g., AdaBoost), neural networks, logistic regression, naive Bayes, memory-based learning, random forests, bagged trees, decision trees, boosted trees, or boosted stumps. When applied to the constituent vectors extracted from the input data items, machine learning model 120 outputs an indication of whether the input data items have a property of interest, such as a Boolean yes / no estimate, a scalar value representing a probability, a vector of scalar values ​​representing multiple properties, or a nonparametric distribution of scalar values ​​representing a discrete and non-empirical fixed number of properties, where the indication is expressed explicitly or implicitly in a Hilbert space or similar infinite-dimensional space.

[0044] In some embodiments, the validation set is formed from additional input data other than the input data in the training set that has already been determined to have or lack the property of interest. The machine learning system 150 applies the trained machine learning model 120 to the data in the validation set to quantify the accuracy of the machine learning model 120. Common metrics applied to measure accuracy include precision = TP / (TP + FP) and recall = TP / (TP + FN), where precision is the number of correct predictions (TP, i.e., true positives) made by the machine learning model 120 out of the total number of predictions made by the machine learning model 120 (TP + FP, i.e., false positives), and recall is the number of correct predictions (TP, i.e., false negatives) made by the machine learning model 120 out of the total number of input data items that had the property of interest (TP + FN, i.e., false negatives). The F-score (F-score = 2 × PR / (P + R)) unifies precision and recall into a single evaluation metric. In some implementations, the machine learning system 150 iteratively retrains the machine learning model 120 until a stopping condition occurs, such as an accuracy measurement indication that the model 120 is sufficiently accurate or after a number of training rounds have been completed.

[0045] In some implementations, the machine learning model 120 includes a neural network. In some implementations, the neural network includes a recurrent neural network (RNN). RNNs generally describe a class of artificial neural networks in which connections between nodes form a directed graph over time, thereby enabling them to exhibit temporally dynamic behavior. Unlike feedforward neural networks, RNNs can use their internal state (memory) to process input sequences. In some implementations, the RNN includes a long short-term memory (LSTM) architecture. LSTM refers to an RNN architecture that has feedback connections and can process entire sequences of data (such as audio or video) rather than just single data points (such as images). The machine learning model 120 may include other types of neural networks, such as convolutional neural networks, radial basis function neural networks, and physical neural networks (e.g., optical neural networks). Exemplary methods for designing and training the machine learning model 120 are discussed in more detail below with reference to FIGS. 2A-4.

[0046] The machine learning model 120 is configured to predict one or more molecular sequences and what immunological properties the predicted one or more molecular sequences will confer for future use based on the received time series dataset 161. As an illustrative example, assume that the received time series dataset 161 includes data representing multiple pathogenic strains, each pathogenic strain known to have circulated at one or more time points between January 1, 2014 and December 31, 2018. The machine learning model 120 may predict one or more molecular sequences (e.g., antigens) that will generate a maximized aggregate biological response (e.g., maximized mean human HAI titer) among all viruses circulating between January 1, 2019 and May 31, 2019, based on the pathogenic strains known to have circulated at one or more time points between January 1, 2014 and December 31, 2018. Additionally or alternatively, the machine learning model 120 may predict one or more molecular sequences that will generate a biological response that will effectively cover (e.g., effectively vaccinate) the greatest number of viruses circulating between January 1, 2019 and May 31, 2019, based on pathogenic strains known to be circulating at one or more points between January 1, 2014 and December 31, 2018. The predicted one or more molecular sequences can be used to design vaccines against viruses circulating in the future (such as between January 1, 2019 and May 31, 2019 in the previous example).

[0047] 2A-2B show a flow diagram of an architecture 200 for designing a system for designing vaccines. The architecture 200 includes a plurality of driver models 210, a translational model 220, and a feedback selection module 230. First, the plurality of driver models 210 are initiated. Each of the plurality of driver models 210 is configured to generate data representing one or more molecular sequences (e.g., antigens) and a prediction as to what immunological properties each of the molecular sequences will confer for use, as discussed with reference to the machine learning model 120 of FIG. 1. In the illustrated embodiment, the plurality of driver models 210 include a first driver model 210a, a second driver model 210b, a third driver model 210c, a fourth driver model 210d, a fifth driver model 210e, a sixth driver model 210f, a seventh driver model 210g, an eighth driver model 210h, a ninth driver model 210i, and a tenth driver model 210j. Although ten driver models are shown, the plurality of driver models 210 may include more or fewer driver models (e.g., five driver models, thirty driver models, one hundred driver models, etc.) One or more of the driver models may be, for example, an RNN as previously described with reference to FIG.

[0048] Translational model 220 is configured to predict biological responses to molecular sequences along multiple translational axes. In the illustrated embodiment, translational model 220 includes ferret HAI translational axis 220a, ferret AF translational axis 220b, mouse HAI axis 220c, mouse AF translational axis 220d, and human replica AF translational axis 220e. While specific translational axes are illustrated, embodiments are not limited to these specific translational axes. For example, translational model may additionally or alternatively include a human HAI translational axis, a human AF translational axis, a human replica HAI axis, or combinations thereof, among others. Some embodiments of translational model 220 are discussed in more detail below with reference to Figures 6-9.

[0049] 2A , each of the driver models of the plurality of driver models 210 is assigned to a particular translational axis of the translational model 220. In the illustrated embodiment, the first driver model 210a and the third driver model 210c are assigned to the ferret HAI translational axis 220a, the second driver model 210b and the sixth driver model 210f are assigned to the ferret AF translational axis 220b, the fourth driver model 210d and the eighth driver model 210h are assigned to the mouse HAI translational axis 220c, the fifth driver model 210e and the ninth driver model 210i are assigned to the mouse AF translational axis 220d, and the seventh driver model 210g and the tenth driver model 210j are assigned to the human replica AF translational axis 220e.

[0050] Each driver model of the plurality of driver models 210 receives a first time series dataset 201. The first time series dataset 201 may include a plurality of molecular sequences and the circulating periods of pathogen strains that include at least one of the plurality of molecular sequences as a natural antigen. As an illustrative example, the first time series dataset 201 may include the molecular sequences and circulating periods of all observed pathogen strains that circulated during a period between January 1, 2014 and December 31, 2018 (also referred to as the "epidemic season"). Based on the received first time series dataset 201, each driver model of the plurality of driver models 220 may generate output data representing one or more molecular sequences. For example, the output data may represent molecular sequences (e.g., antigens) for each outbreak season of the outbreak season. For each outbreak season, molecular sequences can be determined by predicting the molecular sequences that will generate a maximized aggregate biological response across all circulating viruses in that outbreak season and / or that will generate a response that effectively covers (e.g., effectively vaccinates) the maximum number of circulating viruses in that outbreak season based on phylogenetic data from one or more outbreak seasons prior to that outbreak season.

[0051] The translational model 220 can receive output data from each driver model of the plurality of driver models 210 and generate, for each driver model of the plurality of driver models 210, first translational response data representing one or more translational responses corresponding to the particular translational axis assigned to that driver model. In the illustrated example, the translational model 220 receives output data representing one or more predicted molecular sequences from the first driver model 210a and can predict ferret HAI titers for each molecular sequence among the one or more molecular sequences across all pathogen strains circulating in each disease season according to the ferret HAI translational axis 220a (i.e., for each pathogen strain in a particular disease season, predict the immune response of ferrets exposed to that pathogen strain after being immunized with the predicted molecular sequence).

[0052] The first translational response data corresponding to each driver model of the plurality of driver models 210 is received by a feedback selection module 230, which compares the predicted response for each outbreak season to a threshold response. For example, the feedback selection module 230 can aggregate (e.g., average) the predicted biological responses across all viruses for each outbreak season for each driver model, compare the aggregate response to a threshold aggregate response, and generate an error value based on the comparison. Additionally or alternatively, the feedback selection module 230 can compare, for each driver model, the number of viruses effectively vaccinated against to a threshold number for each outbreak season and generate an error value based on the comparison. The feedback selection module 230 can then cause each driver model to adjust one or more parameters (e.g., driver model weights and biases) based on the error value for each outbreak season. This process is repeated for a number of iterations. The number of iterations may be a set number of iterations or may be determined based on a threshold error value (i.e., the process continues until the threshold error value is exceeded). Thus, at one high level: (1) for a particular disease season, each driver model can predict one or more molecular sequences to be used to immunize against the pathogen strain of that particular disease season based on the pathogen strain of the preceding disease season; (2) the performance of each driver model is evaluated for each disease season; and (3) the parameters of each driver model are adjusted based on the model's performance during each disease season.

[0053] After that number of iterations, the performance of each driver model (sometimes referred to herein as a trained driver model) is compared to other driver models assigned to the same translational axis as the driver model, and the driver model exhibiting the best performance is selected to generate a selected set of trained driver models 240. For example, after that number of iterations, the aggregate predicted ferret HAI titers of the molecular sequences predicted by the first driver model 210a are compared to the aggregate predicted ferret HAI titers of the molecular sequences predicted by the third driver model, and the feedback selection module 230 can select the driver model corresponding to the highest aggregate predicted ferret HAI titer (or the greatest number of pathogen strains that have been effectively vaccinated against) over all or part of the disease season. In the illustrated embodiment, the selected set of driver models 240 includes the first driver model 210a, the second driver model 210b, the fifth driver model 210e, the seventh driver model 210g, and the tenth driver model 210j.

[0054] Referring to FIG. 2B , each of the selected set of driver models 240 receives the second time series dataset 202 and generates trained output data representing one or more molecular sequences for a particular disease season based on the second time series dataset 202. Similar to the first time series dataset 201, the second time series dataset 202 may include data representing the molecular sequences and circulating periods of all observed pathogen strains circulating during a given disease season. The disease seasons of the second time series dataset 202 may be the same as or different from the disease seasons of the first time series dataset 201. Each of the driver models in the selected set of driver models 240 can predict one or more molecular sequences (e.g., antigens) for one or more disease seasons. In some embodiments, the predicted one or more molecular sequences are for one of the disease seasons of a temporal period (e.g., the most recent disease season). As an illustrative example, assume that the received second time series dataset 202 includes data representing multiple pathogen strains, each pathogen strain known to have circulated at one or more time points between January 1, 2014 and April 31, 2018. Each of the driver models in the selected set of driver models 240 can predict one or more molecular sequences (e.g., antigens) that generate a maximized aggregate biological response across all viruses circulating between October 1, 2017 and April 31, 2019, based on pathogen strains known to have circulated in the preceding outbreak season between January 1, 2014 and September 30, 2017. Additionally or alternatively, each of the driver models in the selected set of driver models 240 may predict one or more molecular sequences that generate a biological response that effectively covers (e.g., effectively vaccinates) the maximum number of viruses circulating between October 1, 2017 and April 31, 2018 based on the pathogenic strains known to be circulating during the preceding outbreak season between January 1, 2014 and September 30, 2017.

[0055] The translational model 220 receives trained output data from each of the driver models in the selected set of driver models 240 and generates second translational response data for each driver model based on the trained output data. The second translational response data represents one or more translational responses across all translational axes of the translational model 220 for each driver model based on one or more predicted molecular sequences for that driver model. As an illustrative example, the translational model 220 can receive trained output data from a first driver model 210a representing one or more molecular sequences. The translational model 220 can predict ferret HAI titers, ferret AF titers, mouse HAI titers, mouse AF titers, and human replica AF titers for one or more molecular sequences predicted by the first driver model 210a across all strains. The second translational response data for each driver model in the selected set of driver models 240 is received by the feedback selection module 230. The feedback selection module 230 can compare the performance of each driver model along each translational axis and select the driver model that performs best along each axis or combination of axes to generate a selected subset of driver models 250. Using the previous example for illustration, with respect to the ferret HAI axis 220a, the feedback selection module 230 can compare the aggregate HAI titers across all pathogen strains circulating between January 1, 2019 and May 31, 2019 for one or more molecular sequences predicted by each of the driver models in the selected set of driver models 240. The feedback selection module 230 can then select the driver model that is found to have the highest aggregate HAI titer across all pathogen strains. In the illustrated embodiment, the selected subset of driver models 250 includes the second driver model 210b and the tenth driver model 210j.One or more of the selected subset 250 of driver models are included in the machine learning model 120 discussed above with reference to FIG.

[0056] Each of the driver models in the selected subset of driver models 250 is then validated based on observations from real-world experiments. For example, the second translational response data corresponding to the second driver model 210b is compared to biological responses observed in a human HAI experiment (or a ferret HAI experiment, a mouse HAI experiment, etc.) in which human subjects are vaccinated with one or more molecular sequences predicted by the second driver 210b and exposed to one or more of the pathogen strains circulating between October 1, 2017 and April 31, 2018. The predicted and observed responses are compared by feedback selection module 230 to generate an error value, which can determine based on the error value whether one or more of the translational axes corresponding to second driver model 210b (e.g., ferret HAI translational axis 220a, if second driver model 210b was selected based on its performance on ferret HAI translational axis 220a) are good or poor predictors of human response. If the error value meets an error threshold, one or more molecular sequences predicted by second driver model 210b can be used to design a vaccine for at least the disease season between October 1, 2017 and April 31, 2018, or even disease seasons following that disease season. For example, if a real-world ferret HAI experiment was used to validate second driver model 210b, the determined error value can be used to adjust parameters of translational model 220, second driver model 210b, or both.

[0057] Figure 3 shows a flowchart of a method 300 for designing a vaccine. For illustrative purposes, the method 300 is described as being performed by the architecture 200 previously described with reference to Figures 2A-2B. The method includes applying a plurality of driver models to a first time-series data set (block 310), training each driver model using the first time-series data set (block 320), selecting a set of trained driver models (block 330), applying a selected set of the trained driver models to a second time-series data set (block 340), and selecting a subset of the trained driver models (block 350).

[0058] At block 310, each driver model of the plurality of driver models 210 receives the first time series data set 201. Based on the received first time series data set 201, each driver model of the plurality of driver models 220 may generate output data representing one or more molecular sequences.

[0059] At block 320, for each of the driver models 210, the driver model is trained using the translational axes of the translational model 220 assigned to that driver model. Figure 4 shows a flowchart of a method 400 of training one or more driver models for vaccine design. Referring to Figure 4, the method 400 includes receiving output data from each driver model of the plurality of driver models 210 (block 410), applying the translational model 220 to the output data to generate first translational response data for each driver model of the plurality of driver models 210 according to the translational axes assigned to the driver model (block 420), adjusting, for each driver model of the plurality of driver models 210, one or more parameters of the driver model based on the first translational response data corresponding to the driver model (block 430), and repeating blocks 410-430 a number of iterations (block 440).

[0060] At block 330, a selected set of driver models 240 is generated for each translational axis of the translational model 220 based on the behavior of the driver models assigned to that translational axis. For example, after that number of iterations, the aggregate predicted ferret HAI titers of molecular sequences predicted by the first driver model 210a are compared to the aggregate predicted ferret HAI titers of molecular sequences predicted by the third driver model 210c, and the feedback selection module 230 can select the driver model corresponding to the highest aggregate predicted ferret HAI titer (or the greatest number of pathogenic strains effectively vaccinated).

[0061] At block 340, each of the selected set of driver models 240 receives the second time series data set 202 and generates learned output data representing one or more molecular sequences for a particular season of onset based on the second time series data set 202.

[0062] At block 350, the translational model 220 receives trained output data from each of the driver models in the selected set of driver models 240 and generates second translational response data for each driver model based on the trained output data. The second translational response data represents, for each driver model, one or more translational responses across all translational axes of the translational model 220 based on the one or more molecular sequences predicted by that driver model. As an illustrative example, the translational model 220 can receive trained output data from the first driver model 210a representing one or more molecular sequences. The translational model 220 can predict ferret HAI titer, ferret AF titer, mouse HAI titer, mouse AF titer, and human replica AF titer for the one or more molecular sequences predicted by the first driver model 210a. The second translational response data for each driver model in the selected set of driver models 240 is received by the feedback selection module 230. The feedback selection module 230 may compare the performance of each driver model for each translational axis and select the driver model with the best performance in each axis to generate a selected subset 250 of driver models.

[0063] Figure 5 shows a diagram depicting improvements per translational axis compared to traditional techniques for vaccine design. In an exemplary experiment, five different vaccine candidates were selected by a specific instance of the process described above (referred to by the abbreviations MO / 17, OS / 17, MI / 17, ET / 17, and KE / 17, which are homologous to strains A / MISSOURI / 37 / 2017, A / OSORNO / 60580 / 2017, A / MIYAZAKI / 89 / 2017, A / ETHIOPIA / 1877 / 2017, and A / KENYA / 105 / 2017, respectively) and then evaluated against the five different translational axes (shown crossing the x-axis) for the traditionally selected CV A / SINGAPORE / INFIMH160019 / 2016. Each of the five distinct CVs selected by the systems and methods described herein is displayed as a labeled marker on each translational axis, slightly offset within each translational axis for ease of visualization. The Y-axis indicates, for each translational axis, the proportion of March 2018 clinical isolates (later referred to as "seasonal surrogates") reported in the Global Initiative on Sharing All Influenza Data (GISAID) global database as of April 15, 2018. These clinical isolates were predicted to be better protected by a particular antigen than the conventionally selected CV (A / SINGAPORE / INFIMH160019 / 2016), which was the standard of care (SOC) for H3N2 as of March 2018. For example, the leftmost column (ferret HAI) shows that the translational model predicted that A / MISSOURI / 37 / 2017 would raise antibodies in ferrets with uniformly higher HAI titers than the conventionally selected CVs for all of these seasonal surrogate strains. As another example, in the right-most column (human serum antibody forensics (AF)), A / ETHIOPIA / 1877 / 2017 and A / OSORNO / 60580 / 2017 were predicted to be non-inferior to the traditionally selected CVs.Taken together, these results suggested that these five candidates should show diverse and distinct non-inferior patterns of induced immune responses when evaluated by different translational axes.

[0064] Exemplary translational models: 6 illustrates an example of a system 600 for predicting biological responses using machine learning techniques, according to one or more embodiments of the present disclosure. System 600 can be used as a translational model, as discussed above. System 600 includes a computer processor 610. Computer processor 610 includes computer-readable memory 611 and computer-readable instructions 612. System 600 also includes a machine learning system 650. Machine learning system 650 includes a machine learning model 620. Machine learning system 650 can be separate from or integrated with computer processor 610.

[0065] The computer-readable memory 611 (or computer-readable medium) may include any data storage technology type suitable for the local technology environment, including, but not limited to, semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, removable memory, disk memory, flash memory, dynamic random access memory (DRAM), static random access memory (SRAM), electronically erasable programmable read-only memory (EEPROM), etc. In some implementations, the computer-readable memory 611 includes code segments having executable instructions.

[0066] In some embodiments, the computer processor 610 includes a general-purpose processor. In some embodiments, the computer processor 610 includes a central processing unit (CPU). In some embodiments, the computer processor 610 includes at least one application-specific integrated circuit (ASIC). The computer processor 610 may also include a general-purpose programmable microprocessor, a special-purpose programmable microprocessor, a digital signal processor (DSP), a programmable logic array (PLA), a field programmable gate array (FPGA), special-purpose electronic circuitry, or the like, or a combination thereof. The computer processor 610 is configured to execute program code means, such as computer-executable instructions 612. In some embodiments, the computer processor 610 is configured to execute a machine learning model 620.

[0067] The computer processor 610 is configured to acquire first molecular sequence data 661 for the first molecular sequence and second molecular sequence data 662 for the second molecular sequence. The first molecular sequence data 661 may include amino acid sequence data for a candidate antigen (e.g., an inoculum strain). The candidate antigen may correspond to, for example, an H3N1 virus. The second molecular sequence data 662 may include amino acid sequence data for a known viral strain against which protection is sought. For example, the second molecular sequence may be a known viral strain that emerged in 2001. In some embodiments, as described in further detail below with reference to FIG. 9 , the computer processor 610 is also configured to receive non-human biological response data associated with the first and second molecular sequences. The non-human biological response data may include, for example, a biological response readout (e.g., antibody titer) that measures the biological response of a non-human model (e.g., a mouse, a ferret, a human immune system replica, etc.) to the second molecular sequence after inoculation with the first molecular sequence. As discussed in more detail below with reference to Figure 9, in some embodiments, the computer processor 610 can encode the first molecular sequence data 661 and the second molecular sequence data 662 as amino acid mismatches. Such data is obtained via one or more means, such as wired or wireless communication with a database (including in a cloud-based environment), fiber optic communication, a universal serial bus (USB), a read-only memory (CD-ROM), etc.

[0068] In machine learning system 650, machine learning techniques are applied to train machine learning models 620, which, when applied to input data, generate indications of whether or not an input data item has a relevant property, such as the probability that the input data item has a particular Boolean property, or an estimate of a scalar property.

[0069] As part of training the machine learning model 620, the machine learning system 650 can form a training set of input data by identifying a positive training set of input data items determined to have the property of interest, and in some embodiments, a negative training set of input data items that lack the property of interest.

[0070] The machine learning system 650 extracts configuration values ​​from the input data of the training set, where these configurations are variables that are deemed potentially relevant to whether the input data item has the associated property. The ordered list of configurations of the input data is referred to herein as a configuration vector of the input data. In some implementations, the machine learning system 650 applies dimensionality reduction (e.g., by linear discriminant analysis (LDA), principal component analysis (PCA), learned deep configurations from neural networks, etc.) to reduce the amount of data in the configuration vector of the input data to a smaller, more representative set of data.

[0071] In some embodiments, the machine learning system 650 uses supervised machine learning to train the machine learning model 620, with the constituent vectors from the positive and negative training sets as input. Various machine learning techniques are used in some embodiments, such as linear support vector machines (linear SVMs), boosting other algorithms (e.g., AdaBoost), neural networks, logistic regression, naive Bayes, memory-based learning, random forests, bagged trees, decision trees, boosted trees, or boosted stumps. When applied to the constituent vectors extracted from the input data items, the machine learning model 620 outputs an indication of whether the input data items have a property of interest, such as a Boolean yes / no estimate, a scalar value representing a probability, a vector of scalar values ​​representing multiple properties, or a nonparametric distribution of scalar values ​​representing a discrete and non-empirical fixed number of properties, where the indication is expressed explicitly or implicitly in a Hilbert space or similar infinite-dimensional space.

[0072] In some embodiments, the validation set is formed from additional input data other than the input data in the training set that have already been determined to have or lack the property of interest. The machine learning system 650 applies the trained machine learning model 620 to the data in the validation set to quantify the accuracy of the machine learning model 620. Common metrics applied to measure accuracy include: precision = TP / (TP + FP) and recall = TP / (TP + FN), where precision is the number of correct predictions (TP, i.e., true positives) made by the machine learning model 620 out of the total number of predictions made by the machine learning model 620 (TP + FP, i.e., false positives), and recall is the number of correct predictions (TP, i.e., false negatives) made by the machine learning model 620 out of the total number of input data items that had the property of interest (TP + FN, i.e., false negatives). The F-score (F-score = 2 × PR / (P + R)) unifies precision and recall into a single evaluation metric. In some implementations, the machine learning system 650 iteratively retrains the machine learning model 620 until a stopping condition occurs, such as an accuracy measurement indication that the model 620 is sufficiently accurate or after a number of training rounds have been performed.

[0073] In some implementations, the machine learning model 620 includes a neural network. In some implementations, the neural network includes a convolutional neural network. The machine learning model 620 may include other types of neural networks, such as a recurrent neural network, a radial basis function neural network, a physical neural network (e.g., an optical neural network), etc. Specific methods for training a machine learning model according to one or more implementations of the present disclosure are discussed in more detail below with reference to FIGS. 8-9.

[0074] The machine learning model 620 is configured to predict a biological response 663 to a second molecular sequence based on the received data. For example, assume that the first molecular sequence data 661 represents the amino acid sequence of a candidate antigen to be used as a vaccination, and the second molecular sequence data 662 represents the amino acid sequence of a viral strain known to have been circulating in 2012. The machine learning model 620 can predict the biological response (e.g., antibody titer) that the human immune system would produce after encountering the second molecular sequence (e.g., a known viral strain) if the human immune system had been inoculated with the first molecular sequence (i.e., the candidate antigen).

[0075] 7 is a flowchart illustrating an example of a method 700 for predicting a biological response using machine learning techniques, according to one or more embodiments of the present disclosure. For illustrative purposes, method 700 is described as being performed by system 600 for predicting a biological response using machine learning techniques, previously discussed with reference to FIG. 6. Method 700 includes receiving first sequence data for a first molecular sequence (block 710), receiving second sequence data for a second molecular sequence (block 720), and predicting a biological response to the second molecular sequence (block 730).

[0076] At block 710, the computer processor 710 receives first molecular sequence data 161 for a first molecular sequence. As previously indicated, the first molecular sequence data 161 may include amino acid sequence data for a candidate antigen (e.g., an inoculum strain). For example, the candidate antigen may correspond to an H3N1 virus.

[0077] At block 720, the computer processor 720 receives second molecular sequence data 662 for a second molecular sequence. The second molecular sequence data 662 may include amino acid sequence data of a known virus strain against which protection is sought. For example, the second molecular sequence may be a known virus strain that emerged in 2001.

[0078] In some embodiments, the method 700 further includes coding the first molecular sequence data 661 and the second molecular sequence data 662 as amino acid mismatches. For example, regions of similarity between the first and second molecular sequences are compared, and a value of "1" is coded for each non-matching pair of amino acids within the region, and a value of "0" is coded for each matching pair of amino acids within the region. In this way, a dissimilarity between the first and second molecular sequences, defined by non-matching amino acids at positions within the regions of similarity between the molecular sequences, is provided to the machine learning model 620.

[0079] In some embodiments, method 700 further includes receiving non-human biological response data associated with the first and second molecular sequences. The non-human biological response data may include, for example, a biological response readout (e.g., antibody titer) that measures the biological response of a non-human model (e.g., a mouse, a ferret, a replica human immune system, etc.) to the second molecular sequence after being inoculated with the first molecular sequence.

[0080] At block 730, the machine learning model 620 predicts a biological response to the second molecular sequence based on the received data. For example, the machine learning model 620 can predict a biological response (e.g., antibody titer) that a human immune system would produce after encountering a second molecular sequence (i.e., a known virus strain) if the human immune system had been inoculated with the first molecular sequence (i.e., the candidate antigen). In some embodiments, the machine learning model 620 is configured to predict a non-human biological response to the second molecular sequence. For example, the machine learning model can predict an antibody titer that an animal's immune system (e.g., a mouse, a ferret, etc.) would produce after encountering the second molecular sequence if the animal's immune system had been inoculated with the first molecular sequence.

[0081] How to train a machine learning model to predict biological responses: A method for training a machine learning model 620 for predicting a biological response will now be described. FIG. 8 illustrates an example of data used to train a machine learning model for predicting a biological response according to one or more embodiments of the present disclosure. As illustrated, data from thousands (or millions, billions, etc.) of experiments is used to build a comprehensive repository of biological response readout data and viral sequence data, for example, from ferret, mouse, and in vitro human immune system replica (e.g., MIMIC®) models. In the illustrated embodiment, the data includes antigen sequence data, viral sequence data, and biological response readouts evaluated by hemagglutination inhibition assay (HAI) and antibody forensics (AF). The viral sequence data includes a panel of known viral strains (referred to as the "readout" panel). The experiments are divided into batches called "cycles" (e.g., cycle 1 and cycle 2). In each cycle, the model system is challenged with selected molecular sequences (e.g., H3 proteins, vaccine formulations, etc.) and evaluated for their ability to generate an immune response against the panel of "readout" viral strains (referred to as the "readout panel"). The viral readout panel is selected to represent a broad sampling of influenza strains that have circulated during a defined time period (e.g., 1950-2016).

[0082] To correlate model experiments with human results, human sera are compared to a "readout" panel. In the illustrated example, not all antigen-strain / readout-strain pairs tested in the model system necessarily have corresponding pairs in human serum measurements. This is because human samples may be collected from vaccinated individuals during a period that does not cover the entire year used for each cycle. Therefore, the machine learning model is limited to only the antigens and readouts tested in human sera, and a vector of human readout titers is selected as the target vector for the machine learning model. The human AF readout can be from human sera collected 21 days after vaccination, a time period that is usually sufficient for subjects to seroconvert after vaccination.

[0083] Using data obtained from the foregoing experiments, a model is trained to predict biological responses. In some embodiments, a linear model is used.

[0084] FIG. 9 is a flow diagram of an example of training a machine learning model for predicting biological responses, according to one or more embodiments of the present disclosure. As shown, a data matrix 900 is first prepared, with each row corresponding to a pair of viral antigens, such as the H3 region of an antigen strain and a "readout" strain. The columns (or structure) of the matrix include dedicated columns for ferret model AF readout titers 902 and mouse model AF readout titers 903. In some embodiments, the missing titer data is entered as the average value of the column. However, any number of standard methods can be used to enter the missing titer data. The sequence column 901 represents a representation of the amino acid sequence difference (SeqDiff) between the antigen strain and the "readout" strain within a selected region, which in the illustrated example includes the H3 regions of the antigen strain and the "readout" strain. The SeqDiff is generated by examining, at each position in the H3 amino acid sequence alignment, whether that amino acid is the same or different between the antigen strain and the "readout" strain. If the amino acid between the two strains is not the same, a "1" is coded. If an amino acid between the two strains is the same, a "0" is coded. Coding two sequences as amino acid mismatches essentially creates a protein Hamming distance metric, which generally reflects the number of positions where the corresponding amino acids differ. In some embodiments, columns that are consistently "0" across the entire training set are discarded. Columns 901, 902, 903 of each row are correlated with the corresponding human titer 904 using linear regression.

[0085] Columns 902, 903 containing readout titers are, for example, z-score transformed before fitting a linear regression model. The z-score can represent linearly transformed data values ​​with a mean of zero and a standard deviation of one, indicating how many standard deviations above or below the mean an observation is. Because the coding of the SeqDiff representation can be sparse, principal component analysis (PCA) is sometimes used to reduce the dimensionality of the SeqDiff vector to five components. PCA refers to a statistical procedure that uses an orthogonal transformation to convert a set of observed values ​​of potentially correlated variables into a set of values ​​of linearly uncorrelated variables, called principal components. PCA is used to highlight variability, highlight strong patterns in a dataset, and reduce large sets of variables to smaller sets without losing significant information from the larger sets. Linear models are trained on various combinations of the data to better understand the relative ability of mouse and ferret titers and sequence data to predict human response.

[0086] As previously described, machine learning models are constructed as linear models for predicting biological responses, but nonlinear relationships between data structures and human biological responses may exist. Accordingly, using data from the aforementioned experiments, models using deep neural networks or other nonlinear models are constructed that: 1) exploit nonlinear relationships in the data to make relatively accurate predictions compared to the aforementioned linear models; and 2) can simultaneously predict both animal and human titers. Predicting all titers together takes advantage of the recognition that strong signals of immune response are directly encoded in the protein sequences of antigen and "readout" strains. By training a model to predict both human and animal titers from sequence alone, the machine learning model is forced to explore sequence-function relationships that drive immunogenicity across species. In statistical terms, this is called "borrowing strength," and it allows the model to better leverage the large amount of data available in certain models (e.g., the ferret model) to generate more robust predictions of human responses. This approach can accommodate a larger number of viral antigens and the construction of a data matrix with over 13,000 example rows. Similar to the linear model, the SeqDiff representation of the H3 region for each virus strain and readout pair is used as input data.

[0087] In some embodiments, the target vector is a linear model of human titers, while a nonlinear neural network model can represent a multi-target regression problem for, for example, seven output columns (ferret HAI and AF titers, mouse HAI and AF titers, MIMIC AF, human HAI, and human AF). The detection limit for HAI experiments can typically be 40 (or 1:40 in dilution), so all measurements below this value are set to 40. Similarly, if an AF measurement is below 10,000, it is set to 10,000. HAI is expressed as log2(titer / 10), and AF is expressed as log2(titer). Human data and human replica data can have an additional level of complexity when measurements are taken at the time of inoculation (Day 0) and after seroconversion (Day 21). Thus, human and human replica titers are expressed as the log2-fold change of Day 21 / Day 0. If a titer value is missing from the target vector, the value is set to zero, and the neural network loss function is masked at that position. This ensures that predictions for missing values ​​do not contribute to the fitness of the model during training.

[0088] In some implementations, a neural network with two 128-node dense layers with relu activation and a 7-node dense output layer is used. A portion of the data (e.g., 15 percent of the data) is randomly set aside as a test set, and the neural network is trained for a number of epochs (e.g., 400, 500, 1000, etc.). In some implementations, the following parameters are used: learning rate = 0.001; weight decay = 0.0001; batch size = 128.

[0089] In some embodiments, an L2 loss function is used for the human replica, human AF, and human HAI target vectors. Generally, the L2 loss function minimizes the squared difference between the estimated target value and the existing target value. In some embodiments, the Huber loss function is used for the ferret data and the mouse data. Generally, the Huber loss function is used in robust regression and, at least in some cases, may be less sensitive to data outliers than the L2 loss function. To further bias the model, an explicit weighting approach is used to apply an additional penalty to misclassified human samples. For example, each target loss at each epoch of training is multiplied by the following weights: ferret HAI = 0.8; ferret AF = 1; mouse HAI = 1; mouse AF = 1; human HAI = 2; human AF = 2; MIMIC = 1.5.

[0090] Although the above description occasionally describes pathogenic strains in the context of influenza strains for illustrative purposes, the term pathogen is broadly construed to encompass any infectious agent. For example, a pathogenic strain may refer to, among other things, a viral strain, a bacterial strain, a protozoan strain, a prion strain, a viroid strain, or a fungal strain. A pathogenic strain may correspond to respiratory syncytial virus and other paramyxoviruses. A pathogenic strain may correspond to, among other things, pertussis, diphtheria, or tetanus.

[0091] While the above description sometimes describes a disease season in relation to influenza season, the term disease season is broadly interpreted to encompass any discrete time interval. For example, a disease season may refer to, among other things, a particular month, a particular week, a particular set of weeks, a particular set of months, or a particular set of days. Furthermore, successive disease seasons may be constant or variable. For example, two successive disease seasons may both be one month in length, or one disease season may be one month in length and the second disease season may be four days in length.

[0092] While the above description describes specific translational axes / biological responses, such as ferret HAI titers and mouse AF titers, embodiments are not so limited. For example, one biological response / translational axis can correspond to antibody characterization, such as affinity and / or avidity for a panel of specific antigens and / or antigen fragments (e.g., protein arrays, phage display libraries, etc.), functional profiling, such as to determine anti-drug antibodies, immune-complement interactions (e.g., phagocytosis, inflammation, membrane attack), antibody-dependent cellular cytotoxicity (ADCC) or similar Fc-mediated effector functions, profiling of formed immune complexes (e.g., receptor binding profiles), immunoprecipitation assays, or a combination thereof. One biological response / translational axis can correspond to antibody competition, where one target is bound to another antibody or antiserum. One biological response / translational axis can correspond to that of antibody characterization mentioned above as well as antiserum characterization, which can accommodate functional assays (such as microneutralization assay, hemagglutination inhibition, and neuraminidase inhibition), binding assays (such as hemagglutination assay), enzymatic reaction assays (such as enzyme-linked lectin assay (ELLA)), ligand binding assays (such as binding of sialic acid derivatives and their mimetics), and fluorescent readout assays (such as 20-(4-methylumbelliferyl)-aDN-acetylneuraminic acid (MUNANA) cleavage).

[0093] One biological response / translational axis can correspond to in vivo evaluation utilizing either monoclonal or polyclonal antibodies by passive transfer and / or exogenous expression or transfer achieved by one or more of the following: transfection or endogenous expression via retroviral infection, host genome modification such as with CRISPR, bibody-to-body fluid transfer, or a combination thereof. One biological response / translational axis can correspond to in vivo evaluation of immunity raised by immunization to assess antigenicity. One biological response / translational axis can correspond to characterization such as binding / affinity measurements of linear peptide antigens on major histocompatibility complex (MHC) class I and class II, and even to assess productive T cell epitope display for recognition by T cells. One biological response / translational axis can correspond to characterization such as affinity to panels of antigen fragments (e.g., protein arrays, phage display libraries, etc.) to identify recognized epitopes. One biological response / translational axis can correspond to ex vivo and / or in vitro functional profiling, such as to determine T cell responses and / or mediated responses. One biological response / translational axis can correspond to in vivo and / or in situ measurement of proliferation (e.g., abundance in a tissue compartment) of T cells associated with adaptive responses (e.g., αβ or γδ T cells) in response to natural infection and / or challenge and / or immunization. One biological response / translational axis can correspond to in vitro and / or ex vivo measurement of the specificity of recognition by T cells associated with adaptive responses (e.g., αβ or γδ T cells) in response to natural infection and / or challenge and / or immunization as measured by competition with other epitopes.

[0094] One biological response / translational axis can correspond to in situ, ex vivo, and / or in vivo assessment of morphological or physiological changes in response to the pathogen to be protected against or to tissue formation, tissue repair, or tissue invasion by a surrogate such as a pseudotyped virus or bacterium. One biological response / translational axis can correspond to differences in in situ, ex vivo protein, gene expression, and / or non-coding RNA levels in response to other antigens and / or physiological states characterized by biomarkers such as age, sex, frailty, nominal serostatus, race, haplotype, geographic location, etc. One biological response / translational axis can correspond to in situ assessment of defense, infection, or other overall physiological responses to naturally occurring or transmitted infection in humans or model organisms such as, but not limited to, mice, rats, rabbits, ferrets, guinea pigs, pigs, cows, chickens, sheep, dolphins, bats, dogs, cats, zebrafish and other bony fish, and non-human primates such as monkeys and apes.

[0095] With respect to responses to intentional infection (i.e., challenge) with homotypic and / or heterotypic infectious agents, including controlled human challenge studies, one biological response / translational axis can correspond to in situ, ex vivo, and / or in vivo assessment of proteins or metabolites present in blood or tissues, where the proteins can be cytokines, hormones, or signaling molecules, and the metabolites can be vitamins, cofactors, or other metabolic by-products. One biological response / translational axis can correspond to in situ, ex vivo, and / or in vivo assessment of the microbiome, which can be affected by or influence the immune response. One biological response / translational axis can correspond to functional profiling ex vivo, in vitro phenotypic, and / or functional T cell response profiling (receptor expression, cytokine production, cytotoxicity) in response to challenge with antigen alone or in combination with innate immune cells (natural killer (NK) cells, dendritic cells (DCs), neutrophils, macrophages, monocytes, etc.). One biological response / translational axis can correspond to epigenetic analysis performed using samples collected or generated using the techniques or methods described above.

[0096] Although the above description describes several methods and data for training machine learning models to predict biological responses, other methods and data are also used. For example, a neural network model may include more or fewer layers than the models described above, and each layer may have more or fewer nodes.

[0097] In the foregoing description, embodiments have been described with reference to numerous specific details that may vary from embodiment to embodiment. Accordingly, the specification and drawings should be regarded in an illustrative, rather than a limiting sense. The sole and exclusive indicator of the scope of the present disclosure and what applicants deem to be the scope of the present disclosure is the set of claims originating from this application, in the specific form from such claims, including any subsequent amendments, and their literal and equivalent scope. Any definitions expressly set forth herein for terms contained in such claims shall govern the meaning of such terms as used in the claims. Additionally, when the term "further comprising" is used in the foregoing description or the appended claims, what follows this phrase may be additional steps or entities, or substeps / subentities of a previously recited step or entity.

Claims

1. A method implemented by one or more computers, comprising: obtaining, for each of a plurality of molecular sequences, a time sequence data set that includes data defining (i) the molecular sequence and (ii) one or more circulating periods of pathogenic strains that include the molecular sequence as a native antigen; Using machine learning training techniques, in each of a plurality of training iterations, operations including: i) processing the time sequence dataset through a driver neural network to generate output data representing one or more candidate molecular sequences; ii) for each candidate molecular sequence generated by the driver neural network in the training iterations, processing the data defining the candidate molecular sequence with a translational neural network to generate a biological response defining a predicted biological response to the pathogen of an entity vaccinated with the candidate molecular sequence; and iii) Machine learning training techniques are used to train one or more of the driver neural networks. adjusting the parameters to optimize a loss function that depends on the biological response of each candidate molecular sequence generated by the driver neural network in the training iterations; training the driver neural network by executing: outputting the trained driver neural network; The method comprising:

2. The method of claim 1 , wherein the driver neural network comprises a recurrent neural network.

3. The method of claim 1 , wherein the driver neural network comprises a long-short-term memory recurrent neural network.

4. The method of claim 1 , wherein the output data representing the one or more candidate molecular sequences includes one or more candidate molecular sequences for each of a plurality of seasons of onset.

5. For each of the multiple onset seasons, the candidate molecular sequences corresponding to the onset season are 5. The method of claim 4, predicted to achieve a maximized aggregate biological response across all circulating pathogen strains.

6. 5. The method of claim 4, wherein for each of a plurality of disease seasons, a candidate molecular sequence corresponding to the disease season is predicted to generate a biological response that effectively immunizes against a maximum number of viruses circulating in the disease season.

7. 10. The method of claim 1, wherein the entity is: a ferret, a mouse, a human replica, or a human.

8. The method of claim 1 , wherein training of the driver neural network is performed for a predetermined number of training iterations.

9. The method of claim 1 , wherein training of the driver neural network is performed until a termination criterion based on a predetermined error value is met.

10. 10. The method of claim 1, wherein for each candidate molecular sequence, the predicted biological response of an entity vaccinated with the candidate molecular sequence is characterized as a predicted biological response measured by a hemagglutination inhibition assay.

11. 10. The method of claim 1, wherein for each candidate molecular sequence, a predicted biological response of an entity vaccinated with the candidate molecular sequence is characterized as a predicted biological response measured by enzyme-linked immunosorbent assay.

12. The method of claim 1 , wherein each candidate molecule sequence defines a respective antigen.

13. 10. The method of claim 1, wherein the translational neural network generates, for each candidate molecular sequence, a respective biological response to each of multiple strains of the pathogen.

14. For each candidate molecular sequence: generating an aggregate biological response by aggregating the respective biological responses for each of the plurality of strains of the pathogen; Here, the loss function depends on the respective aggregate biological responses for each candidate molecular sequence. To do, 14. The method of claim 13, further comprising:

15. Generating an aggregate biological response for each candidate molecular sequence comprises the steps of: determining an average biological response; The method of claim 14, comprising:

16. For each candidate molecule sequence: determining, based on the biological responses for the plurality of strains of the pathogen, the number of strains of the pathogen for which vaccinating the entity with the candidate molecular sequence achieves at least a threshold biological response in the entity; where the loss function depends, for each candidate molecular sequence, on the number of strains of the pathogen for which at least a threshold biological response is achieved in the entity by vaccinating the entity with the candidate molecular sequence.

14. The method of claim 13, further comprising:

17. Training the driver neural network 10. The method of claim 1, comprising training an ensemble of The method further comprises selecting, from the ensemble of driver neural networks, an appropriate subset of the ensemble of driver neural networks having the highest performance measure; and Selecting molecular sequences for a vaccine against a pathogen using a suitable subset of the ensemble of driver neural networks. The method of claim 1 , comprising:

18. 10. The method of claim 1, Further, after training the driver neural network, using the trained driver neural network to select a molecular sequence for a vaccine against a pathogen; evaluating the performance of each of the trained driver neural networks across each of a plurality of translational axes, each of the plurality of translational axes corresponding to (i) a respective species of the entity to be vaccinated, or (ii) a respective measure of the biological response of the entity to be vaccinated, or (iii) both; Selecting a trained driver neural network from among the ensemble of trained driver neural networks based at least in part on the performance of each of the trained driver neural networks across each of a plurality of translational axes. The method comprising:

19. 20. The method of claim 18, further comprising: experimentally evaluating the performance of a vaccine containing the selected molecular sequence; The method.

20. For each candidate molecular sequence, processing the data defining the candidate molecular sequence with a translational neural network includes: jointly processing data defining (i) candidate molecular sequences and (ii) molecular sequences of strains of pathogens using a translational neural network; The method of claim 1 , comprising:

21. 21. The method of claim 20, the candidate molecule sequence comprises a candidate molecule amino acid sequence; the molecular sequence of the pathogen strain comprises the amino acid sequence of the pathogen strain; and the data defining (i) the candidate molecular sequence and (ii) the molecular sequence of the pathogen strain includes data identifying amino acid mismatches between corresponding positions of the candidate molecular amino acid sequence and the pathogen strain amino acid sequence; The method.

22. 10. The method of claim 1, wherein the translational neural network is parametrized by a set of translational neural network parameters having respective values ​​determined by machine learning training techniques.

23. The method of claim 1 , wherein the translational neural network comprises a recurrent neural network.

24. 10. The method of claim 1, further comprising: generating a vaccine comprising the selected molecular sequence; The method comprising:

25. The system includes: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, the one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform operations, such as: obtaining, for each of a plurality of molecular sequences, a time-sequence data set comprising data defining (i) the molecular sequence and (ii) one or more circulation times of pathogenic strains that contain the molecular sequence as a native antigen; Using machine learning training techniques, in each of a plurality of training iterations, operations including: i) processing the time sequence dataset through a driver neural network to generate output data representing one or more candidate molecular sequences; ii) each candidate generated by the driver neural network in a training iteration For the molecular sequences, processing the data defining the candidate molecular sequences through a translational neural network to generate a biological response defining a predicted biological response to the pathogen of an entity vaccinated with the candidate molecular sequence; and iii) Machine learning training techniques are used to train one or more of the driver neural networks. More than adjusting the parameters to optimize a loss function that depends on the biological response of each candidate molecular sequence generated by the driver neural network in the training iterations; training the driver neural network by executing outputting the trained driver neural network; the storage device, The system comprising:

26. 26. The system of claim 25, wherein the driver neural network comprises a recurrent neural network.

27. 26. The system of claim 25, wherein the driver neural network comprises a long short-term memory recurrent neural network.

28. One or more non-transitory computer storage media having stored thereon instructions for causing one or more computers to perform one or more computer-executed operations, the operations including: obtaining, for each of a plurality of molecular sequences, a time-sequence data set comprising data defining (i) the molecular sequence and (ii) one or more circulation times of pathogenic strains that contain the molecular sequence as a native antigen; Using machine learning training techniques, in each of a plurality of training iterations, operations including: i) processing the time sequence dataset through a driver neural network to generate output data representing one or more candidate molecular sequences; ii) Each candidate segment generated by the driver neural network in a training iteration For the child sequences, processing the data defining the candidate molecular sequences through a translational neural network to generate a biological response defining a predicted biological response to the pathogen of an entity vaccinated with the candidate molecular sequence; and iii) tuning one or more parameters of a driver neural network in a machine learning training technique to optimize a loss function that depends on the biological response of each candidate molecular sequence generated by the driver neural network in the training iterations; training the driver neural network by executing After training the driver neural network, using the trained driver neural network to select a molecular sequence for a vaccine against a pathogen; The non-transitory computer storage medium.

29. 30. The non-transitory computer storage medium of claim 28, wherein the driver neural network comprises a recurrent neural network.

30. 30. The non-transitory computer storage medium of claim 28, wherein the driver neural network comprises a long short-term memory recurrent neural network.

Citation Information

Patent Citations

  • Method for predicting change of influenza virus antigen

    CN109448781A