Systems and methods for compound classification using gated transformer networks
Patent Information
- Application Number
- EP2024804673
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-30
- Filing Date
- 2024-10-24
- Publication Date
- 2026-09-09
AI Technical Summary
Existing methods for evaluating compounds in animal trials face challenges in efficiently analyzing and predicting the administration effects of compounds on humans due to limitations in processing and memory resources, especially when dealing with large volumes of multivariate time series data.
The use of gated transformer networks to process and analyze multivariate time series data from animal trials, allowing for the generation of training samples and the training of machine learning models to predict the administration effects of compounds on humans.
This approach enables more efficient and accurate prediction of compound administration effects, supporting drug development by providing early indications of potentially therapeutic compounds and overcoming the limitations of existing methods.
Smart Images

Figure IMGF000032_0001 
Figure 00000057_0000 
Figure 00000058_0000
Abstract
Description
SYSTEMS AND METHODS FOR COMPOUND CLASSIFICATION USING GATEDTRANSFORMER NETWORKSCROSS-REFERENCE TO RELATED APPLICATION
[0001] The present application claims priority to U.S. Provisional App. No. 63 / 594,282, filed on October 30, 2023, which is hereby incorporated by reference in its entirety.TECHNICAL FIELD
[0002] The present disclosure relates to devices, systems, and methods for capturing, analyzing, and augmenting behavioral and physiological data of subjects. In particular, the subjects can be non-human animals and the behavioral and physiological data can be captured during a trial in which a compound is administered.BACKGROUND
[0003] The evaluation of compounds (e.g., as part of a drug development program) can involve non-human animal trials. Such animal trials can provide the safety and efficacy data necessary to support subsequent human trials.
[0004] Behavioral assessments can be an important component of such trials. The effects of experimental substances on animal behavior can provide information about the potential clinical effects of the experimental substances on human subjects. For example, the discovery that chlorpromazine produces differential effects on avoidance and escape behavior in animals encouraged the evaluation of the behavioral effects of other experimental antipsychotic drugs.
[0005] A behavioral assessment can include comparing the behavior of an animal treated with a compound to the normal behavior of the animal (or another animal). A behavioral phenotype can be generated from a set of monitored behaviors. The behavioral phenotype can correlate the administration of the compound with the behavior or physiology of the animal during the trial. Behavioral platforms can automatically collect observational data for the generation of such behavioral phenotypes. As compared to manual collection of observational data, automatic data collection can be more complete, repeatable, reliable, and accurate.
[0006] Observational data may include long, multi-variate time series data, which can involve many data points over many recording channels. Analysis of observational data may be limited by memory and processing resources, such as memory capabilities in a GPU.Some observational data may include sparse data or data in which meaningful events are not distributed evenly over time. As such, analysis of observational data may be time consuming, use large amounts of memory, and inefficient for machine learning purposes, and may therefore benefit from data augmentation.SUMMARY
[0007] Certain embodiments of the present disclosure relate to systems and method for predicting the administration effect of a compound in humans using observational data acquired from non-human animals.
[0008] The disclosed embodiments include a training method. The method can include obtaining multivariate time series data including multiple channels of univariate time series animal behavior data, the multiple channels of univariate time series animal behavior data acquired from a non-human animal during a trial in which the non-human animal was administered a compound. The method can include generating training samples. Generation of training samples can include generating blocks. Generation of a first one of the blocksincluding determining a block duration and starting timepoint and storing in the first one of the blocks a portion of the multivariate time series data beginning at the starting timepoint and having the block duration. Generation of training samples can further include associating the blocks with a class label. The method can further include training a machine learning model to generate an indication of an administration effect using a training dataset including the training samples.
[0009] The disclosed embodiments include a method for predicting compound class labels using animal behavior data. The method can include obtaining multivariate time series data including multiple channels of univariate time series animal behavior data, the multiple channels of univariate time series animal behavior data acquired from a non-human animal during a trial in which the non-human animal was administered a compound. The method can further include generating a sequence of blocks of the multivariate time series data, a first block of the sequence of blocks including a first portion of the multivariate time series data. The method can further include generating a class label by applying the sequence of blocks to a machine learning model, the class label indicating an effect of the compound when administered to humans. The machine learning model can be configured to generate a first encoding by applying the first block to an encoder; generate a first input vector using a set of encodings corresponding to blocks in the sequence of blocks, the set of encodings including the first encoding; and generate the class label by applying the first input vector to a classifier.
[0010] The disclosed embodiments include a system. The system can include at least one non-transitory computer-readable medium storing instructions and at least one processor configured to execute the instructions to perform operations for predicting an administration effect of a compound when the compound is administered to humans using animal behavior data. The operations can include obtaining multivariate time series data including multiplechannels of univariate time series animal behavior data, the multiple channels of univariate time series animal behavior data acquired from a non-human animal during a trial in which the non-human animal was administered the compound. The operations can further include generating a sequence of blocks using the multivariate time series data, a first block of the sequence of blocks including a first portion of the multivariate time series data. The operations can further include generating an indication of the predicted administration effect by applying the sequence of blocks to a machine learning model. The machine learning model configured to: generate a first encoding by applying the first block to an encoder; generate a first input vector using a set of encodings corresponding to blocks in the sequence of blocks, the set of encodings including the first encoding; and generate the indication of the predicted administration effect by applying the first input vector to a classifier.
[0011] The foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the claims. Other systems, methods, and computer-readable media are also discussed within.BRIEF DESCRIPTION OF DRAWINGS
[0012] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate several embodiments and, together with the description, serve to explain the disclosed principles. In the drawings:
[0013] FIG. 1 depicts a system for acquiring (and optionally analyzing) observational data concerning a non-human animal, consistent with disclosed embodiments.
[0014] FIG. 2 depicts a graph 200 of multivariate time series data, consistent with disclosed embodiments.
[0015] FIG. 3 depicts a method for training and using a machine learning model to predict an administration effect of a compound on a human, consistent with disclosed embodiments.
[0016] FIG. 4 illustrates a diagram of an architecture 400 for predicting an administration effect of a compound on humans using observational data obtained from non-human animals, consistent with disclosed embodiments.
[0017] FIG. 5 depicts an exemplary architecture of a machine learning model, consistent with disclosed embodiments.
[0018] FIG. 6 depicts an exemplary system for generating machine learning models, consistent with disclosed embodiments.DETAILED DESCRIPTION
[0019] Reference will now be made in detail to exemplary embodiments, discussed with regard to the accompanying drawings. In some instances, the same reference numbers will be used throughout the drawings and the following description to refer to the same or like parts. Unless otherwise defined, technical or scientific terms have the meaning commonly understood by one of ordinary skill in the art. The disclosed embodiments are described in sufficient detail to enable those skilled in the art to practice the disclosed embodiments. It is to be understood that some embodiments may be utilized and that changes may be made without departing from the scope of the disclosed embodiments. Thus, the materials, methods, and examples are illustrative only and are not intended to be necessarily limiting.
[0020] Machine learning models consistent with disclosed embodiments can enable prediction of the effect of a compound in humans based on an analysis of observational data concerning a non-human animal administered the compound. The non-human animal can be a control or wild-type animal; an animal model of disease, dysfunction, or injury, or the like. For example, animal models of disease can include models of rare disorders, Parkinson’s disease, Huntington’s disease, Alzheimer’s Disease, or other diseases. As an additional example, animal models of dysfunction can include models of autism spectrum disorder,epilepsy, depression, anxiety, anhedonia, apathy, cognitive dysfunction, hyperreactivity, or other dysfunctions. As a further model, animal models of injury can include models of central, peripheral, enteric, somatic, or autonomic nervous system dysfunction or injury. Animal models may be obtained using genetic, environmental, pharmacological, or other manipulations.
[0021] Such predictions can support drug development by providing an early indication of potentially therapeutic compounds. The observational data can include subject data (e.g., behavioral data, physiological data, electrophysiological data, or the like) and / or external data (e.g., environmental data, indications of stimuli or rewards presented to the subject, or the like). In some instances, the non-human animal can be a rodent.
[0022] Consistent with disclosed embodiments, a testing system can include a controlled environment configured with sensors. The testing system can use the sensors to acquire observational data from a non-human animal administered a compound and disposed within the controlled environment. A machine learning model can predict a class of a compound based on the observational data. The compound class can correspond to an effect of the compound when administered to a human.
[0023] Consistent with disclosed embodiments, the sensors can include cameras, electrophysiological recording devices (e.g., electroencephalography (EEG) recording devices, or the like), piezoelectric sensors, infrared detectors, radiofrequency detectors, and the like. In some embodiments, a sensor can include both an emitter such as an infrared beam, a radiofrequency, a source of heat, a source of optical signals, or other such source, as well as a receiver, such as a component configured to receive data or electronic communications. Multiple cameras can be used to acquire depth or 3D information. Theparticular configuration of sensors can be selected or adapted based on the subject and / or the compound administered.
[0024] Consistent with disclosed embodiments, the observational data can include position data, motion data, thermal data, force data, EEG signals, respiration data, or the like. For example, observational data for a mouse can include head, body center, and paw positions (e.g., x, y, z coordinates or the like) and time derivatives thereof (e.g., velocity and acceleration); heart rate; EEG data; eye temperature; and / or other suitable measures. In some embodiments, observational data can be acquired and stored for subsequent processing.
[0025] Time series data, including multivariate and univariate time series data, may involve large numbers of data points. Time series data may be measured over various intervals of time and at varying sampling rates such that the length of time series may span hundreds to thousands, or more, of data points. In some embodiments, a multivariate time series dataset may include between 2 and 100, or more univariate time series datasets. These univariate time series datasets can include measurements. In some embodiments, each measurement can be associated with a measurement time. An association between a measurement and measurement time can be explicit (e.g., the measurement value and measurement time can both be included in the univariant time series dataset) or implicit (e.g., a sequence of measurements, a start or finish time, a sampling frequency or period, or the like can be included in the univariant time series dataset). In some embodiments, univariate time series included in a multivariate time series dataset can include different numbers of measurements, measurements acquired at different times, measurements acquired under different conditions, or any combination of the foregoing. For example, a first univariate time series can include measurements acquired at 10 Hz, while a second univariate time series can include measurements acquired when a condition is satisfied (e.g., heart rate measurements acquired at .1 Hz for 300 seconds following administration of an electric shock to the test subject). Insome embodiments, the number of measurements included in a univariate time series dataset can be between 100 and 100000 measurements, or more. In some embodiments, a univariate time series dataset can correspond to a measurement time of between be between 100 measurements and 20 million measurements, or more. For example, a channel of EEG data collected at a sampling rate of 1 kHz for 30 minutes can include nearly 2 million measurements. In some embodiments, time series data may include multiple channels and / or multiple drug classes, thereby compounding the number of data points. As such, it will be recognized that the amounts of data points may use large amounts of memory, such as memory on a graphics processing unit.
[0026] Disclosed embodiments may involve multiple channels of univariate time series animal behavior data acquired from a non-human animal during a trial in which the nonhuman animal was administered a compound. The univariate time series animal behavior data can be observational data. Administering a compound to a non-human animal may involve providing a compound to the animal, such as by injecting or adding to the food or water provided to the animal.
[0027] Disclosed embodiments may involve screening compounds using an animal model of disease, dysfunction, or injury. An animal model subject can be treated with a compound. A class for the animal model subject can be predicted. As may be appreciated, absent a therapeutic effect of the compound, the animal model subject may be identified as having the disease, dysfunction, or injury. When the compound has a therapeutic effect that addresses the disease, dysfunction, or injury, the animal model subject may instead be identified as being a vehicle, wild-type animal, or similar control subject. In some embodiments, a likelihood or probability of identifying the animal model subject as being such a control subject can be used to quantify the therapeutic effect of the compound.
[0028] Consistent with disclosed embodiments, an administration effect can be the predicted response (e.g., biological, physiological, or psychological response) of a human to the administration of the compound. The effect of administering the compound may or may not be therapeutic. In some embodiments, the administration effect of the compound can be described in terms of classes of drugs. In some embodiments, the administration effect of the compound can be described in terms of a drug class and subclass. In some embodiments, a drug class can represent a coarse categorization and a drug subclass can represent a finer categorization. In some embodiments, a drug subclass can be associated with a mechanism of action or a type of chemical structure. In some embodiments, an administrative effect can represent or reflect an overall activity, potency, or side effect(s) of the compound. In some embodiments, an administrative effect can represent or reflect a phenotype of an animal model of disease, dysfunction, or injury (e.g., following administration of a placebo, therapeutically ineffective compound, or the like to such an animal model). In some embodiments, an administrative effect can represent or reflect a normal behavioral phenotype (or a recovery of a normal phenotype in such an animal model following administration of a therapeutic compound). For example, recovery of a normal behavioral phenotype (e.g., indicated by a predicted class of “wild-type,” “vehicle,” “normal,” “control” or the like) in a non-human animal model of a disease, dysfunction, or injury, can indicate that a human having the disease, dysfunction, or injury would therapeutically benefit from administration of the compound.
[0029] The disclosed embodiments are not limited to any particular listing of drug classes. In some embodiments, such classes can include antidepressant, anxiolytic, antipsychotic, cognitive enhancer, hallucinogen, anticonvulsant, mood stabilizer, psychostimulant, or the like. In some embodiments, such classes can include therapeutic antidepressant, high dose side-effect antidepressant, hallucinogen / psychedelic / treatment-resistant depression treatment,therapeutic antipsychotic, high dose antipsychotic, anxiolytic, sedative, cognitive enhancer / Alzheimer's treatment, psychostimulant / ADHD, mood stabilizer, anticonvulsant, side effect / toxic, anticognitive, other therapeutic area, inactive, sedative, hypnotic, vehicle, anxiogenic, or analgesic. In some embodiments, such classes may include an “unknown” class. The unknown class can include compounds that do not fall into the other drug classes (e.g., active compounds having administration effects distinct from the administration effects associated with the other drug classes). In some embodiments, as described herein, a compound may be predicted to belong to (or be associated with) more than one drug class (or subclass).
[0030] The disclosed embodiments are not limited to any particular listing of drug subclasses. In some embodiments, antidepressants can be further sub-classified as NMDA antagonists, serotonin 2A antagonists and serotonin reuptake inhibitors, Monoamine oxidase A (MAO A) and Monoamine oxidase B (MAOB) inhibitors, norepinephrine and dopamine reuptake inhibitors, selective serotonin reuptake inhibitors, tricyclics, or the like. In some embodiments, analgesics can be further sub-classified as opiate agonists, non-steroidal antiinflammatories, or the like. In some embodiments, anxiolytics can be further sub-classified as glutamate receptor 5 agonists, serotonin 1 A receptor agonists, benzodiazepines, or the like. In some embodiments, sedatives (or hypnotics) could be further subclassified being z-drugs, or the like. In some embodiments, antipsychotics could be further subclassified as atypical antipsychotics, typical antipsychotics, or the like. In some embodiments, cognitive enhancers could be further subclassified as nicotinic agonists, NMDA antagonists, cholinesterase inhibitors, or the like.
[0031] FIG. 1 illustrates a system 110 for acquiring (and optionally analyzing) observational data concerning a non-human animal 115, consistent with disclosed embodiments. System 110 may include a control computer 50 communicating 135 with a computer-controlledenclosure 120. System 110 can include sensors configurable to acquire data during an experiment on a non-human animal 115. System 110 can include a computer-controlled enclosure 120 configured to contain the non-human animal 115 during the experiment. In some embodiments, the computer-controlled enclosure 120 can include actuators configurable to interact with non-human animal 115. For example, sensors and / or actuators may include an aversive stimulus probe 118 that may be deployed and retracted, a motor challenge 126 that may be deployed or retracted so as to force non-human animal 115 to walk on a plurality of physical obstacles 124 arranged in an array with a predefined pitch to provide a motor challenge, lighting 104 that may be configured to change illumination intensity levels and / or light wavelengths as applied to the non-human animal 115, a tactile stimulator 116 to administer tactile stimuli to the non-human animal 115, a top camera 106, a top thermal camera 102, a first side camera 128 a second side camera 114, a floor force sensor 122, waterers and feeders 129, and as additional actuators 22 for applying any additional suitable stimuli to the non-human animal 115 and / or any additional sensors.
[0032] In some embodiments, computer-controlled enclosure 120 can include social pod 112 (e.g., shown as an opening in enclosure 120 in FIG. 1 that enables social pod 112 to be affixed to enclosure 120 via the opening). Affixing social pod 112 to computer-controlled enclosure 120 can enable a second non-human animal in social pod 112 to interact with non- human animal 115.
[0033] In some embodiments, computer-controlled enclosure 120 can include a 3D camera. A 3D camera can be implemented using multiple 2D cameras configured and arranged around enclosure 120 to obtain 3D image data. In some embodiments, the 3D camera can be implemented using at least one of top camera 106, the first side camera 128, the second side camera 114, or additional cameras.
[0034] In some embodiments, the first side camera 128 and the second side camera 114 may be oriented to capture images in the computer-controlled enclosure 120 from different perspectives. An angle between the orientation of the first side camera 128 and the second side camera 114 may be any suitable angle to capture movement throughout the computer- controlled enclosure 120. For example, the angle may be 90 degrees, though other angles may be used, such as, e.g., any angle from about 1 degree to about 179 degrees, such that imagery from both the first side camera 128 and the second side camera 114 may be processed to determine movement within the computer-controlled enclosure 120.
[0035] In some embodiments, the plurality of sensors may include sensors associated with some actuators to capture specific responses to the actuators. In some embodiments, the plurality of sensors may be combined with multiple actuators to challenge the test subject to react to various events, which are recorded and analyzed by the system. The resulting ethophysiogram, or collection of physiological and behavioral responses, may create a dataset (e.g., a content-rich dataset) suitable for use with the disclosed systems and methods.
[0036] Consistent with disclosed embodiments, the control computer 150 may include at least one processor 160 for executing suitable computer applications as described herein for performing the analysis of the behavior of the non-human animal 115, at least one memory and / or suitable storage device denoted by memory 162 for storing the computer code and any databases used in the analyses of the acquired data over the predetermined time period, control circuitry 164 for controlling the plurality of actuators in accordance with the experimental plan, sensor interface circuitry 168 for outputting data from the plurality of sensors, image device interface circuitry 170 for receiving the output data from any or all of the cameras and / or thermal cameras, input and output (I / O) devices 172, and / or communication circuitry 192 to enable the control computer 150 to communicate over any suitable communication network.
[0037] In some embodiments, the I / O devices 172 may include, for example, a display 186 and / or a keyboard 184. Keyboard 184 may allow a user or operator of system 110 to input data to the control computer 150. The at least one processor 160 may control a graphic user interface (GUI) 188 displayed on the display 186. The GUI 188 may display any suitable parameters and / or data visualizations related to the experimental session and / or results of analyses of the data acquired by the plurality of sensors coupled to enclosure 120 in accordance with the experimental plan.
[0038] In some embodiments, a head mount 190 may be placed on the subject’s head (e.g., skull). The head mount may include at least one electrode to measure brain electrical activity such as electroencephalogram (EEG) signals, for example. In some embodiments, the head mount 190 may also include at least one accelerometer. The signals from the at least one electrode and / or the at least one accelerometer in the head mount 190 may be coupled via wires to circuitry 108 that can relay the signals for processing to the at least one processor 160.
[0039] In some embodiments, system 110 can be configured to acquire electrophysiological data, such as pharmaco-EEG (pEEG) data or the like. System 110 can be configured to record actigraphy and quantitative pEEG from one or more brain regions of unanesthetized nonhuman animals before and after administration of a compound. Such electrophysiological data can be used, consistent with disclosed embodiments, to identify novel compounds that have a desired effect on electrophysiological activity (e.g., pEEG activity).
[0040] In some embodiments, system 110 can be configured to phenotype animal models of disease, dysfunction, or injury. As electrophysiological data can yield pharmaco-dynamic signatures specific to pharmacological action, such data can be used to evaluate translationalbiomarkers and rapidly screen compounds for potential activity at specific pharmacological targets to provide valuable information for guiding the early stages of drug development.
[0041] FIG. 2 illustrates a graph 200 of multivariate time series data, consistent with disclosed embodiments. The multivariate time series data can include animal behavior data (e.g., observational data) and optionally trial state data, as described herein. The multivariate time series data can include multiple channels of univariate time series data (e.g., first channel 202, second channel 204, third channel 206, fourth channel 208, fifth channel 210, sixth channel 212, etc.). An exemplary number, arrangement, and type of channels depicted in FIG. 2. In some embodiments, a trial may include multiple trial states. The state of the trial can be indicated by a value recorded within a channel of the multivariate time series data. In this example, first channel 202 includes trial state data. The values recorded in first channel 202 may correspond to different trial states. When the trial begins, the trial may be in a first state and a corresponding first value may be recorded in first channel 202. When the trial enters a second state (e.g., application of an adverse stimuli to the non-human animal), a corresponding second value may be recorded in first channel 202. As depicted in FIG. 2, the trial may begin at a start time point 216 and continue until end time point 218. Different trial states within a trial can be indicated by differing values of first channel 202 (e.g., recording period 214 having period start time 220 and period end time 222).
[0042] Consistent with disclosed embodiments, the animal behavior data can include or depend upon measurements from any of the sensors or cameras described herein. In some embodiments, for example, the animal behavior data can include a univariate time series extracted from video data. Each data point in such a univariate time series may correspond to (e.g., be extracted or derived from) at least one frame of the video data. As an example, second channel 204 may include measurements of object velocity, third channel 206 may include measurements of a position or changes in position of a body part of the object, fourthchannel 208 may include measurements of the heading or orientation of the object, fifth channel 210 may include temperature data, and sixth channel 212 may include EEG data of the object.
[0043] In some embodiments, the multivariate time series data can include a univariate time series generated using other univariate time series in the multivariate time series data. For example, a recording channel can include correlation data expressing the correlation between two other recording channels. Such correlation data can provide information or insight regarding specific behaviors. For example, overall patterns or trends of signals may be different due to differing drug types and / or doses of drugs.
[0044] As may be appreciated, a recording channel of the multivariate time series data can include sparse data (or intervals of sparse data). Sparse data may include a high proportion of non-informative measurements (e.g., many zero or near zero values, many null values, or the like), or measurements that otherwise contribute little to analysis. Data 224 can be an example of such an interval. For example, second channel 204 can include head velocity measurements and data 224 may correspond to an interval in which the non-human animal is asleep (and therefore have low or near zero head velocity). In contrast, data 226 can be an example of meaningful data (e.g., an interval in which the non-human animal is awake).Intervals of meaningful data can be unevenly distributed over time. For example, a recording channels can include both intervals of scarce data and intervals of meaningful data. It may be beneficial to reduce sparce data, without removing underlying trends, phenomena, or causations, in order to conserve memory and enable more efficient use of time series data.
[0045] FIG. 3 depicts a method 300 for training and using a machine learning model to predict an administration effect of a compound on a human, consistent with disclosed embodiments. Method 300 can further include using the model to generate other, simplerpredictive models. For convenience of description, method 300 may be described herein as being performed by a machine learning system (e.g., such as computer 150, or the like). However, the disclosed embodiments are not so limited. In some embodiments, a data acquisition system separate from the machine learning system can obtain the training data and generate the training dataset. In some embodiments, a training system separate from the machine learning system can train the machine learning model using a previously generated training dataset. In some embodiments, a production system separate from the machine learning system can obtain the trained machine learning model can generate predictions using the trained machine learning model.
[0046] Method 300 may involve a step 302 of obtaining training data. Consistent with disclosed embodiments, the machine learning system can obtain training data. In some embodiments, the machine learning system can receive, request, or create the training data. For example, the machine learning system can receive or request data from sensors associated with system 110, or from a component of system 110 (e.g., control computer 150 or the like). In some embodiments, the machine learning system can create or generate the training data. For example, the machine learning system can acquire the training data using a testing enclosure (e.g., computer-controlled enclosure 120 or the like).
[0047] Consistent with disclosed embodiments, the training data can be or include multivariate time series data. The multivariate time series data can include multiple channels of univariate time series animal behavior data. In some embodiments, the multivariate time series data may include data from different recording channels. In some embodiments, the univariate time series animal behavior data can include electroencephalography data, an animal body part position time series, an animal body part velocity time series, or an animal temperature time series.
[0048] Method 300 may involve a step 304 of generating a training dataset. Consistent with disclosed embodiments, the machine learning system can generate the training dataset using the obtained training data. In some embodiments, the training dataset can include training samples. Each training sample can include input data associated with a class label. The input data can include a sequence of input block(s). The machine learning system can generate the sequence of inputs block(s) using the obtained training data through one or more operations of data processing, augmentation, filtering, or the like, as described herein.
[0049] Consistent with disclosed embodiments, the machine learning system can generate a sequence of time series blocks from training data. A time series block can include a portion of the training data. The portion can be defined by a starting timepoint and a block duration. The portion can include the subset of the training data acquired, created or recorded at, or associated with, times between the starting timepoint and a final timepoint determined by the starting timepoint and the block duration (e.g., the sum of the starting timepoint and the block duration, or the like). The machine learning system can store the portion in the time series block. The disclosed embodiments are not limited to any particular data structure or format for the time series block.
[0050] A starting timepoint of a time series block earlier in the sequence can be earlier in time than a starting timepoint of a time series block later in the sequence. The starting timepoint of the initial time series block in the sequence can be, or precede, the first timepoint in the multivariate time series data. Alternatively, the starting timepoint of the initial time series block can be later in the multivariate time series data. A final time point of the last series block in the sequence can be or succeed the last timepoint in the multivariate time series data. Alternatively, the final timepoint of the last time series block can be earlier in the multivariate time series data.
[0051] In some embodiments, determining a starting timepoint may involve selecting a beginning point in the time series data. For example, a starting timepoint may include a moment in time, and / or a specific datapoint or frame selected as a beginning. In some examples, the starting timepoint may be randomly determined. Alternatively, the starting timepoint may be deterministic.
[0052] In some embodiments, two or more time series blocks in the sequence can overlap. The starting timepoint of the later block in each pair of overlapping blocks can be earlier than the final timepoint of the earlier block in the pair of overlapping blocks. In various embodiments, two or more time series blocks in the sequence can be contiguous. In some embodiments, two or more time series blocks in the sequence can be separated by a time interval. When multiple pairs of time series blocks overlap (or are separated by a time interval) the overlap (or time interval) can be the same or can differ between pairs. In some embodiments, time series blocks can be generated by sliding a window along the training data. The duration can be the length of the window. In some embodiments, two consecutive windows can overlap by a predetermined amount or percentage.
[0053] In some embodiments, a block duration can depend on the machine learning model architecture, number of blocks, degree of block overlap, the length of the trial, the percentage of the trial included in the training data, some combination of the foregoing, or the like. In some embodiments, block duration may be predetermined (e.g., a default value, a value received from a user or another system, or the like). In some embodiments, determining a block duration can include selecting or identifying an amount of data to allocate to a time series block (e.g., a number of data points, video frames, or the like). In some embodiments, the number of blocks in the input data can depend on the machine learning model architecture, block duration, degree of block overlap, length of the trial, percentage of the trial included in the training data, some combination of the foregoing, or the like.
[0054] In some embodiments, when the training data includes multivariate time series data comprising multiple univariate time series (e.g., multiple recording channels), the portion of the training data stored in a block can include one or more of the univariate time series (e.g., one or more of the multiple recording channels).
[0055] In some embodiments, the machine learning system can perform regularization operations on the training samples. The regularization can be performed to improve the training data or usability of the training data (e.g., by reducing the potential for model overfitting or underfitting). The regularization operations can include permutation, dropout, dilution, combining across trials or animals, or other suitable regularization operations.
[0056] Permutation regularization can include permuting the order of the sequence of blocks (e.g., by swapping or shuffling the order of the blocks in the sequence). The permutation of the block sequence order can be deterministic or random. As may be appreciated, in some embodiments the order of the block sequence can be maintained.
[0057] Dropout regularization can include disregarding a portion of the input when training the machine learning model in step 306. Dropout regularization may prevent overfitting and may also reduce the number of sparse data points in the training data, thereby reducing computational loads.
[0058] In some embodiments, dropout regularization can include selecting a subset of the blocks in the sequence of blocks. The remaining blocks can be disregarded (e.g., the values of the blocks can be set to zero, the weights corresponding to these blocks in the machine learning model can be set to zero, or any other suitable method).
[0059] In some embodiments, dropout regularization can include selecting portions of the blocks in the sequence of blocks. For example, when the blocks each contain multiple univariate time series, dropout regularization can include selecting a subset of combinationsof blocks and univariate time series (e.g., a first subset of the time series in the first block, a second, potentially different subset of the time series in the second block, etc.). The remaining combinations of blocks and univariate time series can be disregarded (e.g., the values of the univariate time series can be set to zero, the weights corresponding to these univariate time series in the machine learning model can be set to zero, or any other suitable method). In some embodiments, individual data points can be disregarded.
[0060] The blocks (or portions of blocks) can be selected randomly or deterministically. In some embodiments, the number or percentage of the blocks selected (e.g., the dropout percentage) can be predetermined. In some embodiments, the dropout percentage can vary during training of a machine learning model.
[0061] Dilution regularization can include setting a portion of parameters in the machine learning model to zero (e.g., setting a portion of the weights in a neural network model to zero.)
[0062] Combining across trials or animals can include generating a training dataset that includes data from different trials of the same compound. The different trials may use the same animal or different animals. In some embodiments, the training dataset can include training samples from different trials. In some embodiments, training samples in the training dataset can include blocks from different trials.
[0063] For convenience of description, these regularization operations are described with reference to the generation of the training dataset in step 304. However, the disclosed embodiments are not so limited. Some regularization operations (e.g., dropout or dilution) may be implemented during the training of the machine learning model in step 306. Some regularization operations (e.g., permutation regularization or combining across trials or animals) may be implemented during the generation of the training dataset in step 304.
[0064] In some embodiments, generating a training sample can include associating the input data with a class label. A class label may be any suitable indication of effect(s) of administering the compound to humans. In various embodiments, the class label can be a name, alphanumeric characters(s), value(s), zero-one- or likelihood-valued vectors, or other suitable indications of the administration effect. For example, the class label “D” can indicate that the compound can act as a therapeutic antidepressant. As another example, a vector can include elements corresponding to effects of administering the compound to humans. In this example, the first position corresponds to “acts as therapeutic antidepressant,” while the second position corresponds to “acts as an analgesic.” In some embodiments, a compound that acts as an analgesic but not a therapeutic antidepressant may have a first value (e.g., zero or false) stored in the vector at the first element and a second value (e.g., one or true) stored in the vector at the second position. In some embodiments, a compound with a high probability of acting as an analgesic and a low probability of acting as a therapeutic antidepressant may have a higher first value stored in the vector at the first element and a lower second value stored in the vector at the second position. The first and second values can indicate the absolute or relative probabilities of the compound acting as an analgesic and as a therapeutic antidepressant, respectively.
[0065] In some embodiments, indication(s) of effects of administering the compound to humans can include a control indication and disease model(s) indications. A disease model indication can indicate that a behavioral phenotype of the animal is consistent with an animal model of disease, dysfunction, or injury. Such indications can be used in applications in which compounds are screened to identify compounds that recover a control or wild-type behavioral phenotype.
[0066] The disclosed embodiments are not limited to a particular method of associating the input data with a class label. For example, the class label can be stored together with the inputdata in the training sample. As an additional example, the class label can be stored separately and associated with the training sample through a reference, index, or the like.
[0067] Method 300 may involve a step 306 of training a machine learning model using the training dataset to generate an indication of an administration effect. In some embodiments, the machine learning model may be or include a neural network model. The neural network model can include an encoder and a classifier. The encoder can be configured to map the input data to a point in a high-dimensional space (e.g., a vector). The classier can be configured to map the point in the high dimensional space to an output.
[0068] In some embodiments, the output can indicate a particular class. For example, the output can indicate that the compound is predicted to belong to a class of compounds having a particular administration effect. For example, the output can be a zero-one valued vector. In this example, an element corresponding to the hallucinogen class can have the value one. All the other elements can have the value zero. In this manner, the output can indicate that the compound is predicted to act as a hallucinogen when administered to humans.
[0069] In some embodiments, the output can indicate a set of classes. For example, the output can indicate a set of output class values. In some embodiments, an output class value can indicate a likelihood that the compound has a particular administration effect. For example, the output can include an output class value of 0.7 for the antidepressant class and an output class value of 0.45 for the hallucinogen class. These output class values can indicate a 70% chance that the compound behaves as an antidepressant when administered to a human and a 45% chance that the compound behaves as a hallucinogen when administered to a human. In some embodiments, an output class value can represent a contribution of the class to the administration effect of the compound. When each class is associated with a particular administration effect profile, the set of output class values can indicate adecomposition of the predicted administration effect profile of the compound in terms of these per-class administration effect profiles. For example, when the set of output class values includes values of 0.4 for the antipsychotic class, 0.4 for the cognitive enhancer class, 0.1 for the antidepressant class, and 0.1 for the mood stabilizer class, the predicted administration effect profile of the compound can be weighted combination of the per-class administration effect profiles for the antipsychotic class (40%), the cognitive enhancer class (40%), the antidepressant class (10%), and the mood stabilizer class (10%).
[0070] In some embodiments, the output can indicate a subclass of the compound. The output can indicate the subclass of the compound in addition to the class of the compound, or in place of the class of the compound. For example, the output can specify both that the compound is an antipsychotic and that the compound is an atypical antipsychotic or can specify only that the compound is an atypical antipsychotic (thus implicitly specifying that the compound is an antipsychotic).
[0071] In some embodiments, when the classifier outputs both a class and a subclass, the classifier can be a two-stage classifier. In some embodiments, a first stage of the classifier can identify the class. The identified class can then be used as an additional input in a second classifier that identifies the subclass. In some embodiments, a collection of second stages can be associated with different classes. A suitable second stage can be selected based on the output of the first stage of the classifier. In some embodiments, the input to the second stage can be obtained from the first stage. In some embodiments, the input to the second stage can be obtained from the encoder. The second stage(s) can be trained together with the first stage or can be trained separately.
[0072] In some embodiments, the output can indicate a likelihood of at least one of a control class or an animal model class. For example, an administration effect can indicate that thebehavioral phenotype of a non-human animal subject — the non-human subject being an animal model of a disease, dysfunction, or injury — matches a control behavioral phenotype when administered the compound. In this manner, the output can predict that a human having the disease, dysfunction, or injury exhibits a therapeutic benefit when administered the compound. In some embodiments, a degree or recovery or therapeutic benefit can be assessed or assigned based on a likelihood value of a control class and / or a likelihood value of the animal model class. For example, the greater the likelihood value of the animal model class, the less therapeutic effect assessed or assigned to the compound. As an additional example, the greater the likelihood value of the control class (or of the ratio of the control class to the animal model class, or the like), the greater the therapeutic benefit assessed or assigned to the compound.
[0073] The machine learning model can be trained on at least a portion of the training dataset generated in step 304. In some embodiments, the machine learning model can be trained in batches. In each batch, a number of training samples can be used to generate outputs. A loss can be generated according to a loss function by comparing, for the training samples in the batch, the output for each training sample to the class label for that training sample. The machine learning model can then be updated according to the loss function. The machine learning model can be trained until a termination condition is satisfied. The termination condition can depend on the performance of the machine learning model, a number of batches (or epochs) used in training, a time, or another suitable termination condition.
[0074] In some embodiments, the trained machine learning model can be used for prediction, as depicted in FIG. 3. Method 300 may involve a step 308 of obtaining prediction data. In some embodiments, a machine learning system can receive, request, or create the prediction data. For example, the machine learning system can receive or request prediction data from sensors associated with system 110, or from a component of system 110 (e.g.,control computer 150 or the like). In some embodiments, the machine learning system can create or generate the prediction data. For example, the machine learning system can acquire the prediction data using a testing enclosure (e.g., computer-controlled enclosure 120 or the like).
[0075] Similar to the training data described herein, the prediction data can be animal behavioral data, consistent with disclosed embodiments. The data can be acquired from a non-human animal during a trial. A compound can be administered to the animal during the trial. The animal behavior data can be or include multivariate time series data. The multivariate time series data can be or include univariate time series data.
[0076] Method 300 may involve a step 310 of generating a prediction sample. In some embodiments, machine learning system can generate the prediction sample. Similar to the generation of the training sample, generation of the prediction sample can include creating a sequence of blocks from the prediction data. In some embodiments, the prediction sample may not be subject to regularization operations, unlike the training sample.
[0077] Method 300 may involve a step 312 of predicting an administration effect using the trained machine learning model generated in step 306 and the prediction sample. As may be appreciated, the machine learning system may be configured to obtain the trained machine learning model (or prediction sample), if necessary. In such embodiments, the machine learning system can receive or retrieve the machine learning model (or prediction sample) from a memory or from another system (e.g., the machine learning system that trained the model or generated the prediction sample). The machine learning system can be configured to provide an indication of the predicted administration effect to a user or another system. For example, the machine learning system can display the predicted administration effect in a graphical user interface or store a report indicating the predicted administration effect in amemory accessible to the machine learning system or provide such a report to another system.
[0078] In some embodiments, the trained machine learning model can be used generate second models, as depicted in FIG. 3. In some embodiments, such secondary models may be more lightweight than the machine learning models trained in step 306. For example, such models may include fewer parameters, have simpler architectures, or otherwise require less computational resources. In some embodiments, such secondary models may support implementation modes not supported by the trained machine learning model. For example, the trained machine learning model requires that a trial be completed before the trial can be analyzed. In contrast, a secondary model can be an online model that accepts training data as that training data is generated. Consequently, the secondary models may be suitable for use in applications (or on devices) that are unsuitable for (or cannot practicably execute) the trained machine learning model. In some embodiments, the features relied upon by the trained machine learning model can indicate underlying physiological effects suitable for additional study. Accordingly, the identification of significant input data may be of clinical or research importance.
[0079] Method 300 may involve a step 314 of identifying significant input data. Consistent with disclosed embodiments, the significant input data can include data determined to be important to the output of the trained machine learning model. For example, significant input data may include specific measurements, channels, or time intervals which may be weighed heavily by weights in the trained machine learning model. Disclosed embodiments involve identifying, using input data and one or more attention scores generated by the machine learning model, a time period, a univariate time series, or a time period of a univariate time series within the first block. The determination of significant input data can be performed manually, automatically (e.g., by the machine learning system), or semi automatically. Insome embodiments, identifying significant input data can include identifying within-channel portions of significant input data (e.g., within a univariate time series) and across-channel portions of significant input data (e.g., across multiple univariate time series).
[0080] For example, as discussed herein, a machine learning model such as model 500, as referenced in FIG. 5, may generate attention scores based on blocks or embeddings of blocks. In some embodiments, significant input data can be identified using such attention scores. For example, high attention score(s) may indicate certain channel(s) or time period (s) have a high impact on the decision of the drug class. As such, the channel or time period may exhibit a signature of the drug class, which may be learned by another, simpler machine learning model.
[0081] In some embodiments, an attention score associated with a temporal transformer can be used to identify a time period in the input data that has a significant impact on the decision of drug class. Given the weights of the temporal transformer, an attention score, and a combined feature set input to the temporal transformer, a subset of the combined feature set can be identified. This subset of the combined feature set can include those datapoints that contribute the most to the attention scores. As may be appreciated, the combined feature set can be generated using the input to the temporal transformer, an encoding layer, and a positional encoding. Given the weights of the encoding layer, the positional encoding, and the subset of the combined feature set, time period(s) of the input to the temporal transformer can be identified. These time period(s) can include those datapoints that contribute the most to the subset of the combined feature set. For example, a time period within a block, or a time period spanning multiple blocks, may be identified as having the most impact on the decision of the drug class.
[0082] In some embodiments, an attention score associated with a channel transformer can be used to identify a channel in the input data that has a significant impact on the decision of drug class. Given the weights of the channel transformer, an attention score, and a feature set input to the temporal transformer, a subset of the feature set can be identified. This feature subset can include those datapoints that contribute the most to the attention scores. As may be appreciated, the feature set can be generated using the input to the channel transformer and an encoding layer. Given the weights of the encoding layer and the feature subset, channel(s) of the input to the channel transformer can be identified. These channel(s) can include those datapoints that contribute the most to the feature subset. For example, a particular univariate channel may be identified as having the most impact on the decision of the drug class. As an example, a certain recording channel, such as second channel 204 corresponding to object velocity (as referenced in FIG. 2), may include data that is a signature of a specific drug.
[0083] Method 300 may involve a step 316 of generating an updated training data set using the identified portions of the input data. For example, when a portion of a univariate time series (or a portion of a multivariate time series) is determined to be significant, a label can be associated with that portion of the univariate time series (or multivariate time series). The portion associated with the label can be included as a positive sample in the updated training dataset, while other unlabeled portions of the univariate time series (or multivariate time series) can be included in the updated training dataset as negative examples. These unlabeled portions can be from the same trial or from a control trial (e.g., a trial in which the animal is administered nothing, another compound, a vehicle, or another suitable control). As an additional example, when a subset of the channels in a multivariant training dataset are identified as contributing to the output of the trained machine learning model, an updated version of the multivariant training dataset can include only those channels identified as significant.
[0084] In some embodiments, the significant input data (e.g., the identified time period, univariate time series, or time period of the univariate time series) can be used by computer 150 to generate an updated training sample. For example, the updated training sample may include univariate time series data which was not presented to the machine learning model during training. In some embodiments, the updated training sample can include the significant input data, or the training sample may be capable of associating the significant input data with a predicted administration effect. For example, the machine learning system can query the updated training sample to predict the indication of the administration effect. For example, some disclosed embodiments can involve generating a training sample, at least in part by associating the time period, univariate time series, or time period of the univariate time series with the indication of administration effect. A training sample may include training data based on other training data, outputs, or classifications from a machine learning model. For example, the training data may be time periods or channels identified as having a high impact on the prediction of the drug class, as described elsewhere in this disclosure. It will be appreciated that the updated training set may enable faster predictions of administration effects of compounds administered to non-human animal 115.
[0085] Method 300 may involve a set 318 of training a secondary machine learning model. In some embodiments, the machine learning system can train the secondary machine learning model using the updated training data generated in step 316. For example, when the updated dataset includes positive and negative samples of univariate time series data, a detector can be trained using the samples. The detector can be trained to detect whether input data has been acquired from an animal administered a compound belong to a particular class. For example, a training dataset can be generated from portions of input data significant in predicting (e.g., by the trained machine learning model) that a compound administered to a non-human animal has the administration effect of a hallucinogen when administered to ahuman. This training dataset can be used to train a detector to identify input data as having been acquired from an animal that had been administered such a compound.
[0086] In some embodiments, the updated training dataset can use intermediate outputs of the machine learning model to generate secondary models. For example, when the machine learning model includes an encoder and a classifier, a training dataset can be generated from the output of the encoder. Each training sample can generate an output vector. The output vectors can then be used to train a machine learning model. For example, the output vectors can be clustered, and the semantic significance of the clusters determined. In some embodiments, class labels associated with the compounds can be used to determine the semantic significance of the clusters. Using such a clustering, additional compounds can be identified without using the classifier in the trained machine learning model. For example, referring to FIG. 5, output 534 from model 500 may include input vectors used for clustering. As an example, augmented encodings 422 or summary output 424 may be applied to a machine learning model which generates clusters of compounds. Each cluster in the cluster of compounds may correspond to an indication of an administration effect, such as corresponding to one or more drug classes or drug subclasses.
[0087] FIG. 4 illustrates a diagram of an architecture 400 for predicting an administration effect of a compound on humans using observational data obtained from non-human animals, consistent with disclosed embodiments. Training a machine learning model using architecture 400 can be performed using a machine learning system, as described herein.
[0088] Multivariate data 402 may include multiple channels of univariate time series data, consistent with disclosed embodiments. Multivariate data 402 may be or include animal behavioral data acquired during one or more trials, such as trials performed on non-human animal 115 using system 110, as described in FIG. 1. As described herein, trials may includeadministering one or more compounds with a known administration effect in humans to nonhuman animal 115. Multivariate data 402 may include a number of channels. In some embodiments, each channel can be or include univariate time series data. For example, the univariate time series data may include animal behavioral data or trial state data. Multivariate data 402 may have dimensions of length of data by number of channels.
[0089] Time series blocks 408 may be generated by selecting portions of multivariate data 402, consistent with disclosed embodiments. The machine learning system can create a time series block by selecting a portion of the multivariant time series with a block duration 406 and a starting timepoint 404. In the example depicts in FIG. 4, the certain time series blocks are consecutive, with the starting timepoint of block bi+1being, or being adjacent to, the final timepoint of blockThe disclosed embodiments are not so limited, however. In the example depicted in FIG. 4, certain time series blocks are overlapping (e.g., block 414). In some embodiments, the start time of each time series block can be randomly selected.
[0090] Input blocks 410 may be selected from time series blocks 408, consistent with disclosed embodiments. Input blocks 410 may include a subset of time series blocks 408. In this example, each input block can include T samples, each sample including a value from C channels. Thus, the input block can include a multivariate time series including C univariate time series. The multivariant time series can be represented as a C by T matrix. In some embodiments, the selection of the subset of input blocks can be random.
[0091] Input blocks 410 may be input to encoder 418, consistent with disclosed embodiments. In some embodiments, encoder 418 may be or include an attention-based neural network model, such as a transformer model. In some embodiments, encoder 418 can include a temporal encoder and a channel encoder. The temporal encoder can process time series values for each channel, while the channel encoder can process the set of channelvalues for each time point. The temporal encodings and channel encodings can be combined to generate an encoder output. Thus, the input blocks 410 can be applied to encoder 418 to generate corresponding encoder outputs 420. Thus encoder 418 can generate a sequence of encoder outputs 420 from a sequence of input blocks 410. As may be appreciated, an encoder output can provide a latent space representation of a corresponding input block. Encoder outputs 420 may have dimensions M + M, where M is a length of the encoding vector.
[0092] In some embodiments, encoder outputs 420 may be concatenated with positional information to generate positionally augmented encodings 422, with dimension M + M. The positional information can include the position of the input block used to generate the output encoding in the sequence of input blocks 410. In some embodiments, the positional information can be a number (e.g., 1, 2, 3, etc.).
[0093] In some embodiments, a summary output 424 can be generated from the positionally augmented encodings 422. The summary output may have dimensions of positionally augmented encodings (e.g., M + M + 1) by number of blocks included in the sequence of input blocks 410.
[0094] In some embodiments, a summary output 424 can be generated from encoder outputs 420, without augmenting the encoder outputs 420 with positional information. In such embodiments, summary output 424 can be the concatenation of the encoder outputs 420, or a function of the encoder outputs 430. For example, summary output 424 can be the length M + M average of the encoder outputs 430. This averaged output could be input to a suitable classifier. As an additional example, a preliminary classifier can be included in encoder 418 (e.g., one or more feedforward layers, or the like). The preliminary classifier can be configured to output a class token for each input block 410. The class tokens can beconcatenated together (with or without positional augmentation) and input a suitable classifier.
[0095] In some embodiments, summary output 424 can be input to classifier 425. Classifier 425 can be any machine learning model configured to generate a classification outcome. For example, classifier 425 may include neural networks, regression models, random forests, or the like. Classification output 426 may be an indication of administration effect. For example, classification output 426 be a vector with elements corresponding to various classes (or subclasses) of drugs. The value of an element (e.g., p_0, p l, etc.) can indicate the likelihood of the compound belongs to the corresponding element, or the contribution of the corresponding administration effect profile to the predicted effect profile of the compound.
[0096] In some embodiments, architecture 400 can be adapted during training to regularize the input training data. As may be appreciated, a training dataset can include multiple training samples, each including input data associated with a class label. The input data can include a sequence of input blocks 410. As may be appreciated, multiple training samples can be generated from the same multivariate data 402 (e.g., using randomized start timepoints and randomized selection of input blocks 410 from time series blocks 408).
[0097] In some embodiments, architecture 400 can be augmented with an activity detector portion. The activity detector portion can complement classifier 425 described above. The activity detector portion of the machine learning model can be implemented using a separate encoder and classifier, a separate classifier, or a separate classifier branch that shares one or more layers with classifier 425 in the machine learning model. The activity detector portion can be configured to classify the compound as “active” (or an equivalent label) or “inactive” (e.g., vehicle, control, default, or some equivalent label). In some embodiments, the output of architecture 400 can depend on the output of the therapeutic prediction portion (e.g., thetherapeutic class(es) predicted for the different dosage levels) and the output of the activity detector portion.
[0098] In some embodiments, the classifier 425 and the activity detector portion can output probability values. When the sum of the probabilities for all the therapeutic class(es), for a dosage level and excluding any “control” or semantically similar class (e.g., “vehicle,” “no response”, “default,” or the like), exceeds the output probability for the “activity” class of the activity detector portion, then the overall output of architecture 400 can be the probabilities output by classifier 425. Otherwise, the overall output of architecture 400 can indicate that the class of the compound is unknown. In some embodiments, a suitable probability value can be assigned to an “unknown - active” class, or a semantic equivalent. In some embodiments, a suitable probability value can be assigned to a “control” or semantically similar class.
[0099] Consistent with disclosed embodiments, input data in a training sample can be regularized through dropout regularization or permutation regularization. In some embodiments, the machine learning system can permute the order of time series blocks 408 (e.g., as by swapping 412 pairs of time series blocks). In some embodiments, dropout regularization can be performed at the block level. For example, one or more input blocks may be omitted from use in training (e.g., block 416). In various embodiments, the values of the input block can be set to zero, the resulting output encoding can be set to zero, the weights corresponding to this outpoint encoding in the classifier can be set to zero, or another suitable method of disregarding this omitted input block can be used. As may be appreciated, dropout blocks can be selected from among the input blocks 410 randomly. The number selected can depend on a dropout percentage or the like.
[0100] In some embodiments, dropout regularization can be performed at the univariate time series level or individual datapoint level. For example, at least a portion of the univariate time series data (or a selection of individual datapoints) in one or more of the input blocks can be disregarded. In various embodiments, the values of the portion of univariate time series data (or individual datapoints) can be set to zero, corresponding weights in the encoder can be set to zero, or another suitable method of disregarding this portion of univariate time series data or individual datapoints) can be used. As may be appreciated, the portions of the univariate time series data or the individual datapoints can be selected from among the input blocks 410 randomly. The size of the portions or the number of datapoints selected can depend on a dropout percentage or the like.
[0101] FIG. 5 depicts an exemplary architecture of a machine learning model 500, consistent with disclosed embodiments. In some embodiments, model 500 can be used to implement encoder 418 of FIG. 4. Machine learning model 500 can be implemented using a machine learning system, as described herein. Model 500 may include one or more transformer networks, such as channel transformer 514 and temporal transformer 512.
[0102] Channel transformer 514 can be a channel encoder configured to generate encodings based on univariate time series data from one or more recording channels, consistent with disclosed embodiments. Channel transformer 514 can be configured to use self-attention to attend, across all time steps, on each channel at the same time point. Pair-wise attention weights can be calculated among all of the channels. As the ordering of channels can be arbitrary, positional encodings may not be used with channel transformer 514 (e.g., to ensure the equal contributions of each channel). Temporal transformer 512 can be a temporal encoder configured to generate encodings based on time-based or position-based data. Temporal transformer 512 can be configured to use self-attention with mask to attend on each time point of a single channel (e.g., but repetitively calculated across all channels). Pair-wiseattention weights can be calculated among all of the time steps. Masking can be used to ensure that attention depends only on current and prior time steps.
[0103] Input 502 to model 500 can be multivariate time series data (e.g., multivariate data 402, or the like). Such data can be behavioral data of non-human animal 115 obtained from system 110. Input 502 may be projected to temporal embedding layer 504 and a channel embedding layer 506.
[0104] Model 500 can be configured to generate a combined temporal feature set by applying input 502 to temporal embedding layer 504. An embedding may be a representation of data in a space, such that data sharing similar characteristics are positioned closer together in the space. Embeddings may capture or extract information from data, such as various features, and map the features in the space. Embeddings and encodings, as referenced herein, may include vector or numerical representations. Embeddings may be lower-dimensional representations of complex information. The output of temporal embedding layer 504 can be combined with positional encoding 510. A positional encoding may represent the relative position of information in a sequence. For example, a positional encoding may represent the position of a data point in a time series relative to other datapoints in the time series. By including the positional information, the combined feature set may provide the benefit of retaining information pertinent to the order of events or data in the sequence. The combined feature set may be input to temporal transformer 512, which may generate an intermediate temporal output 529.
[0105] Model 500 can be configured to generate a combined channel feature set by applying input 502 to channel embedding layer 506. The output of channel embedding layer 506 can be applied to channel transformer 514 to generate intermediate channel output 531.
[0106] Model 500 can include a gating layer 532, consistent with disclosed embodiments.Gating layer 532 can be trained to learn the relative weights of inputs to the gating layer, such as intermediate channel output 531 and intermediate temporal output 529. Gating layer 532 may take the outputs from the channel encoder transformer 514 and the time encoder transformer 512, and calculate the weighted sum of the outputs, in order to determine associations between meaningful information across time periods and channels. Gating layer 532 may include any suitable gating function, such as linear activation, non-linear activation, ReLU, or SoftMax. Gating layer 532 can generate an output 534, which can be an encoding.
[0107] Model 500 may include features which increase accuracy of training or inference. In some embodiments, channel transformer 514 may include one or more self-attention layer complexes. Each self-attention layer complex can include a self-attention layer (e.g., selfattention layer 518) followed by an addition and normalization layer. A final feed-forward complex can include a feedthrough layer (e.g., feedthrough layer 522) followed by an addition and normalization layer. Similarly, in some embodiments, temporal transformer 512 may include one or more self-attention layer complexes. Each self-attention layer complex can include a self-attention layer (e.g., self-attention layer 516) followed by an addition and normalization layer. A final feed-forward complex can include a feedthrough layer (e.g., feedthrough layer 524) followed by an addition and normalization layer.
[0108] FIG. 6 depicts an exemplary computing system 600 suitable for generating machine learning models, consistent with disclosed embodiments. In some embodiments, the machine learning system described herein can be implemented using computing system 600.Similarly, method 300 can be performed, and / or architecture 400 can be implemented using computing system 600. As will be appreciated by one skilled in the art, the components and arrangement of components included in computing system 600 may vary. For example, as compared to the depiction in FIG. 6, computing system 600 may include a larger or smallernumber of processors, I / O devices, or memory units. In addition, computing system 600 may further include other components or devices not depicted that perform or assist in the performance of one or more processes consistent with the disclosed embodiments. The components and arrangements shown in FIG. 6 are not intended to limit the disclosed embodiments, as the components used to implement the disclosed processes and features may vary.
[0109] Processor 610 may comprise known computing processors, including a microprocessor. Processor 610 may constitute a single-core or multiple-core processor that executes parallel processes simultaneously. For example, processor 610 may be a single-core processor configured with virtual processing technologies. In some embodiments, processor 610 may use logical processors to simultaneously execute and control multiple processes. Processor 610 may implement virtual machine technologies, or other known technologies to provide the ability to execute, control, run, manipulate, store, etc., multiple software processes, applications, programs, etc. In another embodiment, processor 610 may include a multiple-core processor arrangement (e.g., dual core, quad core, etc.) configured to provide parallel processing functionalities to allow execution of multiple processes simultaneously. One of ordinary skill in the art would understand that other types of processor arrangements could be implemented that provide for the capabilities disclosed herein. The disclosed embodiments are not limited to any type of processor. Processor 610 may execute various instructions stored in memory 630 to perform various functions of the disclosed embodiments described in greater detail below. Processor 610 may be configured to execute functions written in one or more known programming languages.
[0110] Computer program code for carrying out operations, for example, embodiments may be written in any combination of one or more programming languages, including an object- oriented programming language such as Java, Smalltalk, C++ or the like and conventionalprocedural programming languages, such as the “C” programming language or similar programming languages. The program code may execute entirely on the user’s computer, partly on the user’s computer, as a stand-alone software package, partly on the user’s computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user’s computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider).[OHl] I / O device 620 may include at least one of a display, an LED, a router, a touchscreen, a keyboard, a microphone, a speaker, a haptic device, a camera, a button, a dial, a switch, a knob, a transceiver, an input device, an output device, or another I / O device to perform methods of the disclosed embodiments.
[0112] I / O device 620 may be configured to manage interactions between computing system 600 and other systems using a network. In some aspects, I / O device 620 may be configured to publish data received from other databases or systems not shown. This data may be published in a publication and subscription framework (e.g., using APACHE KAFKA), through a network socket, in response to queries from other systems, or using other known methods. In various aspects, I / O 620 may be configured to provide data or instructions received from other systems. For example, I / O 620 may be configured to receive instructions for generating data models (e.g., type of data model, data model parameters, training data indicators, training parameters, or the like) from another system and provide this information to machine learning framework 635. As an additional example, I / O 620 may be configured to receive data from another system (e.g., in a file, a message in a publication and subscription framework, a network socket, or the like) and provide that data to programs or store that data in, for example, data set 632 or model 634.
[0113] In some embodiments, I / O 620 may include a user interface configured to receive user inputs and provide data to a user (e.g., a data manager). For example, I / O 620 may include a display, a microphone, a speaker, a keyboard, a mouse, a track pad, a button, a dial, a knob, a printer, a light, an LED, a haptic feedback device, a touchscreen and / or other input or output devices.
[0114] Memory 630 may be a volatile or non-volatile, magnetic, semiconductor, optical, removable, non-removable, or other type of storage device or tangible (i.e., non-transitory) computer-readable medium, consistent with disclosed embodiments. As shown, memory 630 may include inference data 633 and training data set 632, including one of at least one of encrypted data or unencrypted data. Memory 630 may also include models 634, including weights and parameters of neural network models.
[0115] Machine learning framework 635 may include one or more programs (e.g., modules, code, scripts, or functions) used to perform methods consistent with disclosed embodiments. Programs may include operating systems (not shown) that perform known operating system functions when executed by one or more processors. Disclosed embodiments may operate and function with computer systems running any type of operating system. Machine learning framework 635 may be written in one or more programming or scripting languages. One or more of such software sections or modules of memory 630 may be integrated into a computer system 600, non-transitory computer-readable media, or existing communications software. Machine learning framework 635 may also be implemented or replicated as firmware or circuit logic.
[0116] Machine learning framework 635 may include programs (scripts, functions, algorithms) to assist creation of, train, implement, store, receive, retrieve, and / or transmit one or more machine learning models. Machine learning framework 635 may be configured toassist creation of, train, implement, store, receive, retrieve, and / or transmit, one or more ensemble machine learning models (e.g., machine learning models comprised of a plurality of machine learning models). In some embodiments, training of a model may terminate when a training criterion is satisfied. Training criteria may include the number of epochs, training time, performance metric values (e.g., an estimate of accuracy in reproducing test data), or the like. Machine learning framework 635 may be configured to adjust model parameters and / or hyperparameters during training. For example, machine learning framework 635 may be configured to modify model parameters and / or hyperparameters (i.e., hyperparameter tuning) using an optimization technique during training, consistent with disclosed embodiments. Hyperparameters may include training hyperparameters, which may affect how training of a model occurs, or architectural hyperparameters, which may affect the structure of a model. Optimization techniques used may include grid searches, random searches, gaussian processes, Bayesian processes, Covariance Matrix Adaptation Evolution Strategy techniques (CMA-ES), derivative-based searches, stochastic hill-climbing, neighborhood searches, adaptive random searches, or the like.
[0117] In some embodiments, machine learning framework 635 may be configured to generate models based on instructions received from another component of computing system 600 and / or a computing component outside computing system 600. For example, machine learning framework 635 can be configured to receive a visual (e.g., graphical) depiction of a machine learning model and parse that graphical depiction into instructions for creating and training a corresponding neural network. Machine learning framework 635 can be configured to select model training parameters. This selection can be based on model performance feedback received from another component of machine learning framework 635. Machine learning framework 635 can be configured to provide trained models and descriptive information concerning the trained models to model memory 630.
[0118] Any computer program instructions may also be stored in a computer readable medium that can direct one or more hardware processors of a computer, other programmable data processing apparatus, or other devices to function in a particular manner, such that the instructions stored in the computer readable medium form an article of manufacture including instructions that implement the function / act specified in the flowchart or block diagram block or blocks.
[0119] The computer program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus or other devices to produce a computer implemented process such that the instructions that execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart or block diagram block or blocks.
[0120] The embodiments may further be described using the following clauses:
[0121] 1. A training method, comprising: obtaining multivariate time series data including multiple channels of univariate time series animal behavior data, the multiple channels of univariate time series animal behavior data acquired from a non-human animal during a trial in which the non-human animal was administered a compound; generating training samples, generation comprising: generating blocks, generation of a first one of the blocks comprising: determining a block duration and starting timepoint; and storing in the first one of the blocks a portion of the multivariate time series data beginning at the starting timepoint and having the block duration; and associating the blocks with a class label; and training a machine learning model to generate an indication of an administration effect using a training dataset including the training samples.
[0122] 2. The method of clause 1, wherein the univariate time series animal behavior data includes: electroencephalography data, an animal body part position time series, an animal body part velocity time series, or an animal temperature time series.
[0123] 3. The method of any one of clauses 1 to 2, wherein the multivariate time series data further includes trial state data.
[0124] 4. The method of any one of clauses 1 to 3, wherein the generated indication of an administration effect comprises a vector of likelihood values.
[0125] 5. The method of any one of clauses 1 to 4, wherein the generated indication of administration effect indicates that the compound is one or more of antidepressant, anxiolytic, antipsychotic, cognitive enhancer, hallucinogen, anticonvulsant, mood stabilizer, or psychostimulant when administered to humans.
[0126] 6. The method of any one of clauses 1 to 4, wherein the generated indication of an administration effect indicates a likelihood of a control class or of an animal model class of disease, dysfunction, or injury.
[0127] 7. The method of any one of clauses 1 to 6, wherein the training method includes regularization of the training samples.
[0128] 8. The method of clause 7, wherein: regularization of the training samples includes permutation regularization of the blocks comprising the training samples.
[0129] 9. The method of any one of clauses 7 to 8, wherein: regularization of the training samples includes dropout regularization of the blocks comprising the training samples or portions of the blocks comprising the training samples.
[0130] 10. The method of any one of clauses 1 to 9, wherein the machine learning model comprises a gated transformer, and training the machine learning model using the trainingdata comprises: generating a first embedding using the training sample; generating a combined feature set using the first embedding and a positional encoding; generating a second embedding using the training sample; applying the combined feature set to a timeencoder transformer; and applying the second embedding to a channel-encoder transformer.
[0131] 11. A system comprising: at least one non-transitory computer-readable medium storing instructions; and at least one processor configured to execute the instructions to perform operations for predicting an administration effect of a compound when the compound is administered to humans using animal behavior data, the operations comprising: obtaining multivariate time series data including multiple channels of univariate time series animal behavior data, the multiple channels of univariate time series animal behavior data acquired from a non-human animal during a trial in which the non-human animal was administered the compound; generating a sequence of blocks using the multivariate time series data, a first block of the sequence of blocks including a first portion of the multivariate time series data; generating an indication of the predicted administration effect by applying the sequence of blocks to a machine learning model, the machine learning model configured to: generate a first encoding by applying the first block to an encoder; generate a first input vector using a set of encodings corresponding to blocks in the sequence of blocks, the set of encodings including the first encoding; and generate the indication of the predicted administration effect by applying the first input vector to a classifier.
[0132] 12. The system of clause 11, wherein: the sequence of blocks includes the first block at a first position in the sequence; and generating the first input vector using the set of encodings comprises: generating a first positionally augmented encoding using the first encoding and the first position, and concatenating a set of positionally augmented encodings, the set of positionally augmented encodings including the first positionally augmentencoding; concatenating the encodings in the set of encoding; or averaging the encodings in the set of encoding.
[0133] 13. The system of any one of clauses 11 to 12, wherein the univariate time series animal behavior data includes: electroencephalography data, an animal body part position time series, an animal body part velocity time series, or an animal temperature time series.
[0134] 14. The system of any one of clauses 11 to 13, wherein the multivariate time series data further includes trial state information.
[0135] 15. The system of any one of clauses 11 to 14, wherein: generation of the first encoding comprises: generating a combined feature set using the first block, a temporal embedding layer, and a positional encoding; generating a channel embedding using the first block and a channel embedding layer; generating an intermediate channel output by applying the channel embedding to a channel transformer; generating an intermediate temporal output by applying the combined feature set to a temporal transformer; and generating the first encoding using a gating layer, the intermediate channel output and the intermediate temporal output.
[0136] 16. The system of clause 15, wherein: the operations further comprise: identifying, using the first block and one or more attention scores generated by the machine learning model, a time period, a univariate time series, or a time period of a univariate time series within the first block; generating a training sample, at least in part by associating the time period, univariate time series, or time period of the univariate time series with the indication of administration effect; and training a second classifier to predict drug classification labels using a training dataset including the training sample.
[0137] 17. The system of clause 16, wherein: identifying the time period, the univariate time series, or the time period of the univariate time series within the first block comprises:generating a first attention score using the combined feature set; identifying, using the first attention score, a feature subset of the combined feature set; and identifying, using the temporal embedding layer and the feature subset, the time period.
[0138] 18. The system of clause 16, wherein: identifying the time period, the univariate time series, or the time period of the univariate time series within the first block comprises: generating a second attention score using the channel embedding; identifying, using the second attention score, a feature subset of the channel embedding; and identifying, using the channel embedding layer and the feature subset, the univariate channel.
[0139] 19. The system of any one of clauses 11 to 18, wherein: the operations further comprise: identifying clusters of compounds by clustering a set of input vectors corresponding to the compounds, the set of input vectors including the first input vector.
[0140] 20. A method for predicting compound class labels using animal behavior data, comprising: obtaining multivariate time series data including multiple channels of univariate time series animal behavior data, the multiple channels of univariate time series animal behavior data acquired from a non-human animal during a trial in which the non-human animal was administered a compound; generating a sequence of blocks of the multivariate time series data, a first block of the sequence of blocks including a first portion of the multivariate time series data; generating a class label by applying the sequence of blocks to a machine learning model, the class label indicating an effect of the compound when administered to humans, the machine learning model configured to: generate a first encoding by applying the first block to an encoder; generate a first input vector using a set of encodings corresponding to blocks in the sequence of blocks, the set of encodings including the first encoding; and generate the class label by applying the first input vector to a classifier.
[0141] 21. The method of clause 20, wherein: the sequence of blocks includes the first block at a first position in the sequence; and generating the first input vector using the set of encodings comprises: generating a first positionally augmented encoding using the first encoding and the first position, and concatenating a set of positionally augmented encodings, the set of positionally augmented encodings including the first positionally augment encoding; concatenating the encodings in the set of encoding; or averaging the encodings in the set of encoding.
[0142] 22. The method of any one of clauses 20 to 21, wherein the univariate time series animal behavior data includes: electroencephalography data, an animal body part position time series, an animal body part velocity time series, or an animal temperature time series; and the multivariate time series data further includes trial state information.
[0143] As used herein, unless specifically stated otherwise, the term “or” encompasses all possible combinations, except where infeasible. For example, if it is stated that a component may include A or B, then, unless specifically stated otherwise or infeasible, the component may include A, or B, or A and B. As a second example, if it is stated that a component may include A, B, or C, then, unless specifically stated otherwise or infeasible, the component may include A, or B, or C, or A and B, or A and C, or B and C, or A and B and C.
[0144] It is understood that the described embodiments are not mutually exclusive, and elements, components, materials, or steps described in connection with one example embodiment may be combined with, or eliminated from, other embodiments in suitable ways to accomplish desired design objectives.
[0145] In the foregoing specification, embodiments have been described with reference to numerous specific details that can vary from implementation to implementation. Certain adaptations and modifications of the described embodiments can be made. Otherembodiments can be apparent to those skilled in the art from consideration of the specification and practice of the subject matter disclosed herein. It is intended that the specification and examples be considered as exemplary only. It is also intended that the sequence of steps shown in figures are only for illustrative purposes and are not intended to be limited to any particular sequence of steps. As such, those skilled in the art can appreciate that these steps can be performed in a different order while implementing the same method.
Claims
WHAT IS CLAIMED IS:
1. A training method, comprising: obtaining multivariate time series data including multiple channels of univariate time series animal behavior data, the multiple channels of univariate time series animal behavior data acquired from a non-human animal during a trial in which the non-human animal was administered a compound; generating training samples, generation comprising: generating blocks, generation of a first one of the blocks comprising: determining a block duration and starting timepoint; and storing in the first one of the blocks a portion of the multivariate time series data beginning at the starting timepoint and having the block duration; and associating the blocks with a class label; and training a machine learning model to generate an indication of an administration effect using a training dataset including the training samples.
2. The method of claim 1, wherein the univariate time series animal behavior data includes: electroencephalography data, an animal body part position time series, an animal body part velocity time series, or an animal temperature time series.
3. The method of claim 1, wherein the multivariate time series data further includes trial state data.
4. The method of claim 1, wherein the generated indication of an administration effect comprises a vector of likelihood values.
5. The method of claim 1, wherein the generated indication of administration effect indicates that the compound is one or more of antidepressant, anxiolytic, antipsychotic, cognitive enhancer, hallucinogen, anticonvulsant, mood stabilizer, or psychostimulant when administered to humans.
6. The method of claim 1, wherein the generated indication of an administration effect indicates a likelihood of a control class or of an animal model class of disease, dysfunction, or injury.
7. The method of claim 1, wherein the training method includes regularization of the training samples.
8. The method of claim 7, wherein: regularization of the training samples includes permutation regularization of the blocks comprising the training samples.
9. The method of claim 7, wherein: regularization of the training samples includes dropout regularization of the blocks comprising the training samples or portions of the blocks comprising the training samples.
10. The method of claim 1, wherein the machine learning model comprises a gated transformer, and training the machine learning model using the training data comprises: generating a first embedding using the training sample;generating a combined feature set using the first embedding and a positional encoding; generating a second embedding using the training sample; applying the combined feature set to a time-encoder transformer; and applying the second embedding to a channel-encoder transformer.
11. A system comprising: at least one non-transitory computer-readable medium storing instructions; and at least one processor configured to execute the instructions to perform operations for predicting an administration effect of a compound when the compound is administered to humans using animal behavior data, the operations comprising: obtaining multivariate time series data including multiple channels of univariate time series animal behavior data, the multiple channels of univariate time series animal behavior data acquired from a non-human animal during a trial in which the non-human animal was administered the compound; generating a sequence of blocks using the multivariate time series data, a first block of the sequence of blocks including a first portion of the multivariate time series data; generating an indication of the predicted administration effect by applying the sequence of blocks to a machine learning model, the machine learning model configured to: generate a first encoding by applying the first block to an encoder;generate a first input vector using a set of encodings corresponding to blocks in the sequence of blocks, the set of encodings including the first encoding; and generate the indication of the predicted administration effect by applying the first input vector to a classifier.
12. The system of claim 11, wherein: the sequence of blocks includes the first block at a first position in the sequence; and generating the first input vector using the set of encodings comprises: generating a first positionally augmented encoding using the first encoding and the first position, and concatenating a set of positionally augmented encodings, the set of positionally augmented encodings including the first positionally augment encoding; concatenating the encodings in the set of encoding; or averaging the encodings in the set of encoding.
13. The system of claim 11, wherein the univariate time series animal behavior data includes: electroencephalography data, an animal body part position time series, an animal body part velocity time series, or an animal temperature time series.
14. The system of claim 11, wherein the multivariate time series data further includes trial state information.
15. The system of claim 11, wherein: generation of the first encoding comprises:generating a combined feature set using the first block, a temporal embedding layer, and a positional encoding; generating a channel embedding using the first block and a channel embedding layer; generating an intermediate channel output by applying the channel embedding to a channel transformer; generating an intermediate temporal output by applying the combined feature set to a temporal transformer; and generating the first encoding using a gating layer, the intermediate channel output, and the intermediate temporal output.
16. The system of claim 15, wherein: the operations further comprise: identifying, using the first block and one or more attention scores generated by the machine learning model, a time period, a univariate time series, or a time period of a univariate time series within the first block; generating a training sample, at least in part by associating the time period, univariate time series, or time period of the univariate time series with the indication of administration effect; and training a second classifier to predict drug classification labels using a training dataset including the training sample.
17. The system of claim 16, wherein: identifying the time period, the univariate time series, or the time period of the univariate time series within the first block comprises:generating a first attention score using the combined feature set; identifying, using the first attention score, a feature subset of the combined feature set; and identifying, using the temporal embedding layer and the feature subset, the time period.
18. The system of claim 16, wherein: identifying the time period, the univariate time series, or the time period of the univariate time series within the first block comprises: generating a second attention score using the channel embedding; identifying, using the second attention score, a feature subset of the channel embedding; and identifying, using the channel embedding layer and the feature subset, the univariate channel.
19. The system of claim 11, wherein: the operations further comprise: identifying clusters of compounds by clustering a set of input vectors corresponding to the compounds, the set of input vectors including the first input vector.
20. A method for predicting compound class labels using animal behavior data, comprising:obtaining multivariate time series data including multiple channels of univariate time series animal behavior data, the multiple channels of univariate time series animal behavior data acquired from a non-human animal during a trial in which the non-human animal was administered a compound; generating a sequence of blocks of the multivariate time series data, a first block of the sequence of blocks including a first portion of the multivariate time series data; generating a class label by applying the sequence of blocks to a machine learning model, the class label indicating an effect of the compound when administered to humans, the machine learning model configured to: generate a first encoding by applying the first block to an encoder; generate a first input vector using a set of encodings corresponding to blocks in the sequence of blocks, the set of encodings including the first encoding; and generate the class label by applying the first input vector to a classifier.
21. The method of claim 20, wherein: the sequence of blocks includes the first block at a first position in the sequence; and generating the first input vector using the set of encodings comprises: generating a first positionally augmented encoding using the first encoding and the first position, and concatenating a set of positionally augmented encodings, the set of positionally augmented encodings including the first positionally augment encoding; concatenating the encodings in the set of encoding; or averaging the encodings in the set of encoding.
2. The method of claim 20, wherein the univariate time series animal behavior data includes: electroencephalography data, an animal body part position time series, an animal body part velocity time series, or an animal temperature time series; and the multivariate time series data further includes trial state information.