Decomposed spectra analysis for large model selection and optimization

By comparing label-dependent spectra and reducing model complexity, the proposed method addresses the inefficiencies in existing model selection and optimization processes for natural language processing, resulting in improved efficiency and performance.

JP2025096265APending Publication Date: 2025-06-26テンパスエーアイインコーポレイテッド
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024220229
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-14
Filing Date
2024-12-16
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

The existing methods for selecting and optimizing machine learning models for natural language processing tasks are inefficient, requiring significant time, resources, and computational burden to train and validate multiple models.

Method used

The proposed method compares label-dependent spectra from the outputs of pre-trained models to identify suitable models for downstream tasks and reduces the complexity of pre-trained models using pruning procedures, thereby improving the efficiency of model selection and optimization.

Benefits of technology

This approach reduces the time and effort required to select and optimize machine learning models, improves computational efficiency, and enhances the performance of models for specific natural language processing tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025096265000001_ABST
    Figure 2025096265000001_ABST
Patent Text Reader

Abstract

To provide systems and methods for identifying a model to perform a task and systems and methods for updating architecture of a model to perform a task.SOLUTION: A method for identifying a model to perform a task includes: inputting information on samples in a plurality of samples to each model in a plurality of models, subsets of the samples corresponding to labels; acquiring spectra from outputs of layers in the models by applying parameters for the information, the spectra being dimension reduced to obtain component value sets that correspond to samples and collectively have an explained variance of at least a threshold amount of the total variance; and determining, for each model, a divergence using a mathematical combination of a plurality of distances, and identifying a model having a divergence satisfying a threshold to perform the task, each of the distances being between each label subset of samples relative to all other samples.SELECTED DRAWING: Figure 2A
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to systems and methods for selecting and optimizing machine learning models, and more particularly to systems and methods for use in natural language processing.

Background Art

[0002] The emergence of deep learning models (DLMs) has swept through the natural language processing (NLP) world. The paradigm of obtaining a DLM and then fine-tuning the DLM on a set of labeled data has become widespread in the machine learning context. As a result, many pre-trained and fine-tuned DLMs are generally available. This significantly reduces the time, resources, and data required for a research team to succeed in a task. However, the selection of a pre-trained or fine-tuned model for a downstream fine-tuning task can vary greatly. For this reason, the existence of heuristics for model selection can save a great deal of time and energy compared to the effort and time required to train several different models and select the most performant one.

Summary of the Invention

[0003] In view of the foregoing background, what is needed in the art is an improved method and system for selecting and optimizing models from various possible pre-trained models, particularly for use in natural language processing tasks. The present disclosure addresses these and other problems by comparing label-dependent spectra from the outputs of pre-trained models to identify those pre-trained models that are more suitable for downstream tasks of interest, and by reducing the size and complexity of pre-trained models to subsets thereof having a greater information capacity. The disclosed systems and methods improve the process of obtaining a machine learning model for performing a particular task by reducing the time, effort, and computational burden of training and validating multiple models, and by reducing the complexity of such models using pruning procedures.

[0004] Accordingly, one aspect of the present disclosure provides a method for identifying a model that performs a first category of tasks. In some embodiments, the method is executed on a computer system that includes one or more processors and memory. In some embodiments, the method includes, for each model of a plurality of models, each model of the plurality of models is pre-trained on each task other than the first category of tasks at least partially, and each model of the plurality of models includes a corresponding plurality of layers including a corresponding input layer, a corresponding output layer, and a corresponding plurality of hidden layers for each of the plurality of validation samples, inputting corresponding information into each model, and obtaining, through application of corresponding plurality of parameters of each model to the corresponding information, an output from each hidden layer within the corresponding plurality of hidden layers in the form of a corresponding spectrum including a corresponding plurality of values, the plurality of validation samples including a corresponding label subset of validation samples assigned each label for each of the plurality of labels, thereby obtaining, for each model of the plurality of models, for each of the plurality of validation samples in the plurality of validation samples, a corresponding plurality of spectra having a corresponding total variance across the corresponding plurality of values. In some embodiments, the plurality of validation samples includes a corresponding subset of validation samples for each of the plurality of labels within the plurality of labels.

[0005] In some embodiments, the method further includes, for each model of the plurality of models, performing dimensionality reduction on the corresponding plurality of spectra to obtain a corresponding plurality of sets of component values that collectively have an explained variance of at least a threshold amount of the total variance. In some embodiments, the corresponding plurality of sets of component values includes a set of component values corresponding to each of the plurality of validation samples in the plurality of validation samples.

[0006] In some embodiments, the method is to determine, for each of a plurality of models, a corresponding divergence using a mathematical combination of a corresponding plurality of distances, wherein each of the corresponding plurality of distances represents a respective label within the plurality of labels, and is determined to be between (i) a set of component values for a respective label subset of validation samples assigned each label, and (ii) a set of component values for all other samples within the plurality of samples. In some embodiments, the method further includes identifying a first model having a corresponding divergence that meets a threshold for performing a first task within the plurality of models.

[0007] In some embodiments, the first task includes determining a patient-drug relationship, determining a patient-biomarker association, or determining a disease state.

[0008] In some embodiments, each respective model among the plurality of models is selected from the group consisting of a language model, a transformer model, a large language model (LLM), an encoder, a decoder, an encoder-decoder hybrid model, a generative pre-trained transformer (GPT) model, and a bidirectional encoder representations from transformers (BERT) model. In some embodiments, each respective pre-trained model among the plurality of pre-trained models is selected from the group consisting of: BERT, BERT Base, BERT large, RoBERTa Base, BioBERT Base, RoBERTa BaseTwitter Sentiment Finetune, DeBERTa, ALBERT, RoBERTa, GPT-J, GPT-Neo, GPT-NeoX, Pythia, GPT-NeoX2.0, XLNet, LaMDA, PaLM, Gopher, Sparrow, Chinchilla, Minerva, Bard, GPT-1, GPT-2, GPT-3, CodeX, InstructGPT, ChatGPT, GPT-4, OPT, Galactica, LLaMA, BART, Flan-T5, Flan-UL2, T5, and / or any derivatives or combinations thereof.

[0009] In some embodiments, one or more of the plurality of models are pre-trained using a set of non-specific pre-training samples. In some embodiments, one or more of the plurality of models are pre-trained using a set of domain-specific pre-training samples. In some embodiments, the domain is associated with a first task. In some embodiments, one or more of the plurality of models within the plurality of models are fine-tuned for the first task.

[0010] In some embodiments, the plurality of models further includes an untrained model. In some embodiments, the plurality of models includes at least five models.

[0011] In some embodiments, for each model among the plurality of models, the corresponding plurality of parameters includes at least 1,000 parameters.

[0012] In some embodiments, each respective validation sample among the plurality of validation samples includes all or a portion of an electronic health record (EHR) or an electronic medical record (EMR). In some embodiments, the corresponding information for each respective validation sample among the plurality of validation samples includes one or more corresponding snippets (fragments). In some embodiments, the plurality of validation samples includes at least 100 validation samples.

[0013] In some embodiments, the dimensionality reduction is a principal component analysis algorithm, a random projection algorithm, an independent component analysis algorithm, or a feature selection method. In some embodiments, the dimensionality reduction is a principal component analysis (PCA) reduction, and the dimensionality reduction decomposes a plurality of spectra into respective subsets of principal components.

[0014] In some embodiments, the threshold amount of the total variance is at least 90%, at least 95%, or at least 99% of the total variance.

[0015] In some embodiments, the corresponding plurality of distances is determined in a pairwise fashion between (i) a set of component values for each respective label subset of the validation samples and (ii) a set of component values for the corresponding label subsets for the respective labels within the plurality of labels.

[0016] In some embodiments, the mathematical combination of the corresponding plurality of distances is the sum of the corresponding plurality of distances. In some embodiments, the corresponding divergence is selected from the group consisting of the total variation distance, the Hellinger distance, the Lévy-Prohorov metric, the Wasserstein metric, the Mahalanobis distance, the Amari distance, the Kullback-Leibler divergence, the Renyi divergence, the Jensen-Shannon divergence, the Bhattacharyya distance, the f-divergence, and the discriminability index. In some embodiments, the corresponding divergence is the Jensen-Shannon divergence.

[0017] In some embodiments, each model meets a threshold when it has the largest corresponding divergence among a plurality of models.

[0018] In some embodiments, identifying further includes selecting a subset of models within a plurality of models having the top N largest corresponding divergences. In some embodiments, N is a positive integer from 1 to 5.

[0019] In some embodiments, the method further includes retraining a first model to perform a first task. In some embodiments, retraining includes performing a training procedure using the first model on a plurality of training samples to perform the first task.

[0020] In some embodiments, the method further includes identifying a subset of layers within a plurality of layers of the first model and removing layers other than the subset of layers from the first model prior to retraining.

[0021] Another aspect of the present disclosure provides a method for updating a model architecture to perform a first category task. In some embodiments, the method is executed on a computer system including one or more processors and memory. In some embodiments, the method includes, for each respective validation sample among a plurality of validation samples, inputting corresponding information into the model and obtaining, as an output from each respective layer of a plurality of layers of the model, a corresponding spectrum including a corresponding plurality of values, thereby obtaining a plurality of spectra having a total variance, wherein the model is pre-trained on each task other than the first category task and each layer of the model includes a corresponding set of pre-trained weights.

[0022] In some embodiments, the method comprises performing dimensionality reduction on a plurality of spectra to obtain a plurality of sets of component values that collectively have at least a threshold amount of the explained variance of the total variance, wherein the plurality of sets of component values are obtained such that each of the plurality of sets of component values corresponds to a respective layer of a plurality of layers.

[0023] In some embodiments, the method further comprises determining a first layer within a plurality of layers associated with a set of component values of the set of component values having the highest dimension, and removing each layer within the plurality of layers downstream of the first layer, thereby further comprising updating the architecture of the model to perform a first task.

[0024] In some embodiments, the first task comprises determining a patient-drug relationship, determining a patient-biomarker relationship, or determining a disease state.

[0025] In some embodiments, the model is selected from the group consisting of a language model, a transformer model, a large language model (LLM), an encoder, a decoder, an encoder-decoder hybrid model, a generative pre-trained transformer (GPT) model, and a bidirectional encoder representation from transformers (BERT) model. In some embodiments, the model is selected from the group consisting of: BERT, BERT Base, BERT large, RoBERTa Base, BioBERT Base, RoBERTa BaseTwitter Sentiment Finetune, DeBERTa, ALBERT, RoBERTa, GPT-J, GPT-Neo, GPT-NeoX, Pythia, GPT-NeoX2.0, XLNet, LaMDA, PaLM, Gopher, Sparrow, Chinchilla, Minerva, Bard, GPT-1, GPT-2, GPT-3, CodeX, InstructGPT, ChatGPT, GPT-4, OPT, Galactica, LLaMA, BART, Flan-T5, Flan-UL2, T5, and / or any derivatives or combinations thereof.

[0026] In some embodiments, the model is pre-trained using a set of unspecific pre-training samples. In some embodiments, the model is pre-trained using a set of domain-specific pre-training samples. In some embodiments, the domain is associated with a first task. In some embodiments, the model is fine-tuned for the first task.

[0027] In some embodiments, the plurality of layers includes at least 5 layers, at least 10 layers, or at least 15 layers. In some embodiments, each respective layer of the plurality of layers includes at least 5, at least 10, or at least 15 nodes.

[0028] In some embodiments, the corresponding set of pre-trained weights includes at least 1000 weights. In some embodiments, the model is selected by a method for identifying a model for performing a first task, the method comprising: A) for each respective model of a plurality of models, each respective model of the plurality of models is pre-trained on each task other than the first category task, and each respective model includes a corresponding plurality of layers including a corresponding input layer, a corresponding output layer, and a corresponding plurality of hidden layers, inputting corresponding information into each respective model, and obtaining, through application of the corresponding plurality of parameters of each respective model to the corresponding information, an output from each respective hidden layer within the corresponding plurality of hidden layers in the form of a corresponding spectrum including a corresponding plurality of values, wherein the plurality of validation samples includes a corresponding label subset of validation samples assigned each respective label for each respective label of a plurality of labels, whereby for each respective model of the plurality of models, a corresponding plurality of spectra having corresponding total variances are obtained; B) for each respective model of the plurality of models, performing dimensionality reduction on the corresponding plurality of spectra to obtain a corresponding plurality of sets of component values collectively having at least a threshold amount of the explained variance of the total variance, wherein the corresponding plurality of sets of component values are obtained such that each respective set of component values corresponds to each respective validation sample among the plurality of validation samples; C) for each respective model of the plurality of models, determining a corresponding divergence using a mathematical combination of corresponding plurality of distances, wherein each respective distance of the corresponding plurality of distances represents each respective label within the plurality of labels, and is determined such that it is between (i) the set of component values for the respective label subset of validation samples assigned each respective label and (ii) the set of component values for all other samples within the plurality of samples; and D) identifying, within the plurality of models, a first model having a corresponding divergence that meets a threshold for performing the first task.

[0029] In some embodiments, each respective validation sample among the plurality of validation samples includes all or a portion of an electronic health record (EHR) or an electronic medical record (EMR). In some embodiments, the corresponding information for each respective validation sample among the plurality of validation samples includes one or more corresponding snippets (fragments). In some embodiments, the plurality of validation samples includes at least 100 validation samples.

[0030] In some embodiments, the dimensionality reduction is a principal component analysis algorithm, a random projection algorithm, an independent component analysis algorithm, or a feature selection method. In some embodiments, the dimensionality reduction is a principal component analysis (PCA) reduction, and the dimensionality reduction decomposes a plurality of spectra into respective subsets of principal components. In some embodiments, the threshold amount of total variance is at least 90%, at least 95%, or at least 99% of the total variance.

[0031] In some embodiments, the dimensions include a plurality of principal components determined using dimensionality reduction. In some embodiments, the plurality of principal components includes at least 10, at least 100, or at least 1000 principal components.

[0032] In some embodiments, the model further includes a task-dependent output layer downstream of the plurality of layers, and removing further includes removing the task-dependent output layer.

[0033] In some embodiments, the method further includes retraining the model to perform a first task. In some embodiments, the retraining includes performing a training procedure using a first model on a plurality of training samples to perform the first task.

[0034] Another aspect of the present disclosure provides a computer system. The computer system includes one or more processors and memory addressable by the one or more processors. The memory stores at least one program for execution by the one or more processors. The at least one program includes instructions for performing any of the methods described herein.

[0035] Another aspect of the present disclosure provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores instructions that, when executed by a computer system, cause the computer system to perform any of the methods described herein.

Brief Description of the Drawings

[0036] In the drawings, embodiments of the systems and methods of the present disclosure are illustrated by way of example. It should be clearly understood that the description and drawings are for purposes of illustration only and are not intended as a definition of the limits of the systems and methods of the present disclosure.

[0037]

Figure 1A

Figure 1B

Figure 1C

[0038]

Figure 2A

Figure 2B

Figure 2C

[0039]

Figure 3A

Figure 3B

Figure 3C

[0040]

Figure 4A

Figure 4B

[0041]

Figure 5

[0042]

Figure 6

[0043]

Figure 7

[0044] Like reference numerals refer to corresponding parts throughout several views of the drawings.

DETAILED DESCRIPTION OF THE INVENTION

[0045] Reference is now made in detail to the embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. It will be apparent, however, to one of ordinary skill in the art that the present disclosure may be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to obscure aspects of the embodiments.

[0046] The present disclosure provides a system and method for identifying a model for performing a task such as a classification or prediction task. For each model of a plurality of models, information regarding each verification sample of a plurality of verification samples is input, and a label is assigned to a label subset of the samples. For each model, an output in the form of a corresponding spectrum is obtained from a layer of the model by applying parameters to the information for each verification sample, and thus a plurality of spectra of the model are obtained. The spectra are dimensionally reduced to obtain a set of component values that collectively have at least a threshold amount of explained variance of the total variance, and the set of component values includes a set of component values corresponding to each verification sample among the plurality of verification samples. For each model, divergence is determined using a mathematical combination of a plurality of distances, each distance representing a label, between (i) the set of component values of each label subset to which each respective label is assigned, and (ii) the set of component values for all other samples among the plurality of samples. A model having a divergence that meets a threshold is identified for performing the task.

[0047] Systems and methods are also provided for updating the architecture of a model to perform a task. Each layer of the model includes a plurality of layers and a set of pre-trained weights. The model is input with information about each validation sample in a plurality of validation samples, and an output in the form of a corresponding spectrum including a corresponding plurality of values is obtained from each layer, thus obtaining a plurality of spectra of the model having a total variance. In some embodiments, the model is pre-trained on each task other than the first category task. The spectra are dimensionally reduced to obtain a set of component values that collectively have an explained variance of at least a threshold amount of the total variance, and the set of component values includes a set of component values corresponding to each layer of the plurality of layers. A first layer is determined within the plurality of layers associated with the set of component values having the highest dimension, and each layer downstream of the first layer is removed from the model, thereby updating the architecture of the model to perform the task.

[0048] As described above, many pre-trained and fine-tuned machine learning models are generally available. This availability significantly reduces the time, resources, and amount of data required for a research team to succeed in a particular task of interest (e.g., classification, prediction, etc.). However, the performance of a pre-trained or fine-tuned model can vary significantly depending on the downstream task being performed, and the choice of which pre-trained or fine-tuned model to use for a task can itself require a significant amount of time, resources, and data. For this reason, the existence of heuristics for model selection can save a significant amount of time and energy compared to training several models and selecting the most performant one.

[0049] In view of the above, what is needed in the art is an improved method and system for selecting and optimizing a model from among various possible pre-trained models, particularly for use in natural language processing tasks. The present disclosure addresses these and other problems by comparing label-dependent spectra from the outputs of pre-trained models to identify those pre-trained models that are more suitable for a downstream task of interest, and by reducing the size and complexity of pre-trained models to subsets thereof having a greater information capacity. The disclosed systems and methods improve the process of obtaining a machine learning model for performing a particular task by reducing the time, effort, and computational burden of training and validating multiple models, and by reducing the complexity of such models using pruning procedures.

[0050] Models that better fit the downstream task are generally better at separating data according to the labels of each respective data point. Typically, this can be seen by examining label-dependent statistics in the output of the task-dependent output head. If no downstream training has occurred, this can still be observed by examining the label-dependent spectrum of the data coming from the output of the pre-trained model. In some implementations, the metric for determining label-dependent spectrum separation is Jensen-Shannon (JS) divergence. Since the output spectrum is often multi-dimensional, the JS divergence can be calculated and summed along the dimensions of the spectrum. However, this can be a problem since high-dimensional outputs can have an advantage due to the large number of dimensions contributing to the sum. In such cases, naive JS divergence is not only advantageous for higher-dimensional outputs but also does not consider correlations within the output. To avoid this problem, in some embodiments, the spectrum is decomposed into principal components necessary to account for a contribution rate (e.g., PCA-reduced JS divergence) of a threshold (e.g., 99%).

[0051] Advantageously, as shown in Example 1 below, models with higher PCA-reduced JS divergence (e.g., pre-trained machine learning models) are well-correlated with better downstream classification performance, indicating that such a metric predicts better discrimination of label-dependent data.

[0052] Furthermore, as shown in Example 2 below, the model was found to have a larger information capacity in the intermediate layer. By measuring the dimensions of the PCA-reduced spectra obtained from the outputs of each layer within the model, it is possible to limit the complexity of the pre-trained model to a subset with maximum discriminative power. This advantageously improves the efficiency of training and model use for downstream tasks compared to using the full model by reducing the time, complexity, and resources required to train and execute the model.

[0053]

[0054] Definition. Also, terms such as first, second, etc. may be used herein to describe various elements, but it will be understood that these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the present disclosure, the first object may be referred to as the second object, and similarly, the second object may be referred to as the first object. The first object and the second object are both objects, but not the same object.

[0055] The terms used in this disclosure are for the purpose of describing particular embodiments only and are not intended to be limiting of the invention. As used in the description of the invention and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. As used herein, the term "and / or" refers to any and all possible combinations of one or more of the associated listed items and includes them. As used herein, the terms "comprises" and / or "comprising" identify the presence of the described features, integers, steps, operations, elements, and / or components, but it will be further understood that they do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0056] As used herein, the term "if" may be construed to mean "when" or "upon" or "in response to detecting" or "in response to determining", depending on the context. Similarly, the phrases "when determined" or "when [the described condition or event] is detected" may be construed to mean "upon determination" or "in response to determining" or "upon detection of [the described condition or event]" or "in response to detecting [the described condition or event]", depending on the context.

[0057] As used herein, the term "classification" refers to any number or other character associated with a particular characteristic of a sample or input (e.g., an electronic health record or a portion thereof). For example, in some embodiments, the term "classification" refers to an association with each of a plurality of relationship statuses (e.g., positive, negative, or null), such as an association between an electronic health record or a portion thereof (e.g., each text span of a plurality of text spans) and its respective relationship status. In some embodiments, the term "classification" refers to the relationship status of a subject having a medical entity. For example, in some implementations, a subject is determined to be associated (e.g., positively) or not associated (e.g., negatively) with a medical entity. A classification can be binary (e.g., positive or negative) or can have more levels of classification (e.g., on a scale of 1 - 10 or 0 - 1). The terms "cutoff" and "threshold" can refer to a predetermined number used in an operation. For example, a cutoff size can refer to the size above which a fragment is excluded. A threshold can be a value above or below which a particular classification is applied. Either of these terms can be used in either of these contexts.

[0058] As used interchangeably herein, the terms "classifier" or "model" refer to a machine learning model or algorithm.

[0059] In some embodiments, the model includes an unsupervised learning algorithm. An example of an unsupervised learning algorithm is cluster analysis. In some embodiments, the model includes supervised machine learning. Non-limiting examples of supervised learning algorithms include, but are not limited to, logistic regression, neural networks, support vector machines, naive Bayes algorithms, nearest neighbor algorithms, random forest algorithms, decision tree algorithms, boosted tree algorithms, multinomial logistic regression algorithms, linear models, linear regression, gradient boosting, mixture models, hidden Markov models, Gaussian NB algorithms, linear discriminant analysis, or any combination thereof. In some embodiments, the model is a multinomial classifier algorithm. In some embodiments, the model is a two-stage Stochastic Gradient Descent (SGD) model. In some embodiments, the model is a deep neural network (e.g., a deep wide sample level model).

[0060] Neural network. In some embodiments, the model is a neural network (e.g., a convolutional neural network and / or a residual neural network). Neural network algorithms, also known as artificial neural networks (ANNs), include convolutional and / or residual neural network algorithms (deep learning algorithms). In some embodiments, the neural network is a machine learning algorithm trained to map an input data set to an output data set, and the neural network includes an interconnected group of nodes organized into multiple layers of nodes. For example, in some embodiments, the neural network architecture includes at least an input layer, one or more hidden layers, and an output layer. In some embodiments, the neural network includes any total number of layers and any number of hidden layers, and the hidden layers function as trainable feature extractors that enable mapping a set of input data to an output value or a set of output values. In some embodiments, the deep learning algorithm includes a neural network with multiple hidden layers, e.g., two or more hidden layers. In some cases, each layer of the neural network includes a number of nodes (or "neurons"). In some embodiments, a node receives an input coming directly from either the input data or the output of the nodes in the previous layer and performs a specific operation, e.g., a summation operation. In some embodiments, the connections from the input to the nodes are associated with parameters (e.g., weights and / or weight coefficients). In some embodiments, the node is the input x iSum the products of all pairs and their associated parameters. In some embodiments, the weighted sum is offset by a bias b. In some embodiments, the output of a node or neuron is gated using a threshold or activation function f, which in some cases is a linear or non-linear function. In some embodiments, the activation function is, for example, a rectified linear unit (ReLU) activation function, a Leaky ReLU activation function, or a saturated hyperbolic tangent function, an identity function, a binary step function, a logistic function, an arctangent function, a soft sign function, a parametric rectified linear unit function, an exponential linear unit function, a softplus function, a bent identity function, a soft exponential function, a sinusoid function, a sine function, a Gaussian function, or a sigmoid function, or other functions such as any combination thereof.

[0061] In some implementations, the weighting coefficients, bias values, and thresholds, or other computational parameters of the neural network, are "taught" or "learned" in a training phase using one or more sets of training data. For example, in some implementations, the parameters are trained using the training data set and input data from gradient descent or backpropagation such that the output values calculated by the ANN match the examples included in the training data set. In some embodiments, the parameters are obtained from a backpropagation neural network training process.

[0062] Any of a variety of neural networks are suitable for use according to the present disclosure. Examples include, but are not limited to, feedforward neural networks, radial basis function networks, recurrent neural networks, residual neural networks, convolutional neural networks, residual convolutional neural networks, or any combination thereof. In some embodiments, machine learning utilizes a pre-trained and / or transferred ANN or deep learning architecture. In some implementations, convolutional and / or residual neural networks are used according to the present disclosure.

[0063] For example, a deep neural network model includes an input layer, a plurality of individually parameterized (e.g., weighted) convolutional layers, and an output score layer. Each parameter (e.g., weight) of the convolutional layers and the input layer contribute to a plurality of parameters (e.g., weights) associated with the deep neural network model. In some embodiments, at least 100 parameters, at least 1000 parameters, at least 2000 parameters, or at least 5000 parameters are associated with the deep neural network model. Thus, since the deep neural network model cannot be solved by the human mind, a computer needs to be used. In other words, considering the input to the model, the model output needs to be determined using a computer rather than the human mind in such embodiments. See, for example, Krizhevsky et al., 2012, "Imagenet classification with deep convolutional neural networks", Advances in Neural Information Processing Systems 2, Pereira, Burges, Bottou, Weinberger, eds., pp. 1097-1105, Curran Associates, Inc., Zeiler, 2012, "ADADELTA: an adaptive learning rate method", CoRR, vol. abs / 1212.5701, and Rumelhart et al., 1988, "Neurocomputing: Foundations of research", ch. Learning Representations by Back-propagating Errors, pp. 696-699, Cambridge, MA, USA: MIT Press, each of which is incorporated herein by reference.

[0064] Neural network algorithms that include convolutional neural network algorithms suitable for use as a model are disclosed, for example, in Vincent et al., 2010, "Stacked denoising autoencoder: Learning useful representations in a deep network with a local denoising criterion" J Mach Learn Res 11, pp. 3371-3408, Larochelle et al., 2009, "Exploring strategies for training deep neural networks" J Mach Learn Res 10, pp. 1-40, and Hassoun, 1995, Fundamentals of Artificial Neural Networks, Massachusetts Institute of Technology, each of which is incorporated herein by reference. Additional exemplary neural networks suitable for use as a model are disclosed in Duda et al., 2001, Pattern Classification, Second Edition, John Wiley & Sons, Inc., New York, and Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York, each of which is incorporated herein by reference in its entirety. Additional exemplary neural networks suitable for use as a model are also described in Draghici, 2003, Data Analysis Tools for DNA Microarrays, Chapman & Hall / CRC, and Mount, 2001, Bioinformatics: sequence and genome analysis, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York, each of which is incorporated herein by reference in its entirety.

[0065] Support vector machine. In some embodiments, the model is a support vector machine (SVM). SVM algorithms suitable for use as a model are described, for example, in Cristianini and Shawe-Taylor, 2000, "An Introduction to Support Vector Machines", Cambridge University Press, Cambridge; Boser et al., 1992, "A training algorithm for optimal margin classifiers", Proceedings of the 5th Annual ACM Workshop on Computational Learning Theory, ACM Press, Pittsburgh, Pa., pp. 142-152; Vapnik, 1998, Statistical Learning Theory, Wiley, New York; Mount, 2001, Bioinformatics: sequence and genome analysis, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y.; Duda, Pattern Classification, Second Edition, 2001, John Wiley & Sons, Inc., pp. 259, 262-265; Hastie, 2001, The Elements of Statistical Learning, Springer, New York; and Furey et al., 2000, Bioinformatics 16, 906-914, each of which is incorporated herein by reference in its entirety. When used for classification, the SVM separates a given set of binary-labeled data with a hyperplane that is maximally distant from the labeled data. In certain cases where linear separation is not possible, the SVM functions in combination with a "kernel" technique that automatically implements a non-linear mapping into the feature space. The hyperplane found by the SVM in the feature space corresponds, in some cases, to a non-linear decision boundary in the input space.In some embodiments, a plurality of parameters (e.g., weights) associated with an SVM define a hyperplane. In some embodiments, the hyperplane is defined by at least 10, at least 20, at least 50, or at least 100 parameters, and since the SVM model cannot be solved by a human mind, it needs to be computed by a computer.

[0066] Naive Bayes algorithm. In some embodiments, the model is a Naive Bayes algorithm. A Naive Bayes model suitable for use as a model is disclosed, for example, in Ng et al., 2002, "On discriminative vs. generative classifiers: A comparison of logistic regression and naive Bayes", Advances in Neural Information Processing Systems, 14, which is incorporated herein by reference. A Naive Bayes model is any model in the family of "probability models" based on applying Bayes' theorem using the assumption of strong (naive) independence between features. In some embodiments, they are combined with kernel density estimation. See, for example, Hastie et al., 2001, The elements of statistical learning: data mining, inference, and prediction, eds. Tibshirani and Friedman, Springer, New York, which is incorporated herein by reference.

[0067] Nearest neighbor algorithm. In some embodiments, the model is a nearest neighbor algorithm. In some implementations, the nearest neighbor model is memory-based and does not include a fitting model. For the nearest neighbor, considering a query point x0 (the test subject), the k training points x (r) , r,..., k (here the training subjects) are identified, and then the point x0 is classified using the k nearest neighbors. In some embodiments, the Euclidean distance in the feature space is [Number] It is used to determine distance as. Typically, when a nearest neighbor algorithm is used, the abundance data used to calculate linear discrimination is standardized to have a mean of zero and a variance of 1. In some embodiments, the nearest neighbor rule is refined to address issues of unequal prior distribution of classes, different misclassification costs, and feature selection. Many of these refinements involve some form of weighted voting for the neighborhood. For more information on nearest neighbor analysis, see Duda, Pattern Classification, Second Edition, 2001, John Wiley & Sons, Inc, and Hastie, 2001, The Elements of Statistical Learning, Springer, New York, each of which is incorporated herein by reference.

[0068] The k-nearest neighbor model is a non-parametric machine learning method where the input is composed of the k closest training examples within the feature space. The output is class membership. An object is classified by a plurality of votes from its neighborhood, and the object is assigned to the most common class among its k nearest neighbors (k is a positive integer and typically small). In the case of k = 1, the object is simply assigned to the class of its single nearest neighbor. See Duda et al., 2001, Pattern Classification, Second Edition, John Wiley & Sons, which is incorporated herein by reference. In some embodiments, the number of distance calculations required to solve the k-nearest neighbor model is such that a computer is used to solve the model for a given input because it cannot be performed in a human head.

[0069] Random forest, decision tree, and boosted tree algorithms. In some embodiments, the model is a decision tree. Decision trees suitable for use as a model are generally described in Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York, pp. 395-396, which is hereby incorporated by reference. Tree-based methods partition the feature space into a set of rectangles and then fit a model (such as a constant) to each. In some embodiments, the decision tree is a random forest regression. For example, one particular algorithm is Classification and Regression Trees (CART). Other particular decision tree algorithms include, but are not limited to, ID3, C4.5, MART, and random forest. CART, ID3, and C4.5 are described in Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York, pp. 396-408 and pp. 411-412, which is hereby incorporated by reference. CART, MART, and C4.5 are described in Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York, Chapter 9, which is hereby incorporated by reference. Random forest is described in Breiman, 1999, "Random Forests--Random Features" Technical Report 567, Statistics Department, U.C. Berkeley, September 1999, which is hereby incorporated by reference in its entirety. In some embodiments, the decision tree model includes at least 10, at least 20, at least 50, or at least 100 parameters (e.g., weights and / or decisions), which cannot be solved by the human mind and thus need to be computed by a computer.

[0070] Regression. In some embodiments, the model uses a regression algorithm. In some embodiments, the regression algorithm is any type of regression. For example, in some embodiments, the regression algorithm is logistic regression. In some embodiments, the regression algorithm is logistic regression with lasso, L2, or elastic net regularization. In some embodiments, the extracted features with corresponding regression coefficients that cannot meet the threshold are excluded (removed) from consideration. In some embodiments, the generalization of the logistic regression model for processing multi-category responses is used as the model. The logistic regression algorithm is disclosed in Agresti, An Introduction to Categorical Data Analysis, 1996, Chapter 5, pp. 103-144, John Wiley & Son, New York, which is incorporated herein by reference. In some embodiments, the model utilizes the regression model disclosed in Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York. In some embodiments, the logistic regression model includes at least 10, at least 20, at least 50, at least 100, or at least 1000 parameters (e.g., weights), and since it cannot be solved by the human mind, it needs to be computed by a computer.

[0071] Linear discriminant analysis algorithm. In some embodiments, linear discriminant analysis (LDA), normal discriminant analysis (NDA), or discriminant function analysis is a method used in statistics, pattern recognition, and machine learning to find a linear combination of features that characterize or separate two or more classes of objects or events, which is a generalization of Fisher's linear discriminant. In some embodiments, the resulting combination is used as a model (linear model) in some embodiments of the present disclosure.

[0072] Mixture models and hidden Markov models. In some embodiments, the model is a mixture model such as those described in McLachlan et al., Bioinformatics 18(3):413-422, 2002. In some embodiments, particularly those that include a temporal component, the model is a hidden Markov model such as those described in Schliep et al., 2003, Bioinformatics 19(1):i255-i263.

[0073] Clustering. In some embodiments, the model is an unsupervised clustering model. In some embodiments, the model is a supervised clustering model. Clustering algorithms suitable for use as a model are described, for example, on pages 211 - 256 of Duda and Hart, Pattern Classification and Scene Analysis, 1973, John Wiley & Sons, Inc., New York (hereinafter Duda 1973), which is hereby incorporated by reference in its entirety. As an illustrative example, in some embodiments, the clustering problem is described as the problem of finding natural groups within a dataset. To identify natural groups, two problems are addressed. First, a way to measure the similarity (or dissimilarity) between two samples is determined. This metric (e.g., similarity measure) is used to ensure that samples within one cluster are more similar to each other than samples within other clusters. Second, a mechanism for partitioning the data into clusters using the similarity measure is determined. One way to start a clustering investigation is to define a distance function and compute a matrix of distances between all pairs of samples in the training set. If distance is a good measure of similarity, the distance between reference entries within the same cluster is significantly smaller than the distance between reference entries in different clusters. However, in some implementations, clustering does not use a distance measure. For example, in some embodiments, a non - metric similarity function s(x, x’) is used to compare two vectors x and x’. In some such embodiments, s(x, x’) is a symmetric function that has a large value when x and x’ are “similar” in some way. Once a method for measuring “similarity” or “dissimilarity” between points in the dataset is selected, clustering uses a criterion function that measures the clustering quality of any partition of the data. The partition of the dataset that extracts the criterion function is used to cluster the data.Specific exemplary clustering techniques contemplated for use in the present disclosure include, but are not limited to, hierarchical clustering (agglomerative clustering using the nearest neighbor algorithm, the longest distance algorithm, the average linkage algorithm, the centroid algorithm, or the sum of squares algorithm), k-means clustering, fuzzy k-means clustering algorithms, and Jarvis-Patrick clustering. In some embodiments, clustering includes unsupervised clustering (e.g., the number of clusters is not predetermined and / or the cluster assignments are not pre-determined).

[0074] Ensembles of models and boosting. In some embodiments, an ensemble (two or more) of models is used. In some embodiments, boosting techniques such as AdaBoost are used in combination with many other types of learning algorithms to improve the performance of the model. In this approach, the output of any of the models disclosed herein, or their equivalents, are combined into a weighted sum representing the final output of the boosted model. In some embodiments, multiple outputs from the model are combined using any measure of central tendency known in the art, including, but not limited to, mean, median, mode, weighted mean, weighted median, weighted mode, etc. In some embodiments, multiple outputs are combined using a voting method. In some embodiments, each model within the ensemble of models is weighted or unweighted.

[0075] As used herein, the term "parameter" refers to any coefficient of an internal or external element (e.g., weight and / or hyperparameter) in an algorithm, model, regressor, and / or classifier that can affect (e.g., modify, adjust, and / or regulate) one or more inputs, outputs, and / or functions therein, or similarly any value. For example, in some embodiments, a parameter refers to any coefficient, weight, and / or hyperparameter that can be used to control, modify, adjust, and / or regulate the behavior, learning, and / or performance of an algorithm, model, regressor, and / or classifier. In some cases, a parameter is used to increase or decrease the effect of an input (e.g., a feature) on an algorithm, model, regressor, and / or classifier. By way of non-limiting example, in some embodiments, a parameter is used to increase or decrease the effect of a node (e.g., of a neural network), which includes one or more activation functions. The assignment of a parameter to a particular input, output, and / or function is not limited to any one paradigm for a given algorithm, model, regressor, and / or classifier, but can be used in any suitable algorithm, model, regressor, and / or classifier architecture for desired performance. In some embodiments, a parameter has a fixed value. In some embodiments, the value of a parameter is adjustable manually and / or automatically. In some embodiments, the value of a parameter is modified by a validation and / or training process for an algorithm, model, regressor, and / or classifier (e.g., by an error minimization and / or backpropagation method). In some embodiments, the algorithms, models, regressors, and / or classifiers of the present disclosure include multiple parameters.In some embodiments, the plurality of parameters are n parameters, where n ≥ 2, n ≥ 5, n ≥ 10, n ≥ 25, n ≥ 40, n ≥ 50, n ≥ 75, n ≥ 100, n ≥ 125, n ≥ 150, n ≥ 200, n ≥ 225, n ≥ 250, n ≥ 350, n ≥ 500, n ≥ 600, n ≥ 750, n ≥ 1,000, n ≥ 2,000, n ≥ 4,000, n ≥ 5,000, n ≥ 7,500, n ≥ 10,000, n ≥ 20,000, n ≥ 40,000, n ≥ 75,000, n ≥ 100,000, n ≥ 200,000, n ≥ 500,000, n ≥ 1 × 10. 6 , n ≥ 5 × 10 6 , or n ≥ 1 × 10 7 . Thus, the algorithms, models, regressors, and / or classifiers of the present disclosure cannot be implemented in a human mind. In some embodiments, n is from 10,000 to 1 × 10 7 , from 100,000 to 5 × 10 6 , or from 500,000 to 1 × 10 6 . In some embodiments, the algorithms, models, regressors, and / or classifiers of the present disclosure operate in a k-dimensional space, where k is a positive integer of 5 or more (e.g., 5, 6, 7, 8, 9, 10, etc.). Thus, the algorithms, models, regressors, and / or classifiers of the present disclosure cannot be implemented in a human mind.

[0076] As used herein, the term "untrained model" (e.g., "untrained classifier" and / or "untrained neural network") refers to a machine learning model or algorithm, such as a classifier or neural network, that has not been trained on a target dataset. In some embodiments, "training a model" (e.g., "training a neural network") refers to the process of training an untrained or partially trained model (e.g., "untrained or partially trained neural network"). Further, it will be understood that the term "untrained model" does not exclude the possibility that transfer learning techniques may be used in such training of an untrained or partially trained model. For example, Fernandes et al., 2017, "Transfer Learning with Partial Observability Applied to Cervical Cancer Screening" Pattern Recognition and Image Analysis:8, which is incorporated herein by reference thThe Iberian Conference Proceedings, 243 - 250 provide non - limiting examples of such transfer learning. In cases where transfer learning is used, in the untrained model described above, additional data beyond the data of the primary training dataset is provided. Typically, this additional data is in the form of parameters (e.g., coefficients, weights, and / or hyperparameters) learned from another auxiliary training dataset. Further, although the description of a single auxiliary training dataset is disclosed, it will be understood that there is no limit to the number of auxiliary training datasets that can be used to complement the primary training dataset when training the untrained model in the present disclosure. For example, in some embodiments, two or more auxiliary training datasets, three or more auxiliary training datasets, four or more auxiliary training datasets, or five or more auxiliary training datasets are used to complement the primary training dataset through transfer learning, and each such auxiliary dataset is different from the primary training dataset. In some such embodiments, any manner of transfer learning is used. For example, consider the case where, in addition to the primary training dataset, there is a first auxiliary training dataset and a second auxiliary training dataset. In such a case, the parameters learned from the first auxiliary training dataset (by application of a first model to the first auxiliary training dataset) are applied to the second auxiliary training dataset using transfer learning techniques (e.g., a second model that is the same as or different from the first model), resulting in a trained intermediate model, and then those parameters are applied to the primary training dataset, which, together with the primary training dataset itself, is applied to the untrained model.Alternatively, in another exemplary embodiment, a first set of parameters learned from a first auxiliary training dataset (by application of a first model to the first auxiliary training dataset) and a second set of parameters learned from a second auxiliary training dataset (by application of the same or a different second model from the first model to the second auxiliary training dataset) are each applied individually to separate instances of the primary training dataset (e.g., by separate independent matrix multiplications), and such both applications of the parameters to separate the instances of the primary training dataset, together with the primary training dataset itself (or some reduced form of the primary training dataset such as principal components or regression coefficients learned from the primary training set), are then applied to an untrained model to train the untrained model.

[0077] Exemplary system embodiment.

[0078] Figures 1A - C illustrate a computer system 100 for identifying a model for performing a first category task and / or updating an architecture of a model for performing a first category task, according to some embodiments of the present disclosure. In a typical embodiment, computer system 100 includes one or more computers. For the purposes of the figures of Figures 1A - C, computer system 100 is represented as a single computer including all of the functions of the disclosed computer system 100. However, the present disclosure is not so limited. The functions of computer system 100 can span any number of networked computers and / or can exist on each of several networked computers and / or virtual machines. Those skilled in the art will understand that a wide variety of different computer topologies are possible for computer system 100 and that all such topologies are within the scope of the present disclosure.

[0079] With the foregoing in mind, returning to FIGS. 1A - C, computer system 100 includes one or more processing units (CPUs) 59, a network or other communication interface 84, a user interface 78 (e.g., including an optional display 82 and an optional keyboard 80 or other form of input device), a memory 92 (e.g., random access memory, persistent memory, or a combination thereof), one or more magnetic disk storage and / or persistent devices 90 optionally accessed by one or more controllers 88, one or more communication buses 12 for interconnecting the foregoing components, and a power supply 79 for powering the foregoing components. To the extent that the components of memory 92 are not persistent, the data in memory 92 can be seamlessly shared with non - volatile memory 90 or a portion of memory 92 that is non - volatile or persistent, using known computing techniques such as caching. Memory 92 and / or memory 90 can include mass storage located remotely from central processing unit 59. In other words, some of the data stored in memory 92 and / or memory 90 can actually be external to computer system 100 but hosted on a computer that can be electronically accessed by computer system 100 via network interface 84 over the Internet, an intranet, or other form of network or electronic cable. In some embodiments, computer system 100 utilizes a model executed from memory associated with one or more graphical processing units to improve the speed and performance of the system. In some implementations, computer system 100 utilizes a model executed from memory 92 rather than memory associated with the graphical processing unit.

[0080] In some implementations, the memory 92 of the system 100 stores the following programs, modules, and data structures, or subsets thereof, to identify models that perform first-category tasks and / or update the architectures of models that perform first-category tasks. ● Any operating system 34 that includes procedures for handling various basic system services, and ● Any input / output module 64 for connecting the system 100 to other devices, and ● A sample data store 120 that optionally includes a plurality of validation samples 122 (e.g., 122-1,..., 122-K), and ● For each respective model 132 of a plurality of models (e.g., 132-1,..., 132-M), a model construct 130 that optionally includes a corresponding plurality of parameters 134 (e.g., 134-1-1,..., 134-1-P) and a plurality of layers 136 (e.g., 136-1-1,..., 136-1-L), and ● For each respective model 132 of a plurality of models, an output module 140 that optionally includes a corresponding plurality of spectra 142 (e.g., 142-1,..., 142-S) having total variance, obtained from the output of the model, and ● Optionally, a statistical module 150 that includes: o A plurality of component value sets 152 (e.g., 152-1,..., 152-V) that collectively have at least a threshold amount of explained variance based on dimensionality reduction performed on the plurality of spectra 142, and o For each respective model 132 of a plurality of models, a corresponding plurality of distances 154 (e.g., 154-1,..., 154-D), where each respective distance of the corresponding plurality of distances represents each label within a plurality of labels, and (i) the component value set 152 for each respective label subset of the validation samples 122 to which each label is assigned, and (ii) the component value set for all other samples within the plurality of samples, and o For each respective model 132 of a plurality of models, corresponding divergences 156 (e.g., 156-1, ..., 156-F) obtained using a mathematical combination of corresponding plural distances 154 for each respective model, and / or o Dimensions 158 (e.g., 158-1, ..., 158-G) determined for one or more layers 136 of each respective model 132, are included.

[0081] In some implementations, one or more of the above-described data elements or modules of the computer system 100 are stored in one or more of the previously described memory devices and associated with a series of instructions for performing the above-described functions. The modules, data, or programs (e.g., sets of instructions) identified above need not be implemented as separate software programs, procedures, data sets, or modules, and thus, various subsets of these modules and data may be combined or otherwise rearranged in various embodiments. In some implementations, memory 92 and / or 90 optionally stores a portion of the above-described modules and data structures. Further, in some embodiments, memory 92 and / or 90 stores additional modules and data structures not described above. Details of the modules and data structures identified above are described below with reference to FIGS. 2A-C and 3A-C.

[0082] Exemplary embodiments

[0083] Exemplary embodiments for identifying a model to perform a task

[0084] Figures 2A - C collectively show a flowchart of an exemplary method 200 for identifying a model 132 to perform a first task (e.g., a first category task) according to some embodiments of the present disclosure. The dashed boxes represent optional parts of the method. In some embodiments, this method is executed in a computer system including one or more processors and memory. In some embodiments, the method is implemented by modules of the computer system 100 as detailed elsewhere herein.

[0085] Referring to block 202, in some embodiments, the method inputs corresponding information to each respective model 132 of a plurality of models, where each respective model of the plurality of models is pre - trained at least partially on each task other than the first category task, and each respective model includes a corresponding plurality of layers including a corresponding input layer, a corresponding output layer, and a corresponding plurality of hidden layers, obtains the output from each respective hidden layer 136 within the corresponding plurality of hidden layers in the form of a corresponding plurality of spectra including a corresponding plurality of values through the application of the corresponding plurality of parameters 134 of the respective model to the corresponding information, where the plurality of validation samples include a corresponding label subset of the validation samples assigned to each respective label of a plurality of labels, thereby obtaining, for each respective model of the plurality of models, a corresponding plurality of spectra 142 having a corresponding total variance over the corresponding plurality of values for each respective validation sample in the plurality of validation samples. In some embodiments, the plurality of validation samples include a corresponding label subset of the validation samples for each respective label within the plurality of labels.

[0086] Model.

[0087] In some embodiments, the systems and methods of the present disclosure are implemented to identify a model that has the ability to perform, or is more suitably adapted to perform, a particular task (e.g., a first category task) compared to other models. In some implementations, the model is pre-trained. In some implementations, the model is pre-trained on training data specific to the domain of a particular task. In some implementations, the model is pre-trained to perform a particular task. In some implementations, the model is pre-trained on non-specific training data (e.g., not specific to the domain of the first category task). In some implementations, the model is pre-trained to perform tasks other than a particular task. In this way, even if the available models are not trained to perform a particular task, any number of available pre-trained models can be evaluated to determine which model has the ability to perform, or is more suitably adapted to perform, the particular task. In some embodiments, the task is a category task. For example, in some embodiments, a category task includes assigning a category to an input to the model or a sample thereof. In some embodiments, the category is selected from a set of predetermined categories (e.g., a set of disease types, a set of indications, etc.). In some embodiments, a category task includes outputting a prediction for an input to the model or a sample thereof. In some embodiments, the prediction is selected from a set of possible predictions (e.g., a set of disease types, an indicator in a set of binary indicators, etc.). In some embodiments, a category task includes outputting a characterization of each input to the model or a sample thereof. In some embodiments, the characterization is selected from a set of candidate characterizations (e.g., a set of symptoms, a set of disease types, a set of indications, etc.).

[0088] Referring to block 204, in some embodiments, the first category of tasks includes determining a patient-drug relationship, determining a patient-biomarker association, or determining a disease state. In some embodiments, the disease state includes a diagnosis, prognosis, symptoms, presence or absence of a disease, disease type (e.g., neoplastic disease, cardiovascular disease, endocrine disease, mental disease), disease subtype (e.g., cancer type, subtype, stage classification, and / or tissue of origin), and / or its probability, severity, or indication.

[0089] In some embodiments, the first category of tasks includes determining relationships, predictions, and / or displays in text (e.g., determining a patient-drug relationship in an electronic health record or electronic medical record). In some embodiments, the first category of tasks includes determining relationships, predictions, and / or displays in images (e.g., determining a diagnosis of a disease state in an image of a subject).

[0090] In some embodiments, each model within a plurality of models includes any of the model architectures disclosed herein (see, e.g., the section titled "Model" above). In some embodiments, each respective model of the plurality of models includes any of the model architectures disclosed herein.

[0091] Referring to block 206, in some embodiments, each model of the plurality of models is selected from the group consisting of a language model, a transformer model, a large language model (LLM), an encoder, a decoder, an encoder-decoder hybrid model, a generative pre-trained transformer (GPT) model, and a bidirectional encoder representation from transformers (BERT) model. In some embodiments, each model of the plurality of models is selected from the group consisting of: BERT, BERT Base, BERT large, RoBERTa Base, BioBERT Base, RoBERTa BaseTwitter Sentiment Finetune, DeBERTa, ALBERT, RoBERTa, GPT-J, GPT-Neo, GPT-NeoX, Pythia, GPT-NeoX2.0, XLNet, LaMDA, PaLM, Gopher, Sparrow, Chinchilla, Minerva, Bard, GPT-1, GPT-2, GPT-3, CodeX, InstructGPT, ChatGPT, GPT-4, OPT, Galactica, LLaMA, BART, Flan-T5, Flan-UL2, T5, and / or any derivatives or combinations thereof.

[0092] In some embodiments, the model is an "encoder-style" LLM or a "decoder-style" LLM. Encoder-style and decoder-style model architectures use self-attention layers to encode inputs such as word tokens or snippets. The encoder is designed to learn embeddings that can be used for prediction modeling tasks such as classification, while the decoder is designed to generate new outputs such as new text (e.g., in response to a text query).

[0093] In some embodiments, the transformer model utilizes a multi - head self - attention mechanism. Attention is a learned weighted sum of a set of inputs, and this set can be of any size. Assume that a machine - learning pipeline contains at a certain point a 3D tensor of shape (N, sequence_length, dim_size), and for each data point, there is a collection of sequence_length vectors of length dim_size. These vectors can be anything from token embeddings to hidden states along a recurrent neural network (RNN). The ordering of these vectors is not important, but it is possible to embed that information through positional embeddings. The purpose of attention is to encode the original (N, sequence_length, dim_size) - shaped input into a weighted sum along sequence_length and fold it down to (N, dim_size) where each data point is represented by a single vector. This output can be directly useful as an input to another layer or as an input to a logistic head.

[0094] In some embodiments, rather than taking a naive sum, the attention layer is trained to direct attention to specific inputs when generating this sum. Key in the most important inputs and weight them more. In some implementations, this is done across multiple attention heads, i.e., across co - attention layers that read across the same input and then aggregated into a final summary. A single attention head can be thought of as a retrieval system with a set of keys, queries, and values. The attention mechanism learns to map a query (Q) against a set of keys (K) to obtain the most relevant input values (V). The attention mechanism achieves this by calculating a sum weighted proportionally to the perceived importance of each input (i.e., attention weights). This weighting is performed across all attention heads and then further summarized downstream into a single weighted representation.

[0095] When the attention mechanism is a multi-head attention mechanism, in some embodiments, each snippet or its encoded representation is input to a different attention head. By having multiple heads, the attention mechanism can have more degrees of freedom in attempting to aggregate information. Each individual head can focus on different modes when aggregating. Overall, the heads should converge to the underlying distribution. Thus, the multiple heads help enable the model to focus on different concepts. Examples of attention mechanisms are described in Chaudhari et al., July 12, 2021, "An Attentive Survey of Attention Models", arXiv:1904-02874v3, and Vaswani et al., "Attention is All You Need", 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, California, USA, each of which is incorporated herein by reference. Additional non-limiting models contemplated for use in the present disclosure are, for example, described in Raschka, June 17, 2023, "Understanding Encoder and Decoder LLMs" (available at magazine.sebastianraschka.com / p / understanding-encoder-and-decoder on the Internet), the entirety of which is incorporated herein by reference.

[0096] As will be apparent to those skilled in the art, other public and / or commercially available models suitable for evaluation using the present system and method are contemplated.

[0097] In some embodiments, one or more of the plurality of models are pre-trained using a set of unspecific pre-training samples. As described above, in some implementations, each model within the plurality of models is trained on general domain data. In some implementations, the model is trained on data encompassing a plurality of different domains. In some implementations, the plurality of different domains include the domain of a particular task of interest (e.g., a first category task). In some implementations, the model is trained on data that does not include data related to the domain of the first task.

[0098] In some embodiments, one or more of the plurality of models are pre-trained using a set of domain-specific pre-training samples. In some embodiments, the domain is associated with the first task. In some such implementations, each model within the plurality of models is trained on data associated with the domain of the first category task. For example, in some implementations, the first task is associated with the biomedical domain (e.g., determining a patient-drug relationship), and each model is trained on a corpus of biomedical text (e.g., BioBERT).

[0099] In some embodiments, the domain is not associated with the first task. In some such implementations, each model is trained on data specific to a domain other than the domain of the first task. As an example, in some implementations, the first task is associated with the biomedical domain (e.g., determining a patient-drug relationship), and each model is trained on sentiment (e.g., determining positive, negative, or neutral annotations in text).

[0100] In some embodiments, the domain is the biomedical domain and / or the clinical domain.

[0101] In some embodiments, one or more of the plurality of models are fine-tuned for a task. In some embodiments, one or more of the plurality of models are fine-tuned for a first task. In some embodiments, the fine-tuning is for a task other than the first task. Fine-tuning generally involves updating all or a portion of the model's parameters (e.g., weights) to modify or update the task performed by each model or to modify or update the domain in which each model operates.

[0102] In some embodiments, one or more of the models are pre-trained using a sample type different from the sample type of the plurality of validation samples. For example, in some implementations, each model is pre-trained on its images and / or snippets, and the plurality of validation samples include its text and / or snippets. In some implementations, each model is pre-trained on text and / or text snippets, and the plurality of validation samples include images and / or image snippets. Alternatively or additionally, in some embodiments, one or more of the models are pre-trained using training data of the same type or state as the plurality of validation samples. For example, in some implementations, each model is pre-trained on a corpus of biomedical text, and the plurality of validation samples include text snippets from electronic health records (EHRs) or electronic medical records (EMRs).

[0103] In some embodiments, the plurality of models further include an untrained model (e.g., BERT Base untrained).

[0104] As described above, as will be apparent to those of ordinary skill in the art, any untrained, partially trained, or pre-trained public and / or commercially available model is contemplated for evaluation using the present system and method.

[0105] Referring to block 208, in some embodiments, the plurality of models include at least five models.

[0106] In some embodiments, the plurality of models includes at least 2, at least 5, at least 10, at least 20, at least 50, at least 100, at least 500, or at least 1000 models. In some embodiments, the plurality of models consists of 5000 or fewer, 1000 or fewer, 500 or fewer, 100 or fewer, 50 or fewer, or 10 or fewer models. In some embodiments, the plurality of models consists of 2 - 20, 5 - 100, 50 - 300, 200 - 1000, or 800 - 5000 models. In some embodiments, the plurality of models falls within another range that starts with 2 or more models and ends with 5000 or fewer models.

[0107] In some embodiments, each model within the plurality of models includes a corresponding plurality of parameters. Parameters suitable for use in this disclosure are further described elsewhere in this specification (see, for example, the above - defined section entitled "Parameters"). In some embodiments, the corresponding plurality of parameters includes a plurality of weights for each model.

[0108] In some embodiments, the plurality of parameters include at least 10, at least 50, at least 100, at least 500, at least 1000, at least 2000, at least 5000, at least 10,000, at least 20,000, at least 50,000, at least 100,000, at least 200,000, at least 500,000, at least 1 million, at least 2 million, at least 3 million, at least 4 million, at least 5 million, at least 10 million, at least 100 million, at least 1 billion, at least 10 billion, at least 100 billion, at least 1000 billion, or at least 1 trillion parameters. In some embodiments, the plurality of parameters include parameters of 10 trillion or less, 1 trillion or less, 10 billion or less, 1 billion or less, 10 million or less, 5 million or less, 4 million or less, 1 million or less, 500,000 or less, 100,000 or less, 50,000 or less, 10,000 or less, 5000 or less, 1000 or less, or 500 or less. In some embodiments, the plurality of parameters consist of parameters in the range of 10 to 5000, 500 to 10,000, 10,000 to 500,000, 20,000 to 1 million, 1 million to 10 billion, 10 billion to 1000 billion, or 100 billion to 10 trillion. In some embodiments, the plurality of parameters fall within another range that starts with 10 parameters or more and ends with 10 trillion parameters or less.

[0109] In some embodiments, for each model within the plurality of models, the corresponding plurality of weights include at least 1000 weights.

[0110] In some embodiments, the plurality of weights includes weights of at least 10, at least 50, at least 100, at least 500, at least 1000, at least 2000, at least 5000, at least 10,000, at least 20,000, at least 50,000, at least 100,000, at least 200,000, at least 500,000, at least 1 million, at least 2 million, at least 3 million, at least 4 million, at least 5 million, at least 10 million, at least 100 million, at least 1 billion, at least 10 billion, at least 100 billion, at least 1000 billion, or at least 1 trillion. In some embodiments, the plurality of weights includes weights in a number of 10 trillion or less, 1 trillion or less, 10 billion or less, 1 billion or less, 10 million or less, 5 million or less, 4 million or less, 1 million or less, 500,000 or less, 100,000 or less, 50,000 or less, 10,000 or less, 5000 or less, 1000 or less, 500 or less. In some embodiments, the plurality of weights consists of weights in the range of 10 to 5000, 500 to 10,000, 10,000 to 500,000, 20,000 to 1 million, 1 million to 10 billion, 10 billion to 1000 billion, or 100 billion to 10 trillion. In some embodiments, the plurality of weights begins with more than 10 weights and ends with 10 trillion weights or less and falls within another range.

[0111] Verification sample.

[0112] In some embodiments, each respective verification sample among the plurality of verification samples includes all or part of an electronic health record (EHR) or an electronic medical record (EMR).

[0113] To generate an electronic medical record (EMR), electronic health records (EHRs) or handwritten records, which are later digitized, include patient records that contain interactions between a patient and a healthcare provider. In some implementations, the EHRs and EMRs are stored in an electronic medical system selected for a healthcare provider. These EHRs and EMRs typically have structured data that includes medical codes used by the healthcare provider for billing purposes, and unstructured data that includes clinical notes and observations made by physicians, physician assistants, nurses, and others during patient examinations. EHRs and EMRs theoretically hold vast amounts of clinical data that can be utilized for the greater good of public health. Advantageously, such rich clinical data can be used to generate models for predicting disease risk, predicting treatment outcomes, recommending personalized therapies, predicting disease-free survival periods after treatment, predicting disease recurrence, and so on. In some embodiments, the plurality of validation samples includes clinical notes.

[0114] In some embodiments, each respective validation sample of the plurality of validation samples wholly includes an EHR or an EMR. In some embodiments, each respective validation sample of the plurality of validation samples includes a portion of an EHR or an EMR.

[0115] As will be apparent to those skilled in the art, other sample types are contemplated for use in the present disclosure that are appropriate for a particular task. In some implementations, the plurality of validation samples include text. In some implementations, the plurality of validation samples include images. In some implementations, each validation sample among the plurality of validation samples is in the form of a tensor or other representation. In some implementations, each validation sample among the plurality of validation samples is embedded, encoded, scaled, and / or transformed before being input to the model. In some embodiments, each validation sample among the plurality of validation samples is segmented or split (e.g., into patches). The segmented input is further described below. For example, as shown in FIG. 4A, the input to the model can be obtained from text or an image, where the text and / or image is flattened, split into patches, and embedded before being input to the model.

[0116] In some embodiments, the corresponding information for each respective validation sample among the plurality of validation samples includes one or more corresponding snippets (fragments).

[0117] For example, in some implementations, the input is too large to be supplied to the model as an input. Thus, in some embodiments, the method further includes segmenting or splitting the input into a plurality of snippets, where each snippet corresponds to a portion of the input (e.g., a short snippet of text and / or an image patch). In some implementations, the snippets are equal or approximately equal in size, shape, and / or length. In some implementations, the first snippet and the second snippet within the plurality of snippets have different sizes, shapes, and / or lengths. In some embodiments, one or more snippets are ranked, padded, and / or trimmed (e.g., ranking text according to some medically relevant words in each snippet). In some embodiments, the plurality of snippets per input is limited to a corresponding number of snippets and / or the input portion per snippet (e.g., 512 snippets with a maximum size of 256 words, for a total of 131,072 words).

[0118] In some embodiments, each snippet is a portion of a document or image that is smaller than the whole. In some embodiments, each snippet is a portion that surrounds an instance of a reference or corresponding surface form (e.g., a defined number of words or characters before and / or after an instance of the reference or corresponding surface form). For example, if the reference includes the term "PARP inhibitor" and each document includes the sentence "PARP inhibitors can be used in the treatment of breast cancer and ovarian cancer", the system extracts 100 words before and after the term "PARP inhibitor" to generate a single snippet.

[0119] In some embodiments, raw text is split using regular expression filtering to obtain snippets. An example of a regular expression syntax that can be used to split raw text into sentences is "r'\s{2,}|(?<!\w\.\w.)(?<![A-Z][a-z]\.)(?<=\.|\?)\s'". In some embodiments, certain punctuation marks are not identified as snippet boundaries. For example, the period at the end of the abbreviation "Dr." for a doctor can be excluded (e.g., "dr.XX"). An example of a regular expression syntax that helps to not identify certain punctuation marks as snippet boundaries is described, for example, in Section 3.2.2 of Rokach et al., Information Retrieval Journal, 11(6):499-538 (2008), the content of which is hereby incorporated by reference in its entirety for all purposes. In some embodiments, a machine learning model is used to split the input into snippets. Natural language processing (NLP) libraries for generating snippets (e.g., sentences), including Google SyntaxNet, Stanford CoreNLP, NLTK Phyton library, and spaCy, are known in the art as described by Haris et al., Journal of Information Technology and Computer Science, 5(3):279-92, which is hereby incorporated by reference in its entirety for all purposes.

[0120] In some embodiments, a plurality of validation samples collectively represent a plurality of labels. In some such embodiments, each respective validation sample among the plurality of validation samples includes a respective label. In some embodiments, a plurality of validation samples includes a corresponding label subset of the validation samples for each respective label within the plurality of labels.

[0121] For example, in some embodiments, each validation sample among a plurality of validation samples includes a label indicating the presence or absence of a disease state. Thus, in some such embodiments, a first label subset of the validation samples among the plurality of validation samples includes those validation samples labeled "present," and a second label subset of the validation samples among the plurality of validation samples includes those validation samples labeled "absent."

[0122] In some implementations, the label for a validation sample is task-dependent, as would be apparent to one of ordinary skill in the art. For example, if the first category task is to identify a patient-drug relationship, the plurality of labels includes corresponding labels indicating the association between each respective validation sample among the plurality of validation samples and the patient-drug relationship. In some implementations, if the first category task is to determine a disease state, the plurality of labels includes corresponding labels indicating the presence (e.g., positive) or absence (e.g., negative) of the disease state for each respective validation sample among the plurality of validation samples. In some embodiments, if the first category task is a classification task, the plurality of labels includes one or more classes (e.g., for skin lesion classification as described below in Example 2, the plurality of labels includes keratosis, benign keratosis-like lesions, basal cell carcinoma, dermatofibroma, vascular lesions, melanoma, and / or nevus).

[0123] In some embodiments, the plurality of labels includes at least 2, at least 3, at least 5, at least 10, at least 50, at least 100, at least 200, or at least 300 labels. In some embodiments, the plurality of labels is composed of 500 or fewer, 300 or fewer, 100 or fewer, 50 or fewer, or 10 or fewer labels. In some embodiments, the plurality of labels is composed of 2 - 10, 5 - 30, 20 - 100, 80 - 300, or 200 - 500 labels. In some embodiments, the plurality of labels falls within another range that starts with 2 or more labels and ends with 500 or fewer labels.

[0124] In some embodiments, the plurality of verification samples includes at least 100 verification samples.

[0125] In some embodiments, the plurality of verification samples includes at least 10, at least 50, at least 100, at least 500, at least 1000, at least 2000, at least 5000, at least 10,000, at least 20,000, at least 50,000, at least 100,000, at least 200,000, at least 500,000, at least 1 million, at least 2 million, at least 3 million, at least 4 million, or at least 5 million verification samples. In some embodiments, the plurality of verification samples includes 10 million or fewer, 5 million or fewer, 4 million or fewer, 1 million or fewer, 500,000 or fewer, 100,000 or fewer, 50,000 or fewer, 10,000 or fewer, 5000 or fewer, 1000 or fewer, or 100 or fewer verification samples. In some embodiments, the plurality of verification samples consists of 10 - 5000, 500 - 10,000, 10,000 - 500,000, 20,000 - 1 million, 1 million - 5 million, or 2 million - 10 million verification samples. In some embodiments, the plurality of verification samples begins with 10 or more verification samples and ends with 10 million or fewer verification samples and falls within another range.

[0126] Obtain the output.

[0127] Referring to block 210, in some embodiments, for each respective model of the plurality of models, the output is obtained from each respective hidden layer within the plurality of hidden layers of each respective model.

[0128] Hidden layers and nodes (e.g., neurons) suitable for use in the present disclosure are described in further detail elsewhere in this specification (see, e.g., the section entitled "Definitions: Neural Networks" above). In some embodiments, each respective model includes a plurality of hidden layers and, as input, takes the output of the final hidden layer and an output layer (e.g., a classifier layer) that generates a task-dependent output (e.g., classification). For example, FIGS. 4A-B show exemplary schematic diagrams of a model that includes a plurality of hidden layers followed by an output layer (e.g., a classifier), and each respective hidden layer includes a plurality of nodes.

[0129] As described above, in some embodiments, a model includes a group of interconnected nodes organized into multiple layers of nodes. For example, FIG. 7 shows an exemplary schematic diagram of a fully connected neural network where the model includes at least an input layer 702, one or more hidden layers 704, and an output layer 706, and the hidden layer refers to the layer between the input layer and the output layer. In some embodiments, the model includes any total number of layers and any number of hidden layers, and the hidden layers function as trainable feature extractors that enable mapping a set of input data to an output value or a set of output values. In some cases, each layer of the neural network includes a number of nodes (or "neurons"). In some embodiments, a node receives an input directly from either the input data or the output of a node in a previous layer and performs a particular operation, e.g., a summation operation. For example, referring to the exemplary model of FIG. 7, each respective layer (e.g., layer 702, 704, or 706) includes one or more nodes 708 (e.g., 708-1-A, 708-1-B, 708-1-C, 708-2-A, 708-2-B, 708-2-C, 708-2-D, 708-3). Nodes within the hidden layer 704 may be referred to as hidden nodes or hidden neurons. In some embodiments, the connections from the input to the nodes are associated with parameters (e.g., weights and / or weight coefficients). In some embodiments, a node has an input x iSum all of the pairwise products and their associated parameters. In a fully connected neural network as shown in Figure 7, all possible connections between two layers exist, and as a result, all inputs from the previous layer affect all outputs of the subsequent layer. Each input from the previous layer to a particular node is weighted according to the parameter associated with that node, and thus, at node 708-2-A, the weight of this node only affects the output generated for this node and does not affect the outputs of nodes 708-2-B or 708-2-C.

[0130] In some embodiments, the inputs to the model and / or its respective nodes are in the form of embeddings. Generally, an embedding refers to a representation (e.g., in tensor form) of an object such as a sequence (e.g., text, snippet, image, and / or patch). In some embodiments, an embedding is obtained by mapping a discrete or categorical variable to a vector of continuous values. In some implementations, an embedding captures the semantic relationships or context between elements of the representation (e.g., snippets between a series of text or image patches). For example, Figure 4A illustrates an input where the identity and position of each of a plurality of image patches are embedded as a tensor of values. Embeddings can also be used to represent the semantic context in text inputs such as "start of sentence", "end of sentence", and various text snippets (e.g., "the", "dog", "is", etc.). Methods, models, and algorithms suitable for embeddings for use in the present disclosure include, but are not limited to, examples such as principal component analysis (PCA), singular value decomposition (SVD), Word2Vec, Sequence2Vec, Gene2Vec, kmer2vec, seq2seq, and / or BERT, which are known in the art. See, for example, Mokhtarani, "Embeddings in Machine Learning: Everything You Need to Know" (2021), available at featureform.com / post / the-definitive-guide-to-embeddings on the Internet.

[0131] In some embodiments, the spectral-form output comprises a plurality of values. In some embodiments, each value of the plurality of values is an embedding. For example, as described above, in some embodiments, the model outputs a plurality of values generated by performing operations on input data from one or more output nodes in the output layer and / or from one or more hidden nodes in each hidden layer to one or more output nodes or hidden nodes. In other words, in some implementations, the spectrum is a set of embedded values output from the nodes of a particular layer of the model. In some embodiments, the spectral-form output comprises a plurality of values in tensor or vector form.

[0132] In some embodiments, the plurality of values consists of at least 2, at least 5, at least 10, at least 20, at least 50, at least 100, at least 500, at least 1000, at least 10,000, or at least 100,000 values. In some embodiments, the plurality of values comprises 1 million or fewer, 100,000 or fewer, 10,000 or fewer, 1000 or fewer, 500 or fewer, 100 or fewer, 50 or fewer, or 10 or fewer values. In some embodiments, the plurality of values consists of 2 to 20, 5 to 100, 50 to 300, 200 to 500, 300 to 1000, 1000 to 10,000, or 10,000 to 1 million values. In some embodiments, the plurality of values begins with a value of 2 or more and ends with a value of 1 million or less and falls within another range.

[0133] Referring back to block 202, in some embodiments, the output in the form of a spectrum is obtained from the last hidden layer in a plurality of layers. For example, FIG. 4A shows measurements obtained at the last hidden layer in a plurality of N layers (e.g., before the classifier or head layer) (shown by the dashed circle). Similarly, FIG. 4B shows measurements obtained using the output from the last hidden layer within a plurality of hidden layers before the classifier layer (the measurements are shown by the dashed circle). In some embodiments, the output is obtained from any hidden layer within a plurality of layers. In some embodiments, the output is obtained from the same hidden layer for each respective model of a plurality of models (e.g., the first hidden layer of each model, the second hidden layer of each model, the second-to-last hidden layer of each model, the last hidden layer of each model, etc.). In some embodiments, for the first model within a plurality of models, the output is obtained from a different hidden layer than for the second model within the plurality of models (e.g., the last hidden layer for the first model and the second-to-last hidden layer for the second model).

[0134] In some embodiments, each respective model within a plurality of models includes at least 2 layers, at least 5 layers, at least 10 layers, at least 20 layers, at least 50 layers, at least 100 layers, or at least 500 layers. In some embodiments, the plurality of hidden layers is composed of 1000 or fewer, 500 or fewer, 100 or fewer, 50 or fewer, or 10 or fewer layers. In some embodiments, the plurality of hidden layers consists of 2 to 20 layers, 5 to 100 layers, 50 to 300 layers, 200 to 500 layers, or 300 to 1000 layers. In some embodiments, the plurality of hidden layers begins with 2 or fewer layers and ends within another range that is 1000 or fewer layers.

[0135] Referring to block 212, in some embodiments, for each of the plurality of models, for each of the plurality of hidden layers within the plurality of models, the output is obtained from each of the nodes within the plurality of nodes for each of the hidden layers. In some embodiments, each of the plurality of hidden layers within the plurality of hidden layers includes a plurality of nodes. In some embodiments, the output is obtained from any of the nodes within the plurality of nodes. In some embodiments, the output is obtained from the same node for each of the plurality of models (e.g., the first node of the selected layer of each model, the second node of the selected layer of each model, the second-to-last node of the selected layer of each model, the last node of the selected layer of each model, etc.). In some embodiments, for the first model of the plurality of models, the output is obtained from different nodes of the selected hidden layer for the second model of the plurality of models (e.g., the last node of the selected layer for the first model, the second-to-last node of the selected layer for the second model).

[0136] In some embodiments, each of the plurality of hidden layers within the plurality of hidden layers includes at least 2, at least 5, at least 10, at least 20, at least 50, at least 100, at least 500, at least 1000, at least 10,000, or at least 100,000 nodes. In some embodiments, the plurality of nodes within each hidden layer includes 1 million or fewer, 100,000 or fewer, 10,000 or fewer, 1000 or fewer, 500 or fewer, 100 or fewer, 50 or fewer, or 10 or fewer nodes. In some embodiments, the plurality of nodes within each hidden layer consists of 2 to 20, 5 to 100, 50 to 300, 200 to 500, 300 to 1000, 1000 to 10,000, or 10,000 to 1 million nodes. In some embodiments, the plurality of nodes within each hidden layer falls within another range that starts with 2 nodes or more and ends with 1 million nodes or fewer.

[0137] In some embodiments, for each respective verification sample among a plurality of verification samples, the corresponding spectrum includes a plurality of dimensions (e.g., the spectrum is multi-dimensional).

[0138] In some embodiments, for each respective verification sample among a plurality of verification samples, the corresponding spectrum includes at least 2, at least 5, at least 10, at least 20, at least 50, at least 100, at least 500, at least 1000, or at least 5000 dimensions. In some embodiments, the corresponding spectrum is composed of dimensions of 10,000 or less, 5000 or less, 1000 or less, 500 or less, 100 or less, 20 or less, or 10 or less. In some embodiments, the corresponding spectrum consists of dimensions of 2 to 20, 10 to 100, 80 to 500, 300 to 2000, or 1000 to 10,000. In some embodiments, the corresponding spectrum falls within another range that starts with 2 or more dimensions and ends with 10,000 dimensions or less.

[0139] In some embodiments, the corresponding spectrum includes a plurality of dimensions, and each respective label among a plurality of labels is represented by each respective dimension of the plurality of dimensions.

[0140] In some embodiments, for each respective verification sample among a plurality of verification samples, the corresponding spectrum includes the distribution of the corresponding probabilities for each respective verification sample across a plurality of labels. For example, for each respective text snippet among a plurality of text snippets, the output from the model may include an indication or probability that the text snippet has or does not have a patient-drug relationship. In some embodiments, the indication is the distribution of probabilities assigned to each respective text snippet, including the probability that the text snippet includes a patient-drug relationship and the probability that the text snippet does not include a patient-drug relationship.

[0141] In some embodiments, each respective dimension of the plurality of dimensions does not represent a label among the plurality of labels.

[0142] Dimensionality reduction.

[0143] Referring to block 213, in some embodiments, the method further includes performing dimensionality reduction on each of a plurality of spectra 142 corresponding to each of a plurality of models 132 to obtain a corresponding plurality of component value sets 152 having an explained variance of at least a threshold amount of the total variance. In some embodiments, the corresponding plurality of component value sets 152 include component value sets corresponding to each of the plurality of validation samples 122 in the plurality of validation samples.

[0144] In some embodiments, any one or more of a variety of dimensionality reduction techniques are used. Examples include, but are not limited to, principal component analysis (PCA), non-negative matrix factorization (NMF), linear discriminant analysis (LDA), diffusion maps, or network (e.g., neural network) techniques such as autoencoders.

[0145] Referring to block 214, in some embodiments, the dimensionality reduction is a principal component analysis algorithm, a random projection algorithm, an independent component analysis algorithm, or a feature selection method.

[0146] In some embodiments, the dimensionality reduction is a principal component algorithm, a random projection algorithm, an independent component analysis algorithm, a feature selection method, a factor analysis algorithm, Sammon mapping, curve component analysis, a stochastic neighbor embedding (SNE) algorithm, an Isomap algorithm, a maximum variance unfolding algorithm, a locally linear embedding algorithm, a t-SNE algorithm, a non-negative matrix factorization algorithm, a kernel principal component analysis algorithm, a graph-based kernel principal component analysis algorithm, a linear discriminant analysis algorithm, a generalized discriminant analysis algorithm, a uniform manifold approximation and projection (UMAP) algorithm, a LargeVis algorithm, a Laplacian eigenmap algorithm, or a Fisher's linear discriminant analysis algorithm. See, for example, Fodor, 2002, "A survey of dimension reduction techniques" Center for Applied Scientific Computing, Lawrence Livermore National, Technical Report UCRL-ID-148494, Cunningham, 2007, "Dimension Reduction" University College Dublin, Technical Report UCD-CSI-2007-7, Zahorian et al., 2011, "Nonlinear Dimensionality Reduction Methods for Use with Automatic Speech Recognition" Speech Technologies. doi:10.5772 / 16863. ISBN 978-953-307-996-7, and Lakshmi et al., 2016, "2016 IEEE 6th International Conference on Advanced Computing (IACC)" pp. 31-34. doi:10.1109 / IACC.2016.16, ISBN978-1-4673-8286-1, which are hereby incorporated by reference in their entirety for all purposes.

[0147] Referring to block 216, in some embodiments, the dimensionality reduction is principal component analysis (PCA) reduction, and the dimensionality reduction decomposes a plurality of spectra into respective subsets of principal components.

[0148] In such embodiments, the number of principal components in the subset of principal components can be limited to a number that accounts for a threshold amount of variance in the data to which the dimensionality reduction is applied (e.g., a threshold amount of the total variance of the output spectra corresponding to a plurality of validation samples).

[0149] In general, different models and / or different sets of validation samples can produce outputs with different dimensions. This can be a problem because higher-dimensional outputs are observed, which has an advantage when evaluating the ability of a model to perform sample label-dependent separation (e.g., when calculating the distance between principal components that account for the variance of a validation set). This can be due to a large number of dimensions contributing in total. Naive divergence measurements (such as Jensen-Shannon (JS) divergence) are not only favorable for higher-dimensional outputs but also do not consider output correlations. This presents further problems because the output spectra can be highly correlated along the final dimension. Thus, without being limited to any one theory of operation, by limiting the dimension of the output spectra to account for a threshold percentage of the variance in the data, it is possible to remove linear dependencies in the data that can unduly skew divergence measurements in favor of higher-dimensional outputs.

[0150] Referring to block 218, in some embodiments, the threshold amount of the total variance is at least 90%, at least 95%, or at least 99% of the total variance. In some embodiments, the threshold amount of the total variance is at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% of the total variance. In some embodiments, the threshold amount of the total variance is 99% or less, 98% or less, 95% or less, 90% or less, 85% or less, or 80% or less of the total variance. In some embodiments, the threshold amount of the total variance is 70% - 80%, 80% - 90%, 85% - 95%, 90% - 99%, or 95% - 100% of the total variance. In some embodiments, the threshold amount of the total variance falls within another range that starts at 70% or more and ends at 100% or less.

[0151] In some embodiments, each respective component value set within the plurality of component value sets corresponds to each respective verification sample within the plurality of verification samples and represents the dimensionally reduced output for each respective verification sample. For example, consider the case where the plurality of spectra obtained from each model is a tensor of shape (N, D), where N is the number of verification samples in the verification set and D is the dimension of the output. The dimensionality reduction then results in a new tensor of shape (N, D_pca), where N is the number of verification samples in the verification set and D_pca is the reduced dimension of the output. Thus, the plurality of component value sets represents the decomposition model output (e.g., PCA reduction) for the plurality of verification samples.

[0152] In some embodiments, each respective principal component in the subset of principal components includes each respective component value in the corresponding component value set for each respective verification sample among the plurality of verification samples.

[0153] In some embodiments, the plurality of component value sets includes at least 100 component value sets.

[0154] In some embodiments, the plurality of component value sets include component value sets of at least 10, at least 50, at least 100, at least 500, at least 1000, at least 2000, at least 5000, at least 10,000, at least 20,000, at least 50,000, at least 100,000, at least 200,000, at least 500,000, at least 1 million, at least 2 million, at least 3 million, at least 4 million, or at least 5 million. In some embodiments, the plurality of component value sets include component value sets of 10 million or less, 5 million or less, 4 million or less, 1 million or less, 500,000 or less, 100,000 or less, 50,000 or less, 10,000 or less, 5000 or less, 1000 or less, or 100 or less. In some embodiments, the plurality of component value sets consist of component value sets of 10 to 5000, 500 to 10,000, 10,000 to 500,000, 20,000 to 1 million, 1 million to 5 million, or 2 million to 10 million. In some embodiments, the plurality of component value sets begin with a component value set of 10 or more and end with a component value set of 10 million or less and fall within another range.

[0155] In some embodiments, the subset of principal components includes at least 2, at least 5, at least 10, at least 20, at least 50, at least 100, at least 500, at least 1000, or at least 5000 principal components. In some embodiments, the subset of principal components includes 10,000 or less, 5000 or less, 1000 or less, 500 or less, 100 or less, 20 or less, or 10 or less principal components. In some embodiments, the subset of principal components is composed of principal components of 2 to 20, 10 to 100, 80 to 500, 300 to 2000, or 1000 to 10,000. In some embodiments, the subset of principal components begins with 2 or more principal components and ends with 10,000 or less principal components and falls within another range.

[0156] In some embodiments, the method does not include performing dimensionality reduction.

[0157] Determine divergence.

[0158] Referring to block 220, in some embodiments, the method includes a corresponding divergence section 156 that uses a mathematical combination of corresponding plural distances 154 for each respective model 132 of a plurality of models, where each respective distance of the corresponding plural distances represents each respective label within a plurality of labels and is between (i) a set of component values for each respective label subset of the validation samples 122 to which each respective label is assigned and (ii) a set of component values for all other samples within the plurality of samples.

[0159] In some embodiments, each respective distance is obtained between different label subsets of the validation samples, and each respective label subset of the validation samples corresponds to each respective label within the plurality of labels. In some embodiments, each respective distance is a statistical distance.

[0160] For example, referring back to the above example, consider a new-shaped tensor (N, D_pca) with dimensionality reduction, where N is the number of validation samples in the validation set and D_pca is the dimensionality reduction of the output. For each respective validation sample of N, for each respective dimension of D_pca, the plural sets of component values include the respective component values for each respective validation sample. Next, in an exemplary embodiment, the component values across the plural dimensions of the first validation sample N1 represent a first distribution of component values, the component values across the plural dimensions of the second validation sample N2 represent a second distribution of component values, the first validation sample has a first label, and the second validation sample has a second label. Then, a distance (e.g., a statistical distance) between the two distributions can be obtained. In some embodiments, the distance is determined to evaluate the ability of the model to separate validation samples between at least a first label and a second label of the plurality of labels at each respective layer for the output.

[0161] In some embodiments, for each respective model among a plurality of models, the corresponding divergence is determined as the sum of distances between distributions of sets of component values for each respective dimension within a plurality of dimensions, with respect to the distribution of sets of component values for each respective dimension within the plurality of dimensions.

[0162] For example, referring back to the above example, consider the tensor (N, D_pca) of the new shape with dimensionality reduction, where N is the number of validation samples in the validation set and D_pca is the dimensionality reduction of the output. For each respective dimension of D_pca, for each respective validation sample of N, the plurality of sets of component values include respective component values for each respective dimension. Next, in an exemplary embodiment, the component values across a plurality of validation samples for the first dimension D1 represent a first distribution of component values, and the component values across a plurality of validation samples for the second dimension D2 represent a second distribution of component values. Then, the distance between the two distributions can be obtained. In some embodiments, the plurality of dimensions are a plurality of components (e.g., dimensionality reduction components). In some embodiments, as described above, one or more dimensions of the plurality of dimensions represent one or more corresponding labels of the plurality of labels.

[0163] In some embodiments, the distance is determined without performing dimensionality reduction (e.g., non-reduced tensor (N, D)).

[0164] Referring to block 222, in some embodiments, each distance among the corresponding plurality of distances is determined in a pairwise fashion between (i) a set of component values for each respective label subset of validation samples and (ii) a set of component values for the corresponding label subset for each respective label among the plurality of labels. In some embodiments, the distance is determined in a pairwise fashion between validation samples within different label subsets and / or between dimensions (e.g., between components and / or labels).

[0165] In some embodiments, for each respective label subset of the validation samples, the corresponding mathematical combination of distances is determined for all other samples within the plurality of samples by summing a plurality of pairwise statistical distances obtained between each respective label subset and each other label subset to which each label within the plurality of labels is assigned.

[0166] Referring to block 224, in some embodiments, the corresponding mathematical combination of distances is the sum of the corresponding plurality of distances.

[0167] Accordingly, in some implementations, divergence is determined using a one-vs.-the-rest approach between each validation sample in subsets of each other except the first subset, between each validation sample within the first subset. As another or additional alternative, in some implementations, divergence is determined using a one-vs.-the-rest approach between each dimension, component, and / or label, for the mutual dimensions, components, and / or labels.

[0168] Referring to block 226, in some embodiments, the corresponding divergence is selected from the group consisting of total variation distance, Hellinger distance, Lévy-Prohorov metric, Wasserstein metric, Mahalanobis distance, Amari distance, Kullback–Leibler divergence, Renyi divergence, Jensen–Shannon divergence, Bhattacharyya distance, f-divergence, and discrimination index. As will be apparent to those skilled in the art, other statistical measures are contemplated for use herein.

[0169] Referring to block 228, in some embodiments, the corresponding divergence is the Jensen-Shannon divergence. The JS divergence is an asymmetric measure that measures the relative entropy or difference in information represented by two distributions. Based on the Kullback-Leibler (KL) divergence, the JS divergence can be considered as a way to measure the distance or similarity between two probability distributions to determine how different the two distributions are from each other. For example, FIGS. 4A - B show obtaining the divergence after the dimensionality reduction step as the PCA-reduced JS divergence (the measurements indicated by the dashed circles).

[0170] Thus, in some embodiments, the method includes obtaining, for each respective model of a plurality of models, a corresponding divergence indicative of how well the model separates validation samples in a label-dependent manner.

[0171] Model selection.

[0172] Referring to block 230, in some embodiments, the method further includes identifying a first model 132 among a plurality of models having corresponding divergences 156 that meet a threshold for performing a first category task. In some embodiments, each model meets the threshold when it has the largest corresponding divergence among the plurality of models. In some embodiments, each model meets the threshold when it has a corresponding divergence within the top N largest corresponding divergences. In some embodiments, N is a positive integer from 1 to 5. In some embodiments, N is at least 1, at least 2, at least 3, or at least 5. In some embodiments, N is 10 or less, 5 or less, or 3 or less. In some embodiments, N is from 1 to 5, from 2 to 8, or from 5 to 10. In some embodiments, N falls within a range starting with 1 or more and ending with 10 or less. In some embodiments, each model meets the threshold when it has a corresponding divergence within the top N percent of the largest corresponding divergences. In some embodiments, N is 1% or less, 5% or less, 10% or less, 20% or less, or 40% or less. In some embodiments, N is at least 50%, at least 40%, at least 20%, at least 10%, or at least 5%. In some embodiments, N is from 5% to 50%, from 2% to 30%, or from 1% to 10%. In some embodiments, N falls within a range starting with 1 or more and ending with 50% or less.

[0173] In some embodiments, identifying further includes selecting a subset of models among a plurality of models having the top N largest corresponding divergences. In some embodiments, N is a positive integer from 1 to 5. In some embodiments, N is at least 1, at least 2, at least 3, or at least 5. In some embodiments, N is 10 or less, 5 or less, or 3 or less. In some embodiments, N is from 1 to 5, from 2 to 8, or from 5 to 10. In some embodiments, N falls within a range starting with 1 or more and ending with 10 or less.

[0174] Referring to block 232, in some embodiments, the method further includes retraining a first model to perform a first category task. In some implementations, retraining includes performing a training procedure using the first model on a plurality of training samples to perform the first task.

[0175] In some embodiments, the method further includes fine-tuning a first model to execute a first task.

[0176] In some embodiments, the method further includes determining a validation score for the first model after retraining and / or fine-tuning. In some embodiments, the validation score is selected from the group consisting of accuracy, recall, and F1 score.

[0177] Referring to block 234, in some embodiments, the method further includes identifying a subset of layers within a plurality of layers of the first model and removing layers other than the subset of layers from the first model before retraining. In some embodiments, identifying the subset of layers includes updating the architecture of the first model.

[0178] In some embodiments, updating the architecture of the model comprises: A) for each respective validation sample among a plurality of validation samples, inputting corresponding information into the model and obtaining a corresponding spectrum comprising a plurality of values as an output from each respective layer of a plurality of layers of the model, thereby obtaining a plurality of spectra having a total variance, wherein the model is pre-trained on each task other than a first category task and the obtaining is such that each layer of the model comprises a corresponding set of pre-trained weights; B) performing dimensionality reduction on the plurality of spectra to obtain a plurality of sets of component values collectively having an explained variance of at least a threshold amount of the total variance, wherein the obtaining is such that the plurality of sets of component values comprises a set of component values corresponding to each respective layer of the plurality of layers; C) determining a first layer among the plurality of layers associated with the set of component values having the highest dimension; and D) removing each layer among the plurality of layers downstream of the first layer, thereby updating the architecture of the model to perform the first task.

[0179] Non-limiting exemplary methods for updating or optimizing the model to perform a first category task are further described below with reference to FIGS. 3A-C.

[0180] Example of an embodiment of updating the model to perform a task.

[0181] FIGS. 3A-C collectively show a flowchart of an exemplary method 300 for updating the architecture of a model 132 to perform a first category task, according to some embodiments of the present disclosure, where the dashed boxes represent optional portions of the method. In some embodiments, the method is executed by a computer system comprising one or more processors and memory. In some embodiments, the method is performed by modules of the computer system 100 as detailed elsewhere in this specification.

[0182] Referring to block 302, in some embodiments, the method includes, for each respective validation sample 122 among a plurality of validation samples, inputting corresponding information into model 132 and obtaining, as an output from each respective layer 136 of a plurality of layers of model 132, a corresponding spectrum 142 that includes a corresponding plurality of values, thereby obtaining a plurality of spectra having total variance, the model being pre-trained on each task other than the first category task, and each layer 136 of the model including a corresponding set of pre-trained weights 134.

[0183] For example, FIG. 4A is obtained using the output from each respective hidden layer of a plurality of N hidden layers. Similarly, FIG. 4B shows measurements obtained using the output from each respective layer of a plurality of hidden layers. The measurements are indicated by solid circles.

[0184] Referring to block 304, in some embodiments, the first category task includes determining a patient-drug relationship, determining a patient-biomarker association, or determining a disease state.

[0185] Referring to block 306, in some embodiments, the model is selected from the group consisting of a language model, a transformer model, a large language model (LLM), an encoder, a decoder, an encoder-decoder hybrid model, a generative pre-trained transformer (GPT) model, and a bidirectional encoder representation from transformers (BERT) model. In some embodiments, the model is selected from the group consisting of: BERT, BERT Base, BERT large, RoBERTa Base, BioBERT Base, RoBERTa BaseTwitter Sentiment Finetune, DeBERTa, ALBERT, RoBERTa, GPT-J, GPT-Neo, GPT-NeoX, Pythia, GPT-NeoX2.0, XLNet, LaMDA, PaLM, Gopher, Sparrow, Chinchilla, Minerva, Bard, GPT-1, GPT-2, GPT-3, CodeX, InstructGPT, ChatGPT, GPT-4, OPT, Galactica, LLaMA, BART, Flan-T5, Flan-UL2, T5, and / or any derivatives or combinations thereof. In some embodiments, the model is any of the models disclosed elsewhere in this specification (see, for example, the section entitled "Definitions: Models" above, and "Example Embodiments for Identifying Models to Perform a Task"). In some embodiments, the model is pre-trained using a set of unspecific pre-training samples.

[0186] In some embodiments, the model is pre-trained using a set of domain-specific pre-training samples. In some embodiments, the domain is associated with a first category of tasks. In some embodiments, the model is fine-tuned for the first category of tasks.

[0187] Referring to block 308, in some embodiments, the plurality of layers includes at least 5 layers, at least 10 layers, or at least 15 layers.

[0188] Referring to block 310, in some embodiments, each respective layer within the plurality of layers includes at least 5, at least 10, or at least 15 nodes.

[0189] Referring to block 312, in some embodiments, for each respective layer of the plurality of layers, the output is obtained from the first node of the plurality of nodes.

[0190] As described above, in some embodiments, the model includes at least 2 layers, at least 5 layers, at least 10 layers, at least 20 layers, at least 50 layers, at least 100 layers, or at least 500 layers. In some embodiments, the model is composed of 1000 or fewer, 500 or fewer, 100 or fewer, 50 or fewer, or 10 or fewer layers. In some embodiments, the model consists of 2 to 20 layers, 5 to 100 layers, 50 to 300 layers, 200 to 500 layers, or 300 to 1000 layers. In some embodiments, the model starts with 2 or more layers and ends with 1000 or fewer layers, falling within a range.

[0191] As described above, in some embodiments, for each hidden layer within the plurality of hidden layers, the output is obtained from each node within the plurality of nodes for each hidden layer. For example, FIG. 4A shows that the output is obtained from a first node of a plurality of nodes (e.g., node 1) within each hidden layer used to obtain a measured value (e.g., the measured value indicated by the solid circle). However, FIG. 4A shows the output obtained from the first node, and as will be apparent to those skilled in the art, the output from any node within each hidden layer is contemplated. In some embodiments, each hidden layer within the plurality of hidden layers includes a plurality of nodes. In some embodiments, the output is obtained from any node within the plurality of nodes. In some embodiments, the output is obtained from the same node for each respective layer of a plurality of layers (e.g., the first node of each layer, the second node of each layer, the second-to-last node of each layer, the last node of each layer, etc.). In some embodiments, for the first layer within the plurality of layers, the output is obtained from different nodes (e.g., the first node of the first layer and the second node of the second layer) for the second layer.

[0192] In some embodiments, each layer within the plurality of layers includes at least 2, at least 5, at least 10, at least 20, at least 50, at least 100, at least 500, at least 1000, at least 10,000, or at least 100,000 nodes. In some embodiments, the plurality of nodes within each layer includes 1 million or fewer, 100,000 or fewer, 10,000 or fewer, 1000 or fewer, 500 or fewer, 100 or fewer, 50 or fewer, or 10 or fewer nodes. In some embodiments, the plurality of nodes within each layer consists of 2 to 20, 5 to 100, 50 to 300, 200 to 500, 300 to 1000, 1000 to 10,000, or 10,000 to 1 million nodes. In some embodiments, the plurality of nodes within each layer begins with 2 nodes or more and ends with 1 million nodes or fewer and falls within another range.

[0193] In some embodiments, the corresponding set of pre-trained weights includes at least 1000 weights.

[0194] In some embodiments, the model is selected by a method for identifying a model that performs a first category task, the method comprising: A) for each respective model of a plurality of models, each respective model of the plurality of models is pre-trained on each respective task other than the first category task, each respective model comprising a corresponding plurality of layers including a corresponding input layer, a corresponding output layer, and a corresponding plurality of hidden layers, inputting corresponding information into each respective model, and obtaining, through application of corresponding plurality of parameters of each respective model to the corresponding information, an output from each respective hidden layer within the corresponding plurality of hidden layers in the form of a corresponding spectrum including a corresponding plurality of values, wherein the plurality of validation samples includes a corresponding label subset of validation samples assigned each respective label for each respective label of a plurality of labels, whereby for each respective model of the plurality of models, a corresponding plurality of spectra having a corresponding total variance is obtained; B) for each respective model of the plurality of models, performing dimensionality reduction on the corresponding plurality of spectra to obtain a corresponding plurality of sets of component values that collectively have at least a threshold amount of explained variance of the total variance, wherein the corresponding plurality of sets of component values is obtained such that it includes a set of component values corresponding to each respective validation sample among the plurality of validation samples; C) for each respective model of the plurality of models, determining a corresponding divergence using a mathematical combination of corresponding plurality of distances, wherein each respective distance of the corresponding plurality of distances represents each respective label within the plurality of labels, and is determined such that it is between (i) a set of component values for a respective label subset of validation samples assigned each respective label and (ii) a set of component values for all other samples within the plurality of samples; and D) identifying, within the plurality of models, a first model having a corresponding divergence that meets a threshold for performing a first task.

[0195] In some embodiments, the model includes or is selected using any of the embodiments disclosed elsewhere in this specification (see, e.g., the section entitled "Exemplary Embodiments for Identifying a Model to Perform a Task" above).

[0196] In some embodiments, each respective validation sample among the plurality of validation samples includes all or a portion of an electronic health record (EHR) or an electronic medical record (EMR). In some embodiments, each respective validation sample among the plurality of validation samples includes any of the embodiments for the validation samples described above (see, e.g., the section entitled "Exemplary Embodiments for Identifying a Model to Perform a Task" above).

[0197] In some embodiments, the corresponding information for each respective validation sample among the plurality of validation samples includes one or more corresponding snippets.

[0198] In some embodiments, the plurality of validation samples includes a corresponding label subset of the validation samples for each respective label among the plurality of labels. In some embodiments, as described above, the plurality of validation samples collectively represent a plurality of labels, and each respective validation sample among the plurality of validation samples includes the corresponding label among the plurality of labels.

[0199] In some embodiments, the plurality of validation samples includes at least 100 validation samples.

[0200] In some embodiments, the output from the model is obtained by applying a corresponding set of pre-trained weights to the information for each validation sample among the plurality of validation samples.

[0201] In some embodiments, for each respective validation sample among the plurality of validation samples, the corresponding spectrum includes a plurality of dimensions.

[0202] Referring to block 313, in some embodiments, the method further includes performing dimensionality reduction on the plurality of spectra 142 to obtain a plurality of sets of component values 152 that collectively have at least a threshold amount of the explained variance of the total variance, and the plurality of sets of component values 152 includes a set of component values corresponding to each respective layer 136 within the plurality of layers.

[0203] Referring to block 314, in some embodiments, the dimensionality reduction is a principal component analysis algorithm, a random projection algorithm, an independent component analysis algorithm, or a feature selection method.

[0204] Referring to block 316, in some embodiments, the dimensionality reduction is a principal component analysis (PCA) reduction, and the dimensionality reduction decomposes the plurality of spectra into respective subsets of principal components.

[0205] Referring to block 318, in some embodiments, the threshold amount of the total variance is at least 90%, at least 95%, or at least 99% of the total variance.

[0206] Referring to block 320, in some embodiments, the method further includes determining a first layer 136 in a plurality of layers associated with a set of component values 152 of the plurality of sets of component values having a highest dimension 158.

[0207] Referring to block 322, in some embodiments, the dimensions include a plurality of principal components determined using dimensionality reduction. In some embodiments, the dimensions are PCA-reduced dimensions. For example, referring to FIGS. 4A - B, in some embodiments, the method includes obtaining PCA-reduced dimensions using the output from each respective hidden layer within a plurality of hidden layers of the model (the measurements indicated by the solid circles at each layer). A layer having the highest or maximum PCA-reduced dimension (e.g., the first layer 136) is determined (e.g., layer M in FIG. 4A). FIG. 6 further illustrates an exemplary plot of the PCA-reduced dimensions obtained for each hidden layer within a plurality of hidden layers in the model, and the maximum PCA-reduced dimension is identified as layer 8. All layers following layer 8 are shown to have lower PCA-reduced dimensions.

[0208] Referring to block 324, in some embodiments, the plurality of principal components include at least 10, at least 100, or at least 1000 principal components.

[0209] In some embodiments, the plurality of principal components include at least 2, at least 5, at least 10, at least 20, at least 50, at least 100, at least 500, at least 1000, or at least 5000 principal components. In some embodiments, the plurality of principal components include 10,000 or fewer, 5000 or fewer, 1000 or fewer, 500 or fewer, 100 or fewer, 20 or fewer, or 10 or fewer principal components. In some embodiments, the plurality of principal components consist of 2 - 20, 10 - 100, 80 - 500, 300 - 2000, or 1000 - 10,000 principal components. In some embodiments, the plurality of principal components fall within another range of principal components that starts with 2 or more principal components and ends with 10,000 or fewer principal components.

[0210] In some embodiments, the dimensions are determined using JS divergence and / or PCA-reduced JS divergence. Non-limiting exemplary methods for determining JS divergence are described elsewhere in this specification (see, e.g., the section entitled "Exemplary Embodiments for Identifying a Model for Performing a Task" above).

[0211] Referring to block 326, in some embodiments, the method further includes removing each layer 136 of a plurality of layers downstream of the first layer, thereby updating the architecture of model 132 to perform a first category task. For example, as shown in FIGS. 4A - B, after determining the layer (e.g., the first layer 136) having the highest or maximum PCA-reduced dimension (e.g., layer M in FIG. 4A, or layer 8 in FIG. 6), each layer having a PCA-reduced dimension that is the maximum PCA-reduced dimension and that follows a layer with a lower PCA-reduced dimension is removed from the model (e.g., all layers after layer M in FIG. 4A, or all layers after layer 8 in FIG. 6).

[0212] As described above, in some embodiments, the model includes a plurality of hidden layers. Without being limited to any one theory of operation, the lower layers may be more likely to facilitate lower-resolution identification or classification, while the upper layers may be more likely to fine-tune or facilitate the model's ability to perform high-resolution identification or classification with higher specificity for the fine details tailored to the intended task or domain of the model. Since such details may not be relevant to the task or domain of interest, it is advantageous to remove such upper layers while retaining the underlying engine subsumed by the lower layers.

[0213] Referring to block 328, in some embodiments, the model further includes a task-dependent output layer downstream of the plurality of layers, and removing further includes removing the task-dependent output layer. For example, in some embodiments, the output layer is a classifier head that generates task-dependent classifications. In some embodiments, the method further includes adding a task-dependent output layer downstream of the plurality of layers, and the task-dependent output layer is specific to a first category task.

[0214] Referring to block 330, in some embodiments, the method further includes retraining the updated model to perform a first category task. For example, as shown in FIG. 4A, in some embodiments, the updated model includes an architecture that includes only the layers up to layer M having a maximum reduced dimension (all layers following layer M are removed). Next, the output from layer M is used as an input to a classifier head to retrain the updated model to perform the first category task.

[0215] In some embodiments, retraining includes performing a training procedure using the first model on a plurality of training samples to perform a first category task.

[0216] Additional embodiments

[0217] Yet another aspect of the present disclosure provides a computer system including one or more processors and a non-transitory computer-readable medium including computer-executable instructions that, when executed by the one or more processors, cause the processor to execute any of the methods and / or embodiments disclosed herein.

[0218] Yet another aspect of the present disclosure provides a non-transitory computer-readable storage medium storing program code instructions that, when executed by a processor, cause the processor to execute any of the methods and / or embodiments disclosed herein.

[0219] Any of the embodiments disclosed herein for selecting a model to perform a first category of tasks (see, e.g., the section entitled "Exemplary Embodiments for Identifying a Model to Perform a Task" above) is similarly intended for use in a method for updating the architecture of a model to perform a first task, as would be apparent to one of ordinary skill in the art. Further, any of the embodiments disclosed herein for updating or optimizing a model to perform a first task (see, e.g., the section entitled "Exemplary Embodiments for Updating a Model to Perform a Task" above) is similarly contemplated for use in a method for selecting a model to perform a first task, as would be apparent to one of ordinary skill in the art.

Example

[0220]

[0221] Exemplary Comparison of Models Identified Using Example 1 - PCA Reduction JS Divergence

[0222] Several pre - trained deep learning models (DLMs) were obtained and evaluated for use in patient - drug relationship modeling. The models included three pre - trained models trained on general domain data (BERT Base, BERT Large, and RoBERTa Base) without fine - tuning, one pre - trained model (BioBERT Base, trained in the biomedical domain) trained on a domain related to the task of interest, one pre - trained model (RoBERTa BaseTwitter Sentiment Finetune) trained on a different domain not related to the task of interest but fine - tuned to perform a similar task, and an untrained model (BERT Base Untrained). The validation samples included text snippets obtained from electronic health records and labeled with various class labels related to the patient - drug relationship.

[0223] For each model, the PCA-reduced JS divergence was obtained using the method disclosed herein. The method involves, for each model of a plurality of models, for each respective validation sample among a plurality of validation samples, inputting corresponding information into each respective model, and obtaining an output from a layer of each respective model in the form of a corresponding spectrum through the application of a corresponding plurality of parameters of each respective model to the corresponding information, thereby obtaining a plurality of spectra corresponding to each respective model. PCA was performed on the spectra of each model, and the PCA dimension was reduced to the number of components that explained 99% of the variance of the data, thus obtaining PCA-reduced spectra. For each respective model of a plurality of models, the JS divergence was determined in a one-versus-the-rest fashion between class labels as the sum of distances between, for each component, a set of component values for each respective subset of validation samples (e.g., each first label), and a corresponding set of component values for a corresponding label subset for each other label among a plurality of labels (e.g., each respective subset of each other validation samples corresponding to labels other than the first label). The PCA-reduced JS divergence for each evaluation model is shown in Table 1.

[0224]

Table 1

[0225] In Table 1, the BERT Large pre-trained model had the highest PCA-reduced JS divergence.

[0226] Next, each model was trained with the training data, and it was evaluated whether the PCA-reduced JS divergence was well correlated with the actual ability of the model to perform patient-drug relationship modeling.

[0227] The training data included text snippets labeled with weak labels "administered", "ordered", "considered", "rejected", and "invalid". To maintain the simplicity of the experiment, multi-labeled examples were removed. Mixed-precision training was performed with a batch size of 64 using the AdamW optimizer, and the learning rate = 1×10-5 and were the default parameters. Training was performed for the necessary number of epochs until the validation F1 plateau was reached or overfitting occurred. After the first epoch, the learning rate was reduced to 1×10 -6 . Overfitting was determined by measuring the Wilcoxon rank-sum test p-value between the non-reduced loss distributions of the validation set and the training set.

[0228] The results of model training and validation are shown in Table 2 and Figure 5.

[0229]

Table 2

[0230] Table 2 and Figure 5 show that the BERT Large model showed the best performance after training, as predicted by PCA-reduced JS divergence. This trend is also shown in Figure 5, where the F1 score was well correlated with PCA-reduced JS divergence across all models evaluated.

[0231] In particular, domain-specific models did not necessarily result in better downstream performance in the same domain. However, task-similar fine-tuning can be more beneficial for downstream performance than domain similarity. The untrained BERT model had somewhat comparable performance to other models.

[0232] The results of the experiment showed that there was actually a correlation between PCA-reduced JS divergence and the macro F1 of the test data. In this scenario, the model selection had a significant impact on the final performance.

[0233] Example 2 - Comparative example of updated and non-updated models using PCA-reduced dimensions

[0234] Pre-trained models were evaluated to identify and optimize a model for performing a computer vision task, namely skin lesion classification.

[0235] A collection of skin lesion images containing labels that describe the types of skin lesions in the images was obtained from a database (Huggingface Datasets hub). The classes are as follows: actinic_keratoses, benign_keratosis-like_lesions, basal_cell_carcinoma, dermatofibroma, vascular_lesions, melanoma, and melanocytic_Nevi.

[0236] Next, several common pre-trained visual models were evaluated for PCA-reduced JS divergence in the manner described in Example 1 above. These models and the corresponding PCA-reduced JS divergences are shown below. google / vit-large-patch32-384: 2559 google / vit-base-patch16-224: 2554 microsoft / beit-large-patch16-224-pt22k-ft22k: 2527 facebook / convnext-xlarge-224-22k: 2518 microsoft / resnet-50: 2466

[0237] As described in Example 1 above, the validation of these models also showed that in this domain and modality, a greater PCA-reduced JS divergence of the model output spectrum still retains predictive power over the final downstream performance.

[0238] The google VIT Large patch 32-384 model was found to have the largest PCA-reduced JS divergence, so this model was selected for further optimization.

[0239] Next, PCA dimensionality reduction of the spectra and PCA-reduced JS divergence for the outputs of each hidden layer were examined to evaluate the advantages of removing a specific layer from the model. The spectra of each layer were collected and PCA was reduced according to the method disclosed herein. Briefly, the corresponding information of the dermatological lesion images was input into the selected model. The outputs were obtained from each respective layer of the model as the corresponding spectra, thereby obtaining a plurality of spectra. PCA was performed on the spectra of each layer, reducing the PCA dimensions to the number of components that explained 99% of the variance of the data, thus obtaining PCA-reduced spectra. Optionally, JS divergence was calculated in a one-versus-the-rest fashion for each PCA-reduced spectrum, and it was further examined whether this metric correlated with the PCA-reduced dimensions.

[0240] As shown in FIG. 6, it was found that the layer with the maximum PCA-reduced dimension was layer 8. Both the eight-hidden-layer model and the full model were trained with the training set of dermatological lesion images and compared as shown in Tables 3 and 4.

[0241] [Table 3]

[0242] [Table 4] JPEG2025096265000007.jpg28170

[0243] As can be seen from Tables 3 and 4, the eight-layer fine-tuned model was far superior to the full fine-tuned model. In other words, as shown in FIG. 6, it was found that the layer that resulted in the highest PCA-reduced dimension also resulted in the strongest results. In fact, the full model when trained resulted in a macroscopic test F1 score of 0.76, while the same model with only the layers up to layer 8 (the layer with the maximum PCA-reduced dimension) resulted in a macroscopic test F1 score of 0.86, which was quite good.

[0244] Advantageously, these results show that the systems and methods of the present disclosure can be used to identify a subset of models that function as well as, or better than, existing pre-trained models for a given task. By identifying and optimizing such models into smaller subsets, training, validating, fine-tuning, and / or using the models (e.g., modeling, predicting, and / or classifying) can be performed in a faster and less expensive manner. Thus, the systems and methods of the present disclosure improve the efficiency of such modeling tasks (e.g., using a subset of layers) compared to existing models (e.g., using existing, full-size pre-trained models).

[0245] Conclusion The foregoing description has been presented for purposes of illustration and is described with reference to specific implementations. However, the following illustrative considerations are not intended to be exhaustive or to limit the embodiments to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The implementations were chosen and described in order to best explain the principles and practical examples, thereby enabling others skilled in the art to best utilize the implementations and various implementations with various modifications suitable for the particular use contemplated.

Claims

1. 1. A method for identifying a model that performs a first category task, comprising: A computer system including one or more processors and a memory, A) inputting corresponding information for each respective validation sample in a plurality of validation samples into a respective model of a plurality of models, each respective model of the plurality of models having been pre-trained at least in part on a respective task other than the first category task, each respective model including a corresponding plurality of layers including a corresponding input layer, a corresponding output layer, and a corresponding plurality of hidden layers, and obtaining, through application of a corresponding plurality of parameters of the respective model to the corresponding information, an output from a respective hidden layer in the corresponding plurality of hidden layers in the form of a corresponding spectrum including a corresponding plurality of values, wherein the plurality of validation samples include, for each respective label of a plurality of labels, a corresponding label subset of validation samples assigned the respective label, thereby obtaining, for each respective model of the plurality of models, a corresponding plurality of spectra having a corresponding total variance across the corresponding plurality of values ​​for each respective validation sample in the plurality of validation samples; B) for each respective model of the plurality of models, performing dimensionality reduction of the corresponding plurality of spectra to obtain a corresponding plurality of component value sets that collectively have an explained variance of at least a threshold amount of the total variance, the corresponding plurality of component value sets including a component value set corresponding to each respective validation sample in the plurality of validation samples; C) for each respective model of the plurality of models, determining a corresponding divergence using a mathematical combination of a corresponding plurality of distances, each respective distance of the corresponding plurality of distances representing a respective label in the plurality of labels and between (i) the set of component values ​​for the respective label subset of the validation samples assigned the respective label and (ii) the set of component values ​​for all other samples in the plurality of samples; D) identifying a first model within the plurality of models having a corresponding divergence that satisfies a threshold for performing the first task.

2. The method of claim 1 , wherein the first task comprises determining a patient-drug relationship, determining a patient-biomarker association, or determining a disease state.

3. 3. The method of claim 1 or 2, wherein each model in the plurality of models is selected from the group consisting of a language model, a Transformer model, a large-scale language model (LLM), an encoder, a decoder, an encoder-decoder hybrid model, a Generative Pre-Trained Transformer (GPT) model, and a Bidirectional Encoder Representation from Transformer (BERT) model.

4. Each model in the plurality of models is BERT, BERT Base, BERT large, RoBERTa Base, BioBERT Base, RoBERTa Base Twitter Sentiment 4. The method of any one of claims 1 to 3, wherein the nucleic acid sequence is selected from the group consisting of Finetune, DeBERTa, ALBERT, RoberTa, GPT-J, GPT-Neo, GPT-NeoX, Pythia, GPT-NeoX2.0, XLNet, LaMDA, PaLM, Gopher, Sparrow, Chinchilla, Minerva, Bard, GPT-1, GPT-2, GPT-3, CodeX, InstructGPT, ChatGPT, GPT-4, OPT, Galactica, LLaMA, BART, Fla-T5, Fla-UL2, and T5.

5. The method of any one of claims 1 to 4, wherein one or more of the plurality of models are pre-trained using a set of non-specific pre-training samples.

6. The method of any one of claims 1 to 5, wherein one or more of the plurality of models are pre-trained using a set of domain-specific pre-training samples.

7. The method of claim 6 , wherein the domain is associated with the first task.

8. The method of any one of claims 1 to 7, wherein one or more models of the plurality of models are fine-tuned for the first task.

9. The method of any one of claims 1 to 8, wherein the plurality of models comprises at least five models.

10. The method of any one of claims 1 to 9, wherein for each model in the plurality of models, the corresponding plurality of parameters comprises at least 1000 parameters.

11. The method of any one of claims 1 to 10, wherein each respective validation sample in the plurality of validation samples comprises all or a portion of an electronic health record (EHR) or electronic medical record (EMR).

12. The method of any one of claims 1 to 11, wherein the corresponding information for each respective validation sample in the plurality of validation samples comprises one or more corresponding snippets.

13. The method of any one of claims 1 to 12, wherein the plurality of validation samples comprises at least 100 validation samples.

14. The method according to any one of claims 1 to 13, wherein the dimensionality reduction is a principal component analysis algorithm, a random projection algorithm, an independent component analysis algorithm, or a feature selection method.

15. The method of any one of claims 1 to 14, wherein the dimensionality reduction is a Principal Component Analysis (PCA) reduction, the dimensionality reduction decomposing the plurality of spectra into a respective subset of principal components.

16. The method of any one of claims 1 to 15, wherein the threshold amount of the total variance is at least 90%, at least 95%, or at least 99% of the total variance.

17. 17. The method of claim 1, wherein the corresponding distances are determined in a pair-wise manner between (i) the set of component values ​​for the respective label subsets of a validation sample and (ii) the set of component values ​​for corresponding label subsets for each other label in the plurality of labels.

18. The method of claim 17 , wherein the mathematical combination of the corresponding plurality of distances is a sum of the corresponding plurality of distances.

19. 19. The method of any one of claims 1 to 18, wherein the corresponding divergence is selected from the group consisting of total variation distance, Hellinger distance, Levy-Prokhorov metric, Wasserstein metric, Mahalanobis distance, Amari distance, Kullback-Leibler divergence, Renyi divergence, Jensen-Shannon divergence, Bhattacharya distance, f-divergence, and discriminability index.

20. The method according to any one of claims 1 to 19, wherein the corresponding divergence is the Jensen-Shannon divergence.

21. A method according to any preceding claim, wherein the threshold is met when a respective model has the greatest corresponding divergence among the plurality of models.

22. 21. The method of any one of claims 1 to 20, wherein said identifying D) further comprises selecting a subset of models in said plurality of models having the top-N largest corresponding divergence.

23. 23. The method of claim 22, wherein N is a positive integer from 1 to 5.

24. The method further comprising: The method of any one of claims 1 to 23, further comprising: E) retraining the first model to perform the first task.

25. The retraining step includes:

25. The method of claim 24, comprising performing a training procedure using the first model on a plurality of training samples to perform the first task.

26. Prior to said retraining, identifying a subset of layers in the plurality of layers of the first model; The method of claim 24 or 25, further comprising removing layers other than the subset of layers from the first model.

27. 1. A method for updating an architecture of a model to perform a first category task, comprising: A computer system including one or more processors and a memory, A) for each respective validation sample in a plurality of validation samples, inputting corresponding information into the model and obtaining a corresponding spectrum including a corresponding plurality of values ​​as output from each respective layer of a plurality of layers of the model, thereby obtaining a plurality of spectra having total variance, wherein the model is pre-trained on a respective task other than the first categorical task, and each layer of the model includes a corresponding set of pre-trained weights; B) performing dimensionality reduction on the plurality of spectra to obtain a plurality of component value sets that collectively have an explained variance of at least a threshold amount of the total variance, the plurality of component value sets including a component value set corresponding to each respective layer of the plurality of layers; C) determining a first stratum in the plurality of strata associated with a component value set in the plurality of component value sets having a highest dimension; D) removing each layer of the plurality of layers downstream of the first layer, thereby updating the architecture of the model to perform the first task.

28. 28. The method of claim 27, wherein the first task comprises determining a patient-drug relationship, determining a patient-biomarker association, or determining a disease state.

29. 29. The method of claim 27 or 28, wherein the model is selected from the group consisting of a language model, a Transformer model, a large-scale language model (LLM), an encoder, a decoder, an encoder-decoder hybrid model, a Generative Pre-Trained Transformer (GPT) model, and a Bidirectional Encoder Representation from Transformer (BERT) model.

30. The model is BERT, BERT Base, BERT large, RoBERTa Base, BioBERT Base, RoBERTa Base Twitter Sentiment 30. The method of any one of claims 27 to 29, wherein the nucleic acid sequence is selected from the group consisting of Finetune, DeBERTa, ALBERT, RoberTa, GPT-J, GPT-Neo, GPT-NeoX, Pythia, GPT-NeoX2.0, XLNet, LaMDA, PaLM, Gopher, Sparrow, Chinchilla, Minerva, Bard, GPT-1, GPT-2, GPT-3, CodeX, InstructGPT, ChatGPT, GPT-4, OPT, Galactica, LLaMA, BART, Fla-T5, Fla-UL2, and T5.

31. The method of any one of claims 27 to 30, wherein the model is pre-trained using a set of non-specific pre-training samples.

32. The method of any one of claims 27 to 31, wherein the model is pre-trained using a set of domain-specific pre-training samples.

33. The method of claim 32 , wherein the domain is associated with the first task.

34. The method of any one of claims 27 to 33, wherein the model is fine-tuned for the first task.

35. The method of any one of claims 27 to 34, wherein the plurality of layers comprises at least 5 layers, at least 10 layers, or at least 15 layers.

36. 36. The method of claim 35, wherein each respective layer of the plurality of layers includes a plurality of nodes of at least 5, at least 10, or at least 15.

37. The method of any one of claims 27 to 36, wherein the corresponding set of pre-trained weights comprises at least 1000 weights.

38. The model is selected by a method for identifying a model that performs a first task, the method comprising: A) inputting corresponding information for each respective validation sample in a plurality of validation samples into a respective model of a plurality of models, each respective model of the plurality of models being pre-trained on a respective task other than the first category task, each respective model including a corresponding plurality of layers including a corresponding input layer, a corresponding output layer, and a corresponding plurality of hidden layers, and obtaining, through application of a corresponding plurality of parameters of the respective model to the corresponding information, an output from a respective hidden layer in the corresponding plurality of hidden layers in the form of a corresponding spectrum including a corresponding plurality of values, wherein the plurality of validation samples include, for each respective label of a plurality of labels, a corresponding label subset of validation samples assigned the respective label, thereby obtaining, for each respective model of the plurality of models, a corresponding plurality of spectra having a corresponding total variance; B) for each respective model of the plurality of models, performing dimensionality reduction of the corresponding plurality of spectra to obtain a corresponding plurality of component value sets that collectively have an explained variance of at least a threshold amount of the total variance, the corresponding plurality of component value sets including a component value set corresponding to each respective validation sample in the plurality of validation samples; C) for each respective model of the plurality of models, determining a corresponding divergence using a mathematical combination of a corresponding plurality of distances, each respective distance of the corresponding plurality of distances representing a respective label in the plurality of labels and between (i) the set of component values ​​for the respective label subset of the validation samples assigned the respective label and (ii) the set of component values ​​for all other samples in the plurality of samples; D) identifying a first model within the plurality of models having a corresponding divergence that satisfies a threshold for performing the first task.

39. 39. The method of any one of claims 27-38, wherein each respective validation sample in the plurality of validation samples comprises all or a portion of an electronic health record (EHR) or electronic medical record (EMR).

40. 40. The method of any one of claims 27 to 39, wherein the corresponding information for each respective validation sample in the plurality of validation samples comprises one or more corresponding snippets.

41. The method of any one of claims 27 to 40, wherein the plurality of validation samples comprises at least 100 validation samples.

42. The method of any one of claims 27 to 41, wherein the dimensionality reduction is a principal component analysis algorithm, a random projection algorithm, an independent component analysis algorithm, or a feature selection method.

43. 43. The method of any one of claims 27 to 42, wherein the dimensionality reduction is a Principal Component Analysis (PCA) reduction, the dimensionality reduction decomposing the plurality of spectra into a respective subset of principal components.

44. 44. The method of any one of claims 27 to 43, wherein the threshold amount of the total variance is at least 90%, at least 95%, or at least 99% of the total variance.

45. The method of any one of claims 27 to 44, wherein the dimensionality comprises a number of principal components determined using the dimensionality reduction.

46. 46. ​​The method of claim 45, wherein the plurality of principal components comprises at least 10, at least 100, or at least 1000 principal components.

47. 47. The method of any one of claims 27 to 46, wherein the model further comprises a task-dependent output layer downstream of the plurality of layers, and wherein the removing further comprises removing the task-dependent output layer.

48. The method further comprising: The method of any one of claims 27 to 47, further comprising: E) retraining an optimization model to perform said first task.

49. The retraining step includes:

49. The method of claim 48, comprising performing a training procedure using the first model on a plurality of training samples to perform the first task.

50. 1. A computer system comprising: one or more processors; and a non-transitory computer readable medium comprising computer executable instructions that, when executed by the one or more processors, cause the processors to perform the method of any one of claims 1 to 49.

51. A non-transitory computer readable storage medium storing program code instructions which, when executed by a processor, cause the processor to perform a method according to any one of claims 1 to 49.