Dynamic mode decomposition for evaluating early stop in training of machine learning models

Dynamic mode decomposition (DMD) is used to predict the failure of large language model training by analyzing observability metrics, allowing for early termination and resource optimization.

WO2025106937A1PCT designated stage expired Publication Date: 2025-05-22MTS IP HLDG LTD +3

Patent Information

Application Number
PCT/US2024/056302
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-15
Filing Date
2024-11-15
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

Training large language models (LLMs) often fails due to poor initialization of parameters, improper model configuration, or defective training data, leading to wastage of computational resources and increased costs.

Method used

The use of dynamic mode decomposition (DMD) to predict future values of observability metrics during the training of LLMs, allowing for early detection of training failures and dynamic adjustment of hyperparameters to prevent failure.

Benefits of technology

DMD enables early termination of training runs likely to fail, saving significant computational resources and reducing costs by identifying potential failures 500 to 2000 iterations in advance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024056302_22052025_PF_FP_ABST
    Figure US2024056302_22052025_PF_FP_ABST
Patent Text Reader

Abstract

A method for early stop using dynamic mode decomposition (DMD) includes determining a best fit linear operator that relates two sets of observability metrics (also referred to herein as "states") associated with the training of a machine learning model, such as an LLM. The best fit linear operator, in turn, is used to predict future values of the observability metrics at future iteration steps in the training. The predicted values are then evaluated against an early stop criterion for each respective metric to assess whether training should continue, training should be terminated, or a corrective action should be applied, e.g., by adjusting one or more hyperparameters related to the training. The methods disclosed herein may be executed using various computing systems. Additionally, a non-transitory computer-readable storage medium (CRM) may store instructions, which when executed by a processor, performs the inventive methods.
Need to check novelty before this filing date? Find Prior Art

Description

DYNAMIC MODE DECOMPOSITION FOR EVALUATING EARLY STOP IN TRAININGOF MACHINE LEARNING MODELSCROSS-REFERENCE TO RELATED APPLICATION(S)

[0001] This application claims the priority benefit, under 35 U.S.C. 119(e), of U.S. Application No. 63 / 599,479, filed November 15, 2023 and entitled, “DYNAMIC MODE DECOMPOSITION FOR EVALUATING EARLY STOP IN TRAINING OF LARGE LANGUAGE MODELS,” which is incorporated herein by reference in its entirety.BACKGROUND

[0002] A large language model (LLM) is a machine learning model trained to perform various natural language processing tasks, such as interpreting textual inputs or generating textual outputs. A LLM is distinguished from a conventional language model, in part, by the model having a significantly larger number of parameters (e.g., billions or trillions of parameters) and the use of a significantly larger training dataset (e.g., petabytes of training data), the combination of which gives the LLM more capability to understand natural language. Herein, the parameters are also referred to as “weights.” The weights represent the strength of interconnections formed between layers of nodes. Accordingly, the collection of weights is referred to herein as a “weight matrix.” As a result, training a LLM typically requires the use of significant computational resources at appreciable expense. For example, LLM’s are often trained using platforms (e.g., a cloud platform, a datacenter) that provide thousands of graphics processing units (GPUs) for periods spanning several months.SUMMARY[0003 J The Inventors have recognized and appreciated that training a LLM, particularly from the ground up, can often fail due to poor initialization of the parameters, an improper model configuration, or defective training data. These failures waste valuable computation time and add to the overall cost of training. The Inventors have thus recognized the ability to detect a failure early in the training of a LLM and terminate the training accordingly can appreciably reduce costs, in part, by reducing computation time spent on a process destined for failure. This may be achieved, in principle, using early stop to evaluate the likelihood of a particular training run succeeding or failing as the LLM is being trained. However, the Inventors have recognized conventional approaches of implementing early stop have several limitations that render them unsuitable for large-scale machine learning models, such as an LLM.[0004| One conventional approach involves the use of heuristics, which are predefined rules used to evaluate the quality of training. As an illustrative example, a heuristic may include comparing a validation loss computed for a particular training iteration against a predetermined threshold set by a developer. If the validation loss is greater than the threshold, then training is terminated. Heuristics are often created in an ad hoc manner based on a developer’s intuition and experience rather than any fundamental scientific principle. As a result, heuristics are seldom generalizable across different models or even different hardware (e.g., a cluster) used for training. This means heuristics are often not reliable indicators of training success or failure. Moreover, heuristics are fundamentally reactive in nature in the sense that the failure of a training run is only detected once that failure occurs. Said another way, heuristics cannot be used to predict ahead of time whether a training run will succeed or fail at future training iterations.

[0005] Another conventional approach is the use of the Calculation Consulting WeightWatcher software tool. The tool is based on random matrix theory and computes eigenvalues of a weight matrix associated with a model using singular value decomposition (SVD). This tool, however, requires the computation of an eigen spectrum for each layer of the weight matrix. The eigen spectrum is used to determine an alpha PL parameter, which can be used to assess the quality of training. Generally, if the alpha PL parameter has a value between 2 and 6, this is an indicator the training is progressing well. For a significantly large weight matrix, such as those found in LLM’s, the computation of the eigen spectrum can be computationally intensive. For instance, the WeightWatcher tool may require more than one hour to compute the eigen spectrum for a single layer within the weight matrix at a particular iteration step, thus prohibiting real time analysis while training the LLM. Additionally, the WeightWatcher tool requires numerous intermediate weight matrices to be stored as checkpoints and each checkpoint to be analyzed by the tool. This not only increases computation time, but also increases storage costs. As an illustrative example, if the WeightWatcher tool is applied to a BigScience Large Open-science Open-access Multilingual (BLOOM) language model with 176 billion parameters, each intermediate weight matrix would require 46 terabytes for storage. Similar to heuristics, the WeightWatcher tool is fundamentally reactive in nature and cannot predict whether training will succeed or fail at future training iterations.

[0006] In view of the foregoing limitations, the present disclosure is directed to various inventive methods for early stop using dynamic mode decomposition (DMD), computing systems to execute the inventive methods, and a non-transitory computer-readable storagemedium (CRM) for storing instructions, which when executed by a processor, performs the inventive methods. Specifically, DMD is used to determine a best fit linear operator that relates two sets of observability metrics (also referred to herein as “states”) associated with the training of an LLM, which may be used to predict future values of the observability metrics at future iteration steps in the training. The predicted values, in turn, may be compared against an early stop criterion for each respective metric to assess whether the training should continue or not. If the predicted values do not satisfy the early stop criteria, the training may continue. Otherwise, if the predicted values do satisfy the early stop criteria, the training may be terminated or a corrective action may be performed before continuing the training, e.g., by adjusting a hyperparameter, such as the step size of a gradient descent algorithm used to train a neural network.

[0007] In one aspect, the early stop criterion for a particular metric may be evaluated using a process capability parameter, which includes an upper limit and a lower limit defining a range of acceptable values for a particular observability metric. The combination of the process capability parameter and the predictive capabilities of DMD may provide a way to dynamically adjust and tune one or more hyperparameters during training, for example, to maintain an observability metric within acceptable limits similar to the way a production line may tune one or more manufacturing process parameters to reduce the number of defects in the manufactured product. Additionally, the process capability parameter may provide a common metric that can be used during different stages of training (e.g., pre-training, fine-tuning). In some implementations, the process capability parameter for a particular metric may have a relatively larger range of acceptable values during pre-training and a relatively smaller range of acceptable values during fine-tuning.

[0008] Compared to conventional approaches to early stop, which are typically reactive in nature, the predictive capabilities of DMD may facilitate the early termination of training runs destined for failure using appreciably fewer iterations. For example, the inventive implementations disclosed herein may predict values of an observability metric 500 to 2000 iterations in advance. If the prediction indicates a training run is likely to fail, e.g., a loss metric is likely to diverge, hundreds or even thousands of iterations worth of compute time may be saved by terminating the training run early. The savings in compute time may, in turn, be allocated to a new training run or an existing training run that is more likely to succeed. In this manner, computational resources may be more efficiently used to facilitate training of a machine learning model, such as an LLM.

[0009] Additionally, the inventive implementations of DMD disclosed herein may require appreciably little in terms of computational resources (e.g., processing power and memory storage) for execution. For example, DMD may only evaluate observability metrics rather than the weight matrix of the model itself when assessing the quality of training. Accordingly, the inventive implementations of DMD disclosed herein may be executed using a central processing unit (CPU) rather than a graphics processing unit (GPU). For example, a GPU cluster typically includes multiple GPUs communicatively coupled to a single CPU (i.e., the host). By using the CPU to execute DMD, the pool of GPU-related resources available to train the model is increased, i.e., since none of the GPUs need to be used to execute DMD. Furthermore, the deployment of DMD onto the CPU may ensure any outgoing data transmitted from the CPU to the GPUs or any incoming data transmitted from the CPUs to the CPU is readily accessible for use as an observability metric when executing DMD.

[0010] In one example implementation, a computing system for evaluating training of a machine learning model includes: a plurality of graphics processing units (GPUs) to train the machine learning model; a central processing unit (CPU), communicatively coupled to each GPU of the plurality of GPUs, to receive data associated with one or more metrics related to the training of the machine learning model from the plurality of GPUs; and a memory, communicatively coupled to the CPU, storing instructions. When the instructions are executed by the CPU during training of the machine learning model, it causes the CPU to: A) acquire a first state, X, containing values for the one or more metrics spanning iteration steps i through i+m-1 of the training wherein z and m are positive integer values; B) acquire a second state, X, containing values for the one or more metrics spanning iteration steps i+z through i+m-1 +z of the training wherein z is a positive integer value; C) determine a best fit linear operator, A, that relates X to X according to a first relation: X ~ AX D) determine a future state, X", using A, the future state containing values for the one or more metrics at an iteration step i+n wherein n is a positive integer value greater than m-l+z and E) evaluate the future state against an early stop criterion.

[0011] The plurality of GPUs may be a first plurality of GPUs, the data may be first data, and the system may further include: a second plurality of GPUs, communicatively coupled to the first plurality of GPUs and the CPU, to receive second data associated with the one or more metrics, generate values for at least one metric of the one or more metrics, and transmit the values for the at least one metric to the CPU. The at least one metric may include an alpha PL parameter (e.g., obtained using the WeightWatcher tool). The one or more metrics may includeat least one of a cross entropy, a gradient norm, a perplexity, a learning rate, a validation loss, a BiLingual Evaluation Understudy (BLEU) score, a Recall-Oriented Understudy for Gisting Evaluation (ROUGE) score, a Fl score, precision, recall, an Area Under the Curve Receiver Operating Characteristics (AUC-ROC) score, an Area Under the Curve Precision Recall Curve (AUC-PRC) score, or an alpha PL parameter. The number of metrics included in each of the first state and the second state may range from 1 to 1000.[00.12] The first state may contain values for the one or more metrics at every iteration step spanning steps i through i+m-1. The first state may contain values for the one or more metrics only for a subset of iteration steps spanning steps i through i+m-1. The subset of iteration steps may include iteration steps at equal increments between the steps i and i+m-1 of the training. The subset of iteration steps may include iteration steps at nonequal increments between the steps z and i+m-1 of the training. The number of iteration steps spanning steps z through i+m- 1 may range from 100 to 2000. The parameter z may range from 1 to 100. The second state, X, may contain at least one value for the one or more metrics interpolated from the data received by the CPU. The data associated with the one or more metrics may include data at iteration steps ii and Z2 greater than zz; and the instructions cause the CPU to, when acquiring the second state: interpolate the data at iteration steps ii and Z2 to determine an estimate of the data at an iteration step is between ii and is for inclusion in the second state.

[0013] The instructions may further cause the CPU to, when determining the best fit linear operator, A: determine A according to a second relation: A = XXA wherein XA is a Moore- Penrose generalized inverse of X. The instructions may further cause the CPU to, when determining the best fit linear operator, A: determine a singular value decomposition of X according to a second relation: X~ UXJX wherein Uis a matrix having columns corresponding to left-singular vectors, 27 is a matrix containing singular values, andis a matrix having columns corresponding to right-singular vectors; determine a projection, A, of A according to a third relation: A = AX'V '1wherein IX is a conjugate transpose of U, V is a conjugate transpose of W, and 27_;is an inverse of 27; execute an eigendecomposition of A according to a fourth relation: AW = WA wherein is a matrix having columns corresponding to eigenvectors and / I is a matrix containing eigenvalues; and reconstruct an eigenvector, , of A according to a fifth relation: A> X'VX1W.[00.14] The difference between n and m-l+z may range from 1 to 2000. The difference between n and m-l+z may equal to 500. The difference between n and m-l+z may equal to 1000. The difference between n and m-l+z may equal to 2000.[0015| The instructions may cause the CPU to, when evaluating the future state against the early stop criterion: determine, for a metric of the one or more metrics, a running mean, «, and a standard deviation, <r, associated with the metric based on the values of the metric in at least one of the first state or the second state; and determine whether a value of the metric in the future state is one of greater than or equal to u 3o or less than or equal to u-3o. The metric may be a gradient norm. The instructions may cause the CPU to, when evaluating the future state against the early stop criterion: determine, for a metric of the one or more metrics, a running mean,and a standard deviation, <7, associated with the metric based on the values of the metric in the future state, the first state, and the second state; determine a CPK value of the metric according to a second relation: wherein UL is a predetermined upper limit and LL is a predetermined lower limit; and determine whether the CPK value is below a threshold. The threshold may be equal to 1.33. The instructions may further cause the CPU to: in response to the CPK value being greater than or equal to the threshold, allow the training of the machine learning model to continue; and in response to the CPK value being less than the threshold, terminate the training of the machine learning model or apply a corrective action to the training of the machine learning model. The threshold may be a first threshold and the instructions may further cause the CPU to, in response to the CPK value being less than the first threshold: determine whether the CPK value is below a second threshold less than the first threshold; in response to the CPK value being greater than or equal to the second threshold, apply a corrective action to the training of the machine learning model; and in response to the CPK value being less than the second threshold, terminate the training of the machine learning model.

[0016] The instructions may further cause the CPU to: in response to determining the future state does not satisfy the early stop criterion, allow the training of the machine learning model to continue. The instructions may further cause the CPU to: in response to determining the future state satisfies the early stop criterion, terminate the training of the machine learning model. The instructions may cause the CPU to, when terminating the training of the machine learning model: generate a signal with instructions for the plurality of GPUs to stop training the machine learning model; and transmit the signal to the plurality of GPUs for execution. The instructions may further cause the CPU to: in response to determining the future state satisfies the early stop criterion, apply a corrective action to the training of the machine learning model. The instructions may cause the CPU to, when applying the corrective action to the training of the machine learning model: generate a signal with instructions for the plurality of GPUs to apply the corrective action; and transmit the signal to the plurality of GPUs for execution. Thecorrective action may include an adjustment to one or more hyperparameters associated with the training of the machine learning model. The one or more hyperparameters may include at least one of a learning rate, a batch size, or a temperature.10017] The system may be a GPU cluster. The system may be configured to execute A) through E) in less than 1 second.

[0018] In another example implementation, a method for evaluating training of a machine learning model using a computing system where the computing system includes: a plurality of graphics processing units (GPUs); and a central processing unit (CPU), communicatively coupled to each GPU of the plurality of GPUs, includes the following steps: A) acquiring a first state, X, containing values for one or more metrics related to the training of the machine learning model from the plurality of GPUs spanning iteration steps i through i+m-1 of the training wherein z and m are positive integer values; B) acquiring a second state, X, containing values for the one or more metrics spanning iteration steps i+z through i+m-1 +z of the training wherein z is a positive integer value; C) determining a best fit linear operator, A, that relates X to X according to a first relation: X ~ AX D) determining a future state, X", using A, the future state containing values for the one or more metrics at an iteration step i+n wherein n is a positive integer value greater than m-l z., and E) evaluating the future state against an early stop criterion.

[0019] The method may further include, before step A): training, by the plurality of GPUs, the machine learning model; and receiving, by the CPU from the plurality of GPUs, data associated with the one or more metrics.

[0020] In yet another example implementation, a non-transitory computer-readable storage medium storing instructions for evaluating a machine learning model, the instructions, when executed, cause at least one processor of a central processing unit (CPU) communicatively coupled to a plurality of graphics processing units (GPUs) to, while the plurality of GPUs is training the machine learning model: A) acquire a first state, X, containing values for one or more metrics related to the training of the machine learning model from the plurality of GPUs spanning iteration steps z through i+m-1 of the training wherein z and m are positive integer values; B) acquire a second state, X, containing values for the one or more metrics spanning iteration steps i+z through i+m-l+z of the training wherein z is a positive integer value; C) determine a best fit linear operator, A, that relates X to X according to a first relation: X ~ AX D) determine a future state, X", using A, the future state containing values for the one or moremetrics at an iteration step i+n wherein n is a positive integer value greater than m-l+z and E) evaluate the future state against an early stop criterion.

[0021] In yet another inventive implementation, a method for evaluating early stop criteria during training of a large language model (LLM) includes: A) acquiring a first state, X associated with the training of the LLM where the first state is a matrix containing values for a set of metrics associated with the training of the LLM spanning iteration steps i through i+m- 1 of the training with z and m being integer values; B) acquiring a second state, X', associated with the training of the LLM where the second state is a matrix containing values for the set of metrics spanning iteration steps z+1 through i+m of the training; C) determining a best fit linear operator, A, that relates X to X according to a first relation: X' ~ AX,' D) determining a future state, X", associated with the training of the LLM using A where the future state is a matrix containing values for the set of metrics at an iteration step i+n wherein n is an integer value greater than m; and E) evaluating the future state against the early stop criteria.

[0022] The step of determining the best fit linear operator, A, may include: determining A according to a second relation: A = X'XA wherein X is a Moore-Penrose generalized inverse of[0023J The step of determining the best fit linear operator, A, may include: determining a singular value decomposition of X according to a second relation: X ~ UXW wherein U is a matrix having columns corresponding to left-singular vectors, X is a matrix containing singular values, and l!is a matrix having columns corresponding to right-singular vectors; determining a projection, A, of A according to a third relation: A = l AX'V ^ wherein IX is a conjugate transpose of U, V is a conjugate transpose of IX and X'' is an inverse of X; executing an eigendecomposition of A according to a fourth relation: AW = WA wherein Wis a matrix having columns corresponding to eigenvectors and A is a matrix containing eigenvalues; and reconstructing an eigenvector, , of A according to a fifth relation: X>=X'VX~IW.[0024| The set of metrics may include at least one of a cross entropy, a gradient norm, a perplexity, a learning rate, a validation loss, a BiLingual Evaluation Understudy (BLEU) score, a Recall-Oriented Understudy for Gisting Evaluation (ROUGE) score, a Fl score, precision, recall, an Area Under the Curve Receiver Operating Characteristics (AUC-ROC) score, or an Area Under the Curve Precision Recall Curve (AUC-PRC) score.[0025| The step of evaluating the future state against the early stop criterion may include: determining, for a metric of the set of metrics, a running mean,and a standard deviation, u,associated with the metric based on the values of the metric in at least one of the first state or the second state; and determining whether a value of the metric in the future state is one of greater than or equal to u 3o or less than or equal to u-3o. The metric may be a gradient norm. The step of evaluating the future state against the early stop criterion includes: determining, for a metric of the set of metrics, a running mean,and a standard deviation, cr, associated with the metric based on the values of the metric in the future state, the first state, and the second state; determining a CPK value of the metric according to a second relation: wherein UL is a predetermined upper limit and LL is a predetermined lower limit; and determining whether the CPK value is below a predetermined threshold. The predetermined threshold may equal to 1.33.

[0026] The method may further include: in response to determining the future state does not satisfy the early stop criterion, continuing training of the LLM. The method may further include: in response to determining the future state satisfies the early stop criterion, terminating training of the LLM.

[0027] In yet another inventive implementation, a non-transitory computer-readable storage medium stores instructions for evaluating early stop criteria during training of a large language model (LLM), the instructions, when executed, cause at least one processor to: A) acquire a first state, A, associated with the training of the LLM where the first state is a matrix containing values for a set of metrics associated with the training of the LLM spanning iteration steps i through i+m-1 of the training with z and m being integer values; B) acquire a second state, X, associated with the training of the LLM where the second state is a matrix containing values for the set of metrics spanning iteration steps z+1 through i+m of the training; C) determine a best fit linear operator, A, that relates X to X' according to a first relation: X ~ AX, D) determine a future state, X", associated with the training of the LLM using A, the future state being a matrix containing values for the set of metrics at an iteration step i+n wherein n is an integer value greater than m; and E) evaluate the future state against the early stop criteria.

[0028] It should be appreciated that all combinations of the foregoing concepts and additional concepts discussed in greater detail below (provided such concepts are not mutually inconsistent) are contemplated as being part of the inventive subject matter disclosed herein. In particular, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as being part of the inventive subject matter disclosed herein. It should also be appreciated that terminology explicitly employed herein that also may appear in any disclosure incorporated by reference should be accorded a meaning most consistent with the particular concepts disclosed herein.BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The skilled artisan will understand that the drawings primarily are for illustrative purposes and are not intended to limit the scope of the inventive subject matter described herein. The drawings are not necessarily to scale; in some instances, various aspects of the inventive subject matter disclosed herein may be shown exaggerated or enlarged in the drawings to facilitate an understanding of different features. In the drawings, like reference characters generally refer to like features (e.g., functionally similar and / or structurally similar elements).

[0030] FIG. 1 shows an illustration of a dynamic mode decomposition (DMD) algorithm applied to forecast future dynamics of a fluid flow.

[0031] FIG. 2 shows a flow chart of an example method for evaluating early stop during training of a large language model (LLM).

[0032] FIG. 3 shows an example DMD algorithm for determining a best-fit linear operator d.

[0033] FIG. 4A shows an example state, X, containing values for observability metrics acquired over multiple iteration steps.

[0034] FIG. 4B shows an example state, X containing values for observability metrics acquired over multiple iteration steps.

[0035] FIG. 5 shows a plot of example raw data for a mean loss metric as a function of iteration step.

[0036] FIG. 6A shows a plot of example data for a metric as a function of iteration step. The additional lines represent an upper limit (UL), a lower limit (LL), and a centerline (CL) for determining a CPK value associated with the data for the metric.

[0037] FIG. 6B shows a plot of the CPK value corresponding to the data in FIG. 6A as a function of iteration step.

[0038] FIG. 6C shows a plot of example CPK values and different upper and lower limits during different stages of training (pre-training and fine-tuning).

[0039] FIG. 7A shows an example runtime monitor using DMD.

[0040] FIG. 7B shows an example architecture for a GPU cluster configured to support the runtime monitor of FIG. 7 A.DETAILED DESCRIPTION

[0041] Following below are more detailed descriptions of various concepts related to, and implementations of methods for evaluating early stop during training of a large language model (LLM) using dynamic mode decomposition (DMD), systems (e.g., a computer, a GPU cluster, a datacenter) configured to execute the methods, and / or a non-transitory computer-readable storage medium (CRM) for storing instructions, which when executed by a processor, performs the methods. It should be appreciated that various concepts introduced above and discussed in greater detail below may be implemented in multiple ways. Examples of specific implementations and applications are provided primarily for illustrative purposes so as to enable those skilled in the art to practice the implementations and alternatives apparent to those skilled in the art.

[0042] The figures and example implementations described below are not meant to limit the scope of the present implementations to a single embodiment. Other implementations are possible by way of interchange of some or all of the described or illustrated elements. Moreover, where certain elements of the disclosed example implementations may be partially or fully implemented using known components, in some instances only those portions of such known components that are necessary for an understanding of the present implementations are described, and detailed descriptions of other portions of such known components are omitted so as not to obscure the present implementations.

[0043] In the discussion below, various examples of methods for evaluating early stop are provided using DMD, wherein a given example or set of examples showcases one or more observability metrics for evaluating the training of a LLM (or, more generally, any machine learning model), methods for determining a best-fit linear operator, and methods for tracking changes in the observability metrics using process capability parameters. It should be appreciated that one or more features discussed in connection with a given example method for evaluating early stop may be employed in other example methods, according to the present disclosure, such that the various steps and / or features disclosed herein may be readily combined in a given method according to the present disclosure (provided that respective features are not mutually inconsistent).

[0044] It should be appreciated that the methods disclosed herein may be executed by one or more processors of a computing system. For example, the computing system may include memory storing one or more instructions associated with the methods, which when executedby the processor(s), cause the processor(s) to execute the method. For example, the electronic device may be a server (e.g., a compute node, a GPU cluster) in a data center or a cloud platform. In another example, the electronic device may be a personal computer, such as a desktop or a laptop. It should also be appreciated that the methods disclosed herein may be stored as instruction(s) in a non-transitory computer-readable storage medium. When the instruction(s) are executed by a computer, the methods disclosed herein are performed. The non-transitory computer-readable storage medium may be digitally or physically distributed to one or more computers. l. Overview of Application of Dynamic Mode Decomposition (DMD) for Training Evaluation

[0045] Dynamic mode decomposition (DMD) is a dimensionality reduction technique that can be used to extrapolate values in a sequential dataset containing non-linear dynamics. This is accomplished, in part, by determining a set of characteristic modes within the dataset and using the collective behavior of these modes to predict values that extend beyond the dataset. The modes may include, but are not limited to, a sinusoid, an exponential decay, and an exponential growth.

[0046] As an illustrative example, FIG. 1 shows the use of DMD to forecast the dynamics of fluid flow. FIG. 1 is reproduced from Kutz et al., “Dynamic Mode Decomposition: Data- Driven Modeling of Complex Systems,” Society for Industrial and Applied Mathematics (SIAM) (November 23, 2016) (hereafter referred to as Kutz), which is incorporated herein by reference in its entirety. In this example, the sequential dataset is a sequential time-series flow field. As shown, an initial dataset of the fluid flow is acquired spanning time steps 1 through m. To predict the fluid flow at future time steps, a regression analysis is performed. First, the dataset is divided between two states: X spanning time steps 1 through m-1; and ” spanning time steps 2 through m. Second, a best-fit linear operator A is found to temporally advance from states A to A’. DMD may determine the characteristic modes of A (e.g., the eigenmodes) without directly computing A. The modes of A, in turn, provide a basis from which the fluid flow at future time steps (i.e., time steps greater than m) may be estimated.

[0047] In the present disclosure, DMD is applied to a sequential dataset associated with the training of a LLM. It should be appreciated that while the inventive implementations of DMD disclosed herein are described in application to the training of a LLM, these inventive implementations are not limited in application only to LLM’s. The inventive implementationsof DMD may be applied during training of any generative artificial intelligence model capable of generating, for example, text, imagery, audio, video, or any combination of the foregoing outputs. More generally, the inventive implementations of DMD may be applied during training of various deep learning models including, but not limited to, a LLM, a diffusion model, a normalizing flow model, a generative adversarial network (GAN) model, and a variational autoencoder (VAE) model. Even more generally, the inventive implementations of DMD disclosed herein may be applied to the training of any machine learning model (e.g., a supervised machine learning model, an unsupervised machine learning model, a semisupervised machine learning model) including, but not limited to, a linear regression model, a logistic regression model, a support vector machine, an artificial neural network (e.g., a deep neural network), a K-means clustering model, and the like.

[0048] Additionally, the inventive implementations of DMD disclosed herein may be applied at any stage of training a LLM or, more generally, any machine learning model. These training stages may include, but are not limited to, pre-training and fine-tuning. In particular, the inventive implementations of DMD disclosed herein may be applied while executing various fine-tuning techniques including, but not limited to, supervised fine-tuning, transfer learning, reinforcement learning from human feedback (RLHF), direct preference optimization (DPO), and the like.[0049| The sequential dataset considered herein may include one or more observability metrics that vary as a function of the iteration step for the training. The observability metrics each provide information on the progress of the training, such as the rate that an objective function is converging and / or the precision and / or the accuracy of the model. DMD may be used to determine the best-fit linear operator for this sequential dataset and, in turn, identify the characteristic modes of each observability metric as it varies as a function of iteration step. The dynamics of these modes may provide a basis to predict values for each observability metric at future iteration steps. These predictions may then be compared against an early stop criterion for that metric to assess whether the training should continue, be terminated, or a corrective action performed.[0050 | Compared to conventional approaches of implementing early stop, such as heuristics or the WeightWatcher tool, the inventive implementations of DMD disclosed herein offer several advantages.[00511 First, the inventive implementations of DMD may be fully data driven. In other words, the inventive implementations of DMD do not require any knowledge of the underlying model dynamics, but instead may rely only on outputs generated when training the model (e.g., the observability metrics). Generally, the observability metrics constitute an appreciably smaller set of data compared to the weight matrix of the model. Thus, the inventive implementations of DMD disclosed herein may require appreciably less memory for storage and appreciably less processing for execution unlike the WeightWatcher tool and some heuristics that require storage and processing of the weight matrix. In some implementations, the processing and storage requirements to execute DMD may be further reduced by configuring DMD to track the evolution of more dominant eigenvalues in the best-fit linear operator, A.

[0052] The relatively low processing and storage requirements associated with the inventive implementations of DMD disclosed herein may, in turn, allow a training run to be evaluated more quickly compared to conventional approaches. For example, the inventive implementations of DMD disclosed herein may readily evaluate training, e.g., by providing predictions of observability metrics at future iterations and evaluating the predictions against various early stop criteria, in real time or near real time (e.g., less than 1 second).

[0053] Additionally, the inventive implementation of DMD may be readily deployed using a central processing unit (CPU) and its associated memory rather than a graphics processing unit (GPU) and its associated memory. For example, a GPU cluster typically includes multiple GPUs communicatively coupled to a single CPU (i.e., the host). By using the CPU to execute DMD, the pool of GPU-related resources available to train the model is increased, i.e., since none of the GPUs need to be used to execute DMD. Furthermore, the deployment of DMD onto the CPU may ensure any outgoing data transmitted from the CPU to the GPUs or any incoming data transmitted from the CPUs to the CPU is readily accessible for use as an observability metric when executing DMD.

[0054] Second, the inventive implementations of DMD are generalizable to any machine learning model executed using any computing system unlike heuristics. The inventive implementations of DMD disclosed herein are used to evaluate one or more observability metrics associated with the training of the machine learning model against various early stop criteria. Although different observability metrics and / or early stop criteria may be used between different models, the implementation of DMD may remain the same.[0055} Third, the inventive implementations of DMD disclosed herein are predictive in the sense that future values of the observability metrics may be extrapolated from existing data using the characteristic modes obtained by DMD. In contrast, conventional approaches including both the WeightWatcher tool and heuristics, are typically descriptive and, hence, limited to evaluate present or past values of the observability metrics in a reactive manner. Thus, the predictive capabilities of DMD as used herein may facilitate the termination of training runs likely to fail using appreciably fewer iterations compared to conventional approaches. For example, the inventive implementations disclosed herein may predict values of an observability metric 500 to 2000 iterations in advance. If the prediction indicates a training run is likely to fail, e.g., a loss metric is likely to diverge, hundreds or even thousands of iterations worth of compute time may be saved by terminating the training run early. The savings in compute time may, in turn, be allocated to a new training run or an existing training run that is more likely to succeed. In this manner, computational resources may be more efficiently used to facilitate training of a machine learning model, such as an LLM.2. Example Methods for Evaluating Early Stop using DMD(0056} FIG. 2 shows an example method 100 for evaluating early stop according to the inventive concepts disclosed herein. As shown, the method 100 begins at step 102 by acquiring a first state, X. The first state, A, may be a matrix that contains values for one or more observability metrics associated with the training of a LLM spanning iteration steps i through i+m-1 of the training where i and m are positive integer values (see, for example, FIG. 4A).[0057| At step 104, a second state, X’, is acquired. The second state, X’, may be a matrix containing values for the observability metric(s) in the first state, A, spanning iteration steps z+7 through i+m of the training (see, for example, FIG. 4B). In other words, the second state, A’, may include data incremented one iteration step forward with respect to the first state, A. It should be appreciated that, in some implementations, the data in the second state, A’, may include data incremented several iteration steps forward with respect to the first state, A. Generally, the second state, A’, may contain data spanning iteration steps i+z through i+m-1 +z where z represents the increment forward with respect to the data in the first state, A. In one non-limiting example, z may be a positive integer value ranging from 1 to 100, including all sub-ranges and values in between. More generally, the value of z may be limited by the manner in which the metrics vary as a function of iteration step. For example, if the value of the metric varies linearly as a function of iteration step, the value of z may be greater since the rate of change of the value is approximately constant. In another example, if the value of the metricvaries nonlinearly as a function of iteration step, the value of z may be smaller to capture the nonlinear variation of the metric.

[0058] At step 106, a best fit linear operator, A, may be determined directly or indirectly (see Section 2.2 for further details). The best fit linear operator, A, relates X to X’ according to the following relation:X’ ~AX (1)

[0059] At step 108, a future state, X”, may be determined based on the best fit linear operator, A. The future state, X”, may be a matrix containing values for the observability metric(s) in the first state, X, predicted for at least one iteration step i+n where n is an integer value greater than m. In some implementations, the future state, X”, may contain values spanning iteration steps i+m through i+n2 where m and ri2 are each integer values, m is greater than m, and ri2 is greater than m. The foregoing assumes z is equal to 1. More generally, n may be an integer value greater than m-l+z.

[0060] At step 110, the future state, X”, may be evaluated against an early stop criterion. In some implementations, if the future state, X”, includes multiple observability metrics, an early stop criterion may be defined for each metric (see Section 2.3 for additional details).

[0061] The evaluation of the future state, X”, against the early stop criterion may ultimately lead to one of several outcomes. If the predicted values in the future state, X”, does not satisfy the early stop criterion, this indicates that training is progressing in a manner that is acceptable. Accordingly, the process of training the LLM may continue. If the predicted values in the future state, X”, does satisfy the early stop criterion, this indicates that training is likely to fail. In some implementations, the training may be terminated, thus avoiding further consumption of valuable computation time for a process destined to fail. In some implementations, a corrective action may be performed, such as adjusting a parameter in an optimization algorithm used to minimize a cost function of the LLM. For example, the corrective action may include adjusting the step size of a gradient descent algorithm used to train a neural network. Training may continue thereafter using the adjusted optimization algorithm. More generally, the correction action may include adjusting one or more hyperparameters associated with training including, but not limited to, a learning rate, a batch size, and a temperature (i.e., a parameter affecting the randomness and / or creativity of the output generated by the machine learning model).

[0062] In some implementations, the choice between allowing a training run to continue, terminating a training run, or applying a corrective action to improve the quality of trainingmay depend on the predicted values in the future state, X” . For example, if the observability metric is a training loss or a validation loss, i.e., the error of the model in reproducing data in a training dataset or a validation dataset, respectively, the magnitude of the loss may be compared against multiple thresholds. If, for example, the loss is greater than a first threshold, the training run may be terminated. If the loss is less than or equal to the first threshold and greater than a second threshold, a corrective action may be applied. If the loss is less than or equal to the second threshold, the training run may be allowed to continue without any corrective action.

[0063] It should be appreciated that, in some implementations, a subset of the actions described may be implemented when evaluating the observability metric(s) against the early stop criterion. For example, the outcomes may be limited to either terminating a training run or applying a corrective action. Referring to the above example where the observability metric is either a training loss or a validation loss, the metric may be compared against a threshold. If the loss is greater than the threshold, the training run may be terminated. If the loss is less than or equal to the threshold, a corrective action may be applied. In this manner, a training run may be continually improved through adjustments of one or more hyperparameters until either the training run is complete or the predicted loss is so great (e.g., greater than the threshold) to trigger early termination.

[0064] It should also be appreciated that the manner in which a metric is compared against a threshold is not limited to the foregoing example. In another example, the observability metric may be represented as a process capability parameter (see section 2.3). Generally, higher values of the process capability parameter indicate closer alignment to acceptable bounds. Accordingly, in some implementations, if the process capability parameter is less than a first threshold, the training run may be terminated. If the process capability parameter is greater than or equal to the first threshold and less than a second threshold, a corrective action may be applied. If the process capability parameter is greater than or equal to the second threshold, the training run may be allowed to continue without any corrective action.

[0065] In some implementations, the thresholds described above may each be predetermined values, i.e., values set by a user. In some implementations, the thresholds may correspond to a desired statistical significance. For example, a threshold for a process capability parameter may be equal to 1.33 such that the smallest difference between the running mean, p, and the lower or upper limits is equal to 4c (see Section 2.3 for further details on the formulation of the process capability parameter).

[0066] The accuracy of the predicted values in the future state, X’ at larger values of n (i.e., farther into the future) depends on several factors. In one example, the amount of data included in states X and X’ and used to determine the best-fit linear operator, A, may influence the accuracy of the predicted values in X” with the collection of more data often leading to more accurate predictions further into the future. In another example, the number of observability metrics may influence the accuracy of the predicted values in X” particularly if the observability metrics are interdependent. The interdependency between any metrics may provide additional information on the training of a particular model, thus leading to more accurate predictions further into the future.

[0067] It should be appreciated that the method 100 may be executed concurrently with the training of a LLM or, more generally, any machine learning model. It should also be appreciated the method 100 may be performed multiple times during the training of a LLM. For example, training may initially progress from iteration steps 1 to mi. The method 100 may then be executed for a first time on the data spanning iteration steps 1 through mi. If the future state, X”, does not satisfy the early stop criterion, training may continue from iteration steps mi to m2. The method 100 may then be executed for a second time on the data spanning either iteration steps mi through m2 or iteration steps 1 through m2. This process may continue until either training is completed or the training is terminated prematurely, i.e., because the future state, X”, satisfies the early stop criterion.

[0068] In one aspect, the methods disclosed herein may provide appreciable cost savings by reducing the amount of wasted computation time consumed during training. As an illustrative example, Bloomberg’s BloombergGPT is a 50 billion parameter specialized LLM model that was originally intended to be trained using 700 billion tokens. Bloomberg allocated 1.9 million GPU hours at a budget of about $1.9 million. During the training phase of the model, many experimental training runs failed to converge, thus consuming part of the budget. As a result, Bloomberg was only able to train around 75% of the 700 billion tokens. This means that out of the $1.9 million budget, about $450,000 was spent on training. The methods disclosed herein for evaluating early stop may appreciably reduce the cost of training by providing a way to terminate training runs early if there are early indications from the observability metrics that convergence is unlikely to occur.

[0069] In another aspect, the methods disclosed herein may be incorporated into a software diagnostic tool (see Section 3). The tool may aid users in characterizing the quality of training for their model and / or executing an early stop in the training if the tool determines convergencefor a particular training run is unlikely to occur. For example, the tool may visually display (e.g., via a plot) the dynamic modes obtained from DMD as signature spectral graphs (see, for example, FIG. 1). The graphs may further be updated in real time or near real time (e.g., less than 1 second) to show how the predicted metrics evolve as the training progresses. In some implementations, the tool may be provided to users (e.g., customers) as a Training as a Service (TaaS).2.1 States X and X’

[0070] The states X and X’ may generally include values for one or more observability metrics associated with the training of the LLM spanning one or more iteration steps. FIGS. 4A and 4B show examples of the first and second states Xand X’. As shown, the first and second states X andX’ include the same observability metrics (e.g., at least 13 metrics in total) and span the same number of iteration steps (e.g., 100 iteration steps) with the difference being X covers iteration steps 1 through m-1 and X’ covers iteration steps 2 through m. In other words, the iteration steps covered by X’ are incremented forward by a single iteration step with respect to X. It should be appreciated that the future state, X”, may also be similarly formatted. For example, referring to the example states X and X’ in FIGS. 4 A and 4B, the future state X” may also include the same observability metrics as X and X’ (e.g., at least 13 metrics in total) spanning one or more future iteration steps.

[0071] The number of iteration steps into the future at which values for observability metrics are predicted using the inventive implementations disclosed herein generally depends on several factors. For example, the size of the model may appreciably influence the noise and variation of certain metrics. In particular, larger models with a higher number of weights are often less noisy and better behaved, thus allowing more reliable predictions further into the future. In another example, the number of observability metrics used in the states X, X’, and / or X” may appreciably influence the predictions on the quality of training. For some models, a larger number of observability metrics may provide a more accurate prediction on the quality of training. As an illustrative non-limiting example, the number of iteration steps into the future that the inventive implementations disclosed herein are used for predictions of one or more observability metrics may range from 1 iteration step to 2000 iteration steps, including all subranges and values in between. This includes, for example, 500 iteration steps into the future, 1000 iteration steps into the future, 1500 iteration steps into the future, and / or 2000 iteration steps into the future. Herein, the number of iteration steps into the future may be determinedwith respect to the most recent iteration step during training. In some implementations, this number may be equal to the difference of n and m-l+z.

[0072] FIGS. 4A and 4B show the data associated with the observability metrics may be recorded and stored for every iteration step. However, it should be appreciated that this is a non-limiting example. For some observability metrics, the data associated with these metrics may be recorded and stored (e.g., in memory) periodically during training. In other words, data may not be available at each iteration step. When the data associated with these metrics is stored for a particular iteration step, this may constitute a checkpoint at that iteration step. Generally, the number of checkpoints in states X and X’ may vary in number depending on, for example, the noise level of the observability metric, the type of observability metric used, the architecture of the model, and so on. In some instances, the number of checkpoints in states X and X’ may be appreciably smaller in number than shown in FIGS. 4 A and 4B. In some implementations, the number of checkpoints and the iteration steps at which checkpoints occur may be predetermined. As an illustrative non-limiting example, the number of checkpoints in states A and A’ may range from 2 to 1000, including all sub-ranges and values in between. For example, the number of checkpoints in states A and A’ may be equal to 3, 100, or 1000.

[0073] In some implementations, checkpoints may be stored every 10, 100, 200, or 500 iteration steps. As an illustrative non-limiting example, the number of iteration steps between checkpoints may range from about 1 iteration step to 1000 iteration steps, including all subranges and values in between. In some implementations, checkpoints may be stored at equal increments (e.g., every 100 iteration steps). In some implementations, checkpoints may be stored at nonequal increments. For example, the number of checkpoints stored at the beginning of training may be greater than the number of checkpoints stored later in training due to greater volatility in the observability metrics at the start of training. As an illustrative example, FIG. 5 shows example data for a mean loss observability metric as a function of iteration step collected during training of a nanoGPT model with 300,000 parameters. As shown, there may be multiple checkpoints during the training, for example, at iteration steps 20, 200, and 1000.

[0074] In implementations where checkpoints are stored incrementally and the states A and A’ only contain data from the checkpoints, the data in A’ may correspond to iteration steps that are at an undesirably large increment with respect to the state A (i.e., the increment z may be too large). To address this issue, interpolation may be used to estimate the values of metrics included in A’ so that the increment z may be reduced. As an illustrative non-limiting example, checkpoints may be stored at iteration steps ii = 500 and 1 = 1000. When applying DMD, itmay be desirable for the data stored in A to correspond to the iteration step ii and the data stored in X’ to correspond to an iteration step is = 600. The data for X may be taken from the checkpoint at ii. The data for X’ may be obtained by interpolating the data at iteration steps ii and is using, for instance, linear regression. It should be appreciated that various interpolation techniques may be used to estimate data for various observability metrics at various iteration steps including, but not limited to, linear regression, polynomial regression, spline interpolation, and the like.

[0075] Additionally, the range of iteration steps covered by the first and second states X and X’ may generally vary on a case-by-case basis depending on the particular model being trained. Notably, it should be appreciated that the range of iteration steps covered by the states X and X’ may not equal the number of checkpoints included in the X and X’ since the checkpoints may only be stored incrementally as described above. In some implementations, the range of iteration steps covered may range from 100 iteration steps to 2000 iteration steps, including all sub-ranges and values in between. In one non-limiting example, the iteration steps included in X and X’ may range from step 100 to step 1000.

[0076] It should be appreciated that the metrics shown in FIGS. 4A and 4B are non-limiting examples. More generally, the states A, X and / or X” may include any number of observability metrics. As an illustrative non-limiting example, the number of observability metrics may range from 1 to 1000, including all sub-ranges and values in between. The observability metrics may generally include any metric relevant to the training of a machine learning model and / or the operation of the computing system used for training. The observability metrics in the states A, A’, and A” related to training may include, but is not limited to, a cross entropy, a gradient norm, a perplexity, a learning rate, a validation loss, a BiLingual Evaluation Understudy (BLEU) score, a Recall-Oriented Understudy for Gisting Evaluation (ROUGE) score, a Fl score, precision, recall, an Area Under the Curve Receiver Operating Characteristics (AUC- ROC) score, an Area Under the Curve Precision Recall Curve (AUC-PRC) score, and an alpha PL parameter representing the shape of the eigenvalue distribution of the weight matrix determined using, for example, the WeightWatcher tool. The observability metrics in the states A, A’, and A” related to the operation of the computing system may include, but is not limited to, network bandwidth, tensor core utilization, a parameter related to the input / output stress of a processor, a parameter related to the health of the hardware, and the like.

[0077] As described above, conventional approaches to early stop, such as the WeightWatcher tool, can be too computationally intensive to use on its own, especially for large models.However, it should be appreciated that this does not necessarily mean such tools are precluded from use in combination with the inventive implementations of DMD disclosed herein. In some implementations, the WeightWatcher tool may provide an observability metric, e.g., the alpha PL parameter, that, in turn, is included in the states X and X’ during execution of the method 100. Although computing the alpha PL parameter may be computationally intensive, the number of times the alpha PL parameter is computed may be appreciably reduced than if the WeightWatcher tool is used on its own for evaluating early stop. Moreover, if the method 100 determines a training run is destined to fail, early termination of the training run may save compute time that would otherwise be spent computing the alpha PL parameter at later iteration steps. More generally, some observability metrics may require additional processing to obtain (see, for example, the data source 202 in FIG. 7A). However, the combination of the predictive capabilities of DMD and the use of relatively few checkpoints containing these observability metrics may make them viable for use in evaluating early stop in real time or near real time (e.g., less than 1 second).2.2 Determination of the Best-Fit Linear Operator, A[(>078) The best-fit linear operator, A, may be determined in several ways. In one example, the best-fit linear operator, A, may be determined directly without any approximation. This may be accomplished using the following relation:A = X (2) where X is the Moore-Penrose generalized inverse of state X, i.e., X = (X X)~]X The direct computation of A may be possible in applications where the size of X and X’ is relatively small such that direct computation does not require significant computation time.

[0079] In another example, the best-fit linear operator, A, may be approximated using method 106a shown in FIG. 3. At step 112, the singular value decomposition (SVD) of Xis determined using the following relation:X~ U V1(3) where U is a matrix having columns corresponding to left-singular vectors, X is a matrix containing singular values, and X is a matrix having columns corresponding to right-singular vectors.

[0080] At step 114, A is approximated asH by determining the r r projection of 4 onto U using the following relation:A = I AU = If'X’V (4) where lr is a conjugate transpose of U, Kis a conjugate transpose of JX, and IP1is an inverse of .

[0081] At step 116, an eigendecomposition of A is performed using the following relation:AW = WA (5) where If is a matrix having columns corresponding to eigenvectors and A is a matrix containing eigenvalues.[0082J At step 118, an eigenvector of A containing the characteristic modes of A is reconstructed using the following relation:(^X’V AV (6)

[0083] As shown above, the method 106a may be used to determine the eigenvector of A without computing^ directly. The method 106a may be suitable in applications where the size ofXandX’ is relatively large such that direct computation of A requires significant computation time.

[0084] Additional details and examples of DMD algorithms, which may be incorporated into the methods disclosed herein, may be found in Kutz.2.3 Early Stop Criterion

[0085] As described above, the future state, X”, may be evaluated against an early stop criterion (or early stop criteria if multiple observability metrics are used) to determine whether training is likely to succeed or not. This may be accomplished by evaluating each observability metric in the future state, X”, against a corresponding early stop criterion for that metric. In some implementations, if any one of the observability metrics in the future state satisfies its corresponding early stop criterion, training may be terminated or a corrective action performed before resuming training.

[0086] In situations where at least one observability metric satisfies its corresponding early stop criterion and at least one other observability metric does not satisfy its corresponding early stop criterion, various approaches may be used to decide whether training should continue or training should be terminated or a corrective action performed. In one example approach, the decision may be made based on a simple majority. If the number of observability metrics that do not satisfy their corresponding early stop criterion is greater than the number ofobservability metrics that do satisfy their corresponding early stop criterion, training may continue. If the number of observability metrics that do not satisfy their corresponding early stop criterion is less than the number of observability metrics that do satisfy their corresponding early stop criterion, training may be terminated or a corrective action performed. In another example approach, the observability metrics may be given different weights to emphasize the importance of some observability metrics over others when determining whether training should continue or not. It should be appreciated that these approaches are non-limiting examples and that users may generally define and customize how they use the future state, X”, and evaluate it against various early stop criterion according to their needs.

[0087] For some observability metrics, the values collected during training may exhibit variability, i.e., the values deviate from a statistical mean. As an illustrative example, FIG. 5 shows example data for a mean loss observability metric as a function of iteration step collected during training of a nanoGPT model with 300,000 parameters. As shown, the mean loss decreases with iteration step, but varies around its running mean. In some implementations, the variability may be an indicator of whether the metric is likely to converge, thus providing an indicator to determine if training is likely to succeed or not.

[0088] Accordingly, an early stop criterion may be defined based on this variability. For example, the data collected for the mean loss metric of FIG. 5 may be used to determine a running standard deviation, c, and a running mean, p. The predicted value may be considered to satisfy the early stop criterion if the predicted value of the metric deviates more than or equal to 3o from the running mean, p, i.e., the predicted value is greater than or equal to p+3o or the predicted value is less than or equal to p-3o. It should be appreciated this condition is not limited to the mean loss metric, but may be applied to other observability metrics exhibiting variability, such as the gradient norm, an alpha PL parameter, and / or the like.

[0089] In some implementations, the observability metrics may be monitored and evaluated using a process capability parameter, CPK. The process capability parameter, CPK, is a numerical parameter that provides a convenient way to evaluate whether an observability metric is likely to fall within an acceptable range of values given the variability of that metric as a function of iteration step. The process capability parameter, CPK, for a particular observability metric may be defined according to the following relation:„ . .CPK= minimum[0090} where UL is an upper limit and LL is a lower limit. The upper limit and the lower limit define the range of values for the observability metric to be considered acceptable. The process capability parameter, CPK, also considers variations in the running mean, p, of the metric with respect to the upper and lower limits as shown in Eq. (7). It should be appreciated that when multiple observability metrics are considered in states X and X’, a CPK value for each observability metric may be determined during training.[009.11 Generally, a higher value for the process capability parameter, CPK, indicates the metric is likely to fall within acceptable bounds. This occurs when the running mean, p, is located closer to the midpoint between the upper and lower limits and / or when the standard deviation is low (i.e., the spread of the values is small). Conversely, a lower value for the process capability parameter, CPK, indicates the metric is likely to fall outside acceptable bounds. This occurs when the running mean, p, is located closer to one of the upper and lower limits and / or when the standard deviation is high (i.e., the spread of the values is large).10092 J As an illustrative example, FIG. 6A shows a plot of an observability metric as a function of iteration step. For this metric, an upper limit (UL) and a lower limit (LL) may be defined before training starts. Accordingly, a centerline (CL) may also be shown to indicate the proximity of the values collected for the metric to the midpoint between the upper and lower limits. The plot in FIG. 6A may be updated continuously during training to track the changes in the observability metric with respect to the upper and lower limits. Based on the data in FIG. 6A, the CPK value may also be obtained as a function of iteration number using Eq. (7) where the running mean, p, and the running standard deviation, c, are computed using previously collected data. FIG. 6B shows a plot of the CPK value corresponding to the data shown in FIG. 6A.[0093| In some implementations, the CPK value for an observability metric may be computed using values from the first state, X, the second state X’, and / or the future state X”. By including values from the future state, X”, the CPK value may be treated as a prediction despite also including data previously collected during training. This CPK value, in turn, may be compared against an early stop criterion. For example, the early stop criterion may be defined to be a threshold where if the CPK value remains above the threshold, training may continue and if the CPK value falls below the threshold, training may be terminated, or a correction action performed. In some implementations, the threshold may be equal to 1.33, i.e., the smallest difference between the running mean, p, and the lower or upper limits is equal to 4o. In someimplementations, the upper limit, the lower limit, and the threshold may each be predetermined values.

[0094] In some implementations, an observability metric formulated as a process capability parameter may be used during multiple stages of training with adjustments to the upper and lower limits to reflect different target ranges for a particular observability metric. For example, FIG. 6C shows a representative observability metric as a process capability parameter with a relatively wider range of acceptable values during pre-training and a narrower range of acceptable values during fine-tuning in order to achieve a higher CPK value in a more gradual manner.3. An Example Runtime Monitor using DMD

[0095] The inventive implementations of DMD disclosed herein may be implemented in a runtime monitor, which is a software diagnostic tool that monitors and evaluates the quality of a training run as a model is being trained. In some implementations, the runtime monitor may operate in real time or near real time (e.g., less than 1 second).

[0096] FIG. 7A shows a non-limiting example of a runtime monitor 210. As shown, the monitor 210 may be operatively coupled to a machine learning model 201 (e.g., an LLM) to evaluate the quality of training as the model 201 is being trained. The monitor 210 may include a metrics collector 220 to receive data associated with one or more observability metrics from the model 201 during training. In some implementations, the data may be divided between two subsets: a) a subset of data 280 that requires relatively little or, in some instances, no processing to determine associated observability metrics (e.g., training loss, validation loss); and b) a subset of data 282 that requires more extensive processing to determine associated observability metrics (e.g., an alpha PL parameter). The subset of data 280 may be transmitted directly from the model 201 to the metrics collector 220. The subset of data 282 may be transmitted from the model 201 to a data source 202. The data source 202 provides computational resources (e.g., an allocation of GPUs) to compute the observability metrics from the subset of data 282. These observability metrics are, in turn, transmitted from the data source 202 to the metrics collector 220 (see data 284).

[0097] The observability metrics received by the metrics collector 220 may be stored in memory (e.g., memory associated with the metrics collector 220) for subsequent retrieval and / or use. For example, one or more observability metrics may be stored corresponding to one or more iteration steps. In some implementations, one or more observability metrics maybe stored for every iteration step up to the most recent iteration step. In some implementations, one or more observability metrics may be stored incrementally, e.g., as checkpoints.

[0098] The monitor 210 further includes a DMD engine 230 operatively coupled to the metrics collector 220. The metrics collector 220 may transmit a subset or, in some instances, all the observability metrics to the DMD engine 230 for evaluation (i.e., see data 286). The DMD engine 230 may predict future values of one or more observability metrics using DMD and evaluate the predicted values of the observability metric(s) against an early stop criterion (or an early stop criteria if multiple observability metrics are used). This may be accomplished, for example, by the DMD engine 230 executing the method 100 described in Section 2. Once the DMD engine 230 is finished evaluating the observability metrics, the DMD engine 230 may generate and transmit a signal 290 to the metrics collector 220 indicating whether the training run should continue, be terminated (e.g., due to a high likelihood of failure), or a corrective action should be applied. The metrics collector 220 may, in turn, generate and transmit a signal 292 to the model 201 containing instructions based on the signal 290. For instance, the signal 292 may contain instructions for a GPU to suspend its activities until further instructions are provided when the signal 290 indicates the training run should be terminated.

[0099] The various components of the monitor 210 (e.g., the metrics collector 220, the DMD engine 230) may be executable by any suitable components of a computing system. The computing system may generally include one or more processors (e.g., a CPU, a GPU) to execute the functions of the monitor 210. The computing system may further include memory to store instructions to cause the processor(s) to execute processes and / or functions associated with the monitor 210 and / or data used by the monitor 210 to evaluate training (e.g., the observability metrics). In some implementations, each processor may have dedicated memory. For example, if the computing system includes one or more CPUs and one or more GPUs, each CPU and each GPU may have memory dedicated to that particular device. Various types of processors may be used including, but not limited to, a general-purpose processor, a Field Programmable Gate Array (FPGA), an Application Specific Integrated Circuit (ASIC), a Digital Signal Processor (DSP), and / or the like. The memory may include, but is not limited to, a random-access memory (RAM), a memory buffer, a hard drive, a database, an erasable programmable read-only memory (EPROM), an electrically erasable read-only memory (EEPROM), a read-only memory (ROM), Flash memory, and / or so forth.

[0100] In some implementations, the computing system supporting the monitor 210 may also support training of the model 201. For example, the computational resources of the computingsystem may be divided such that a first portion of the resources are allocated to train the model 201 and a second portion of the resources are allocated to facilitate operation of the monitor 210. In some implementations, the computing system supporting the monitor 210 may be operatively coupled to another computing system that trains the model 201.[0101 | In one non-limiting example, the inventive implementations of DMD disclosed herein, including the monitor 210 described above, may be executed using a GPU cluster, which is often deployed in large-scale computing systems like a datacenter or a cloud computing system. FIG. 7B shows an example GPU cluster 300. As shown, the GPU cluster 300 may include multiple GPU servers 320a-320n (also referred to as a GPU server 320). Each GPU server 320 may include at least one CPU 322 (often referred to as a “host”). The CPU 322 may be communicatively coupled to multiple GPUs, e.g., GPUs 324a-324n (also referred to as a GPU 324). Each GPU 324 is further communicatively coupled to a corresponding network interface controller (NIC) (also sometimes referred to as a network interface card). For example, GPU 324a is communicatively coupled to NIC 326a, GPU 324b is communicatively coupled to NIC 326b, and GPU 324n is communicatively coupled to NIC 326n.

[0102] The CPUs 322 of respective GPU servers 320 may be communicatively coupled to each other via a frontend network 312. The GPU cluster 300 may also include a backend network 314 to facilitate direct communication between the GPUs 324a-324n (also referred to as a GPU 324) of the GPU servers 320a-320n. For example, the GPUs 324a-324n of the GPU servers 320a-320n may be communicatively coupled together via the NICs 326a-326n and the backend network 314.

[0103] As discussed above, the inventive implementations of DMD disclosed herein may require appreciably little in terms of computational resources (e.g., processing power and memory storage) for execution. Accordingly, in some implementations, the functions of the monitor 210 may be executed using a CPU rather than a GPU. Referring to the example in FIG. 7B, the functions of the monitor 210 may be executed using one of the CPUs 322. In this manner, the pool of GPU-related resources (e.g., the GPUs 324) available to train the model is increased, i.e., since none of the GPUs 324 need to support the monitor 210. Additionally, since the CPU 322 is communicatively coupled to every GPU 324 for a given GPU server 320, the monitor 210 may have access to any outgoing data transmitted from the CPU 322 to the GPUs 324 for that GPU server 320 or any incoming data transmitted from the GPUs 324 to the CPU 322 for that GPU server 320 is readily accessible for use as an observability metric when executing the functions of the monitor 210 (e.g., executing the DMD engine 230). It should beappreciated that various functions related to the model 201 (e.g., training the model 201) and the data source 202 (e.g., computing observability metrics, such as the alpha PL parameter) may be supported by the GPUs 324.4. Conclusion

[0104] All parameters, dimensions, materials, and configurations described herein are meant to be example and the actual parameters, dimensions, materials, and / or configurations will depend upon the specific application or applications for which the inventive teachings is / are used. It is to be understood that the foregoing embodiments are presented primarily by way of example and that, within the scope of the appended claims and equivalents thereto, inventive embodiments may be practiced otherwise than as specifically described and claimed. Inventive embodiments of the present disclosure are directed to each individual feature, system, article, material, kit, and / or method described herein.

[0105] In addition, any combination of two or more such features, systems, articles, materials, kits, and / or methods, if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent, is included within the inventive scope of the present disclosure. Other substitutions, modifications, changes, and omissions may be made in the design, operating conditions and arrangement of respective elements of the example implementations without departing from the scope of the present disclosure. The use of a numerical range does not preclude equivalents that fall outside the range that fulfill the same function, in the same way, to produce the same result.

[0106] The above-described embodiments can be implemented in multiple ways. For example, embodiments may be implemented using hardware, software or a combination thereof. When implemented in software, the software code can be executed on a suitable processor or collection of processors, whether provided in a single computer or distributed among multiple computers.

[0107] Further, it should be appreciated that a computer may be embodied in any of a number of forms, such as a rack-mounted computer, a desktop computer, a laptop computer, or a tablet computer. Additionally, a computer may be embedded in a device not generally regarded as a computer but with suitable processing capabilities, including a Personal Digital Assistant (PDA), a smart phone or any other suitable portable or fixed electronic device.

[0108] Also, a computer may have one or more input and output devices. These devices can be used, among other things, to present a user interface. Examples of output devices that canbe used to provide a user interface include printers or display screens for visual presentation of output and speakers or other sound generating devices for audible presentation of output. Examples of input devices that can be used for a user interface include keyboards, and pointing devices, such as mice, touch pads, and digitizing tablets. As another example, a computer may receive input information through speech recognition or in other audible format.[0109J Such computers may be interconnected by one or more networks in a suitable form, including a local area network or a wide area network, such as an enterprise network, an intelligent network (IN) or the Internet. Such networks may be based on a suitable technology, may operate according to a suitable protocol, and may include wireless networks, wired networks or fiber optic networks. The various methods or processes outlined herein may be coded as software that is executable on one or more processors that employ any one of a variety of operating systems or platforms. Additionally, such software may be written using any of a number of suitable programming languages and / or programming or scripting tools, and also may be compiled as executable machine language code or intermediate code that is executed on a framework or virtual machine. Some implementations may specifically employ one or more of a particular operating system or platform and a particular programming language and / or scripting tool to facilitate execution.

[0111] Also, various inventive concepts may be embodied as one or more methods, of which an example has been provided. The acts performed as part of the method may be ordered in any suitable way. Accordingly, embodiments may be constructed in which acts are performed in an order different than illustrated, which may include performing some acts simultaneously, even though shown as sequential acts in illustrative embodiments.

[0112] All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entirety.

[0113] All definitions, as defined and used herein, should be understood to control over dictionary definitions, definitions in documents incorporated by reference, and / or ordinary meanings of the defined terms.

[0114] The indefinite articles “a” and “an,” as used herein in the specification and in the claims, unless clearly indicated to the contrary, should be understood to mean “at least one.”

[0115] The phrase “and / or,” as used herein in the specification and in the claims, should be understood to mean “either or both” of the elements so conjoined, i.e., elements that areconjunctively present in some cases and disjunctively present in other cases. Multiple elements listed with “and / or” should be construed in the same fashion, i.e., “one or more” of the elements so conjoined. Other elements may optionally be present other than the elements specifically identified by the “and / or” clause, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, a reference to “A and / or B”, when used in conjunction with open-ended language such as “comprising” can refer, in one embodiment, to A only (optionally including elements other than B); in another embodiment, to B only (optionally including elements other than A); in yet another embodiment, to both A and B (optionally including other elements); etc.

[0116] As used herein in the specification and in the claims, “or” should be understood to have the same meaning as “and / or” as defined above. For example, when separating items in a list, “or” or “and / or” shall be interpreted as being inclusive, i.e., the inclusion of at least one, but also including more than one, of a number or list of elements, and, optionally, additional unlisted items. Only terms clearly indicated to the contrary, such as “only one of’ or “exactly one of,” or, when used in the claims, “consisting of,” will refer to the inclusion of exactly one element of a number or list of elements. In general, the term “or” as used herein shall only be interpreted as indicating exclusive alternatives (i.e., “one or the other but not both”) when preceded by terms of exclusivity, such as “either,” “one of,” “only one of,” or “exactly one of.” “Consisting essentially of,” when used in the claims, shall have its ordinary meaning as used in the field of patent law.

[0117] As used herein in the specification and in the claims, the phrase “at least one,” in reference to a list of one or more elements, should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed within the list of elements and not excluding any combinations of elements in the list of elements. This definition also allows that elements may optionally be present other than the elements specifically identified within the list of elements to which the phrase “at least one” refers, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, “at least one of A and B” (or, equivalently, “at least one of A or B,” or, equivalently “at least one of A and / or B”) can refer, in one embodiment, to at least one, optionally including more than one, A, with no B present (and optionally including elements other than B); in another embodiment, to at least one, optionally including more than one, B, with no A present (and optionally including elements other than A); in yet another embodiment, to at least one, optionally including morethan one, A, and at least one, optionally including more than one, B (and optionally including other elements); etc.[Oi l 8] In the claims, as well as in the specification above, all transitional phrases such as “comprising,” “including,” “carrying,” “having,” “containing,” “involving,” “holding,” “composed of,” and the like are to be understood to be open-ended, i.e., to mean including but not limited to. Only the transitional phrases “consisting of’ and “consisting essentially of’ shall be closed or semi-closed transitional phrases, respectively, as set forth in the United States Patent Office Manual of Patent Examining Procedures, Section 2111.03.

Claims

CLAIMS1. A computing system for evaluating training of a machine learning model, the system comprising: a plurality of graphics processing units (GPUs) to train the machine learning model; a central processing unit (CPU), communicatively coupled to each GPU of the plurality of GPUs, to receive data associated with one or more metrics related to the training of the machine learning model from the plurality of GPUs; and a memory, communicatively coupled to the CPU, storing instructions that, when executed by the CPU during training of the machine learning model, cause the CPU to:A) acquire a first state, X, containing values for the one or more metrics spanning iteration steps i through i+m-1 of the training wherein z and m are positive integer values;B) acquire a second state, X’, containing values for the one or more metrics spanning iteration steps z+z through i+m-l+z of the training wherein z is a positive integer value;C) determine a best fit linear operator, A, that relates X to X’ according to a first relation:X’ ~AXD) determine a future state, X”, usings, the future state containing values for the one or more metrics at an iteration step i+n wherein n is a positive integer value greater than m-l+z andE) evaluate the future state against an early stop criterion.

2. The system of claim 1, wherein: the plurality of GPUs is a first plurality of GPUs and the data is first data; and the system further comprises: a second plurality of GPUs, communicatively coupled to the first plurality of GPUs and the CPU, to receive second data associated with the one or more metrics, generate values for at least one metric of the one or more metrics, and transmit the values for the at least one metric to the CPU.

3. The system of claim 2, wherein the at least one metric comprises an alpha PL parameter.

4. The system of claim 1, wherein the one or more metrics comprises at least one of a cross entropy, a gradient norm, a perplexity, a learning rate, a validation loss, a BiLingual Evaluation Understudy (BLEU) score, a Recall-Oriented Understudy for Gisting Evaluation (ROUGE) score, a Fl score, precision, recall, an Area Under the Curve Receiver Operating Characteristics (AUC-ROC) score, an Area Under the Curve Precision Recall Curve (AUC- PRC) score, or an alpha PL parameter.

5. The system of claim 1, wherein a number of metrics included in each of the first state and the second state ranges from 1 to 1000.

6. The system of claim 1, wherein the first state contains values for the one or more metrics at every iteration step spanning steps i through i+m-1.

7. The system of claim 1, wherein the first state contains values for the one or more metrics only for a subset of iteration steps spanning steps i through i+m-1.

8. The system of claim 7, wherein the subset of iteration steps includes iteration steps at equal increments between the steps i and i+m-1 of the training.

9. The system of claim 7, wherein the subset of iteration steps includes iteration steps at nonequal increments between the steps i and i+m-1 of the training.

10. The system of claim 1, wherein a number of iteration steps spanning steps i through i+m-1 ranges from 100 to 2000.

11. The system of claim 1, wherein z ranges from 1 to 100.

12. The system of claim 1, wherein the second state, X’, contains at least one value for the one or more metrics interpolated from the data received by the CPU.

13. The system of claim 1, wherein: the data associated with the one or more metrics includes data at iteration steps ii and 72 greater than zy; and the instructions cause the CPU to, when acquiring the second state:interpolate the data at iteration steps ii and U to determine an estimate of the data at an iteration step is between ii and 12 for inclusion in the second state.

14. The system of claim 1, wherein the instructions further cause the CPU to, when determining the best fit linear operator, A: determine d according to a second relation:A = X’Xi- wherein is a Moore-Penrose generalized inverse of X.

15. The system of claim 1, wherein the instructions further cause the CPU to, when determining the best fit linear operator, A: determine a singular value decomposition of X according to a second relation:X^ UXW wherein Uis a matrix having columns corresponding to left-singular vectors, X is a matrix containing singular values, and W is a matrix having columns corresponding to rightsingular vectors; determine a projection, d, of A according to a third relation:A = IXX’VX1wherein IX is a conjugate transpose of U, Cis a conjugate transpose of W, and X"1is an inverse of 27; execute an eigendecomposition of A according to a fourth relation:AW = WA wherein is a matrix having columns corresponding to eigenvectors and A is a matrix containing eigenvalues; and reconstruct an eigenvector, , ofd according to a fifth relation: =X’VX]W16. The system of claim 1, wherein a difference between n and m-l+z ranges from 1 to 2000.

17. The system of claim 1, wherein a difference between n and m-l+z is equal to 500.

18. The system of claim 1, wherein a difference between n and m-l+z is equal to 100019. The system of claim 1, wherein a difference between n and m-l+z is equal to 2000.

20. The system of claim 1, wherein the instructions cause the CPU to, when evaluating the future state against the early stop criterion: determine, for a metric of the one or more metrics, a running mean,and a standard deviation, <r, associated with the metric based on the values of the metric in at least one of the first state or the second state; and determine whether a value of the metric in the future state is one of greater than or equal to / / +3<T or less than or equal to u- o.

21. The system of claim 20, wherein the metric is a gradient norm.

22. The system of claim 1, wherein the instructions cause the CPU to, when evaluating the future state against the early stop criterion: determine, for a metric of the one or more metrics, a running mean,and a standard deviation, <7, associated with the metric based on the values of the metric in the future state, the first state, and the second state; determine a CPK value of the metric according to a second relation:wherein UL is a predetermined upper limit and LL is a predetermined lower limit; and determine whether the CPK value is below a threshold.

23. The system of claim 22, wherein the threshold is equal to 1.33.

24. The system of claim 22, wherein the instructions further cause the CPU to: in response to the CPK value being greater than or equal to the threshold, allow the training of the machine learning model to continue; and in response to the CPK value being less than the threshold, terminate the training of the machine learning model or apply a corrective action to the training of the machine learning model.

25. The system of claim 24, wherein: the threshold is a first threshold; andthe instructions further cause the CPU to, in response to the CPK value being less than the first threshold: determine whether the CPK value is below a second threshold less than the first threshold; in response to the CPK value being greater than or equal to the second threshold, apply a corrective action to the training of the machine learning model; and in response to the CPK value being less than the second threshold, terminate the training of the machine learning model.

26. The system of claim 1, wherein the instructions further cause the CPU to: in response to determining the future state does not satisfy the early stop criterion, allow the training of the machine learning model to continue.

27. The system of claim 1, wherein the instructions further cause the CPU to: in response to determining the future state satisfies the early stop criterion, terminate the training of the machine learning model.

28. The system of claim 27, wherein the instructions cause the CPU to, when terminating the training of the machine learning model: generate a signal with instructions for the plurality of GPUs to stop training the machine learning model; and transmit the signal to the plurality of GPUs for execution.

29. The system of claim 1, wherein the instructions further cause the CPU to: in response to determining the future state satisfies the early stop criterion, apply a corrective action to the training of the machine learning model.

30. The system of claim 29, wherein the instructions cause the CPU to, when applying the corrective action to the training of the machine learning model: generate a signal with instructions for the plurality of GPUs to apply the corrective action; and transmit the signal to the plurality of GPUs for execution.

31. The system of claim 29, wherein the corrective action includes an adjustment to one or more hyperparameters associated with the training of the machine learning model.

32. The system of claim 31, wherein the one or more hyperparameters comprises at least one of a learning rate, a batch size, or a temperature.

33. The system of claim 1, wherein the system is a GPU cluster.

34. The system of claim 1, wherein the system is configured to execute A) through E) in less than 1 second.

35. A method for evaluating training of a machine learning model using a computing system, the computing system comprising: a plurality of graphics processing units (GPUs); and a central processing unit (CPU), communicatively coupled to each GPU of the plurality of GPUs, the method comprising:A) acquiring a first state, X, containing values for one or more metrics related to the training of the machine learning model from the plurality of GPUs spanning iteration steps i through i+m-1 of the training wherein z and m are positive integer values;B) acquiring a second state, X’, containing values for the one or more metrics spanning iteration steps z+z through i+m-l+z of the training wherein z is a positive integer value;C) determining a best fit linear operator, A, that relates X to X’ according to a first relation:X’ ~AXD) determining a future state, X”, usings, the future state containing values for the one or more metrics at an iteration step i+n wherein n is a positive integer value greater than m-l+z andE) evaluating the future state against an early stop criterion.

36. The method of claim 35, further comprising, before step A): training, by the plurality of GPUs, the machine learning model; andreceiving, by the CPU from the plurality of GPUs, data associated with the one or more metrics.

37. A non-transitory computer-readable storage medium storing instructions for evaluating a machine learning model, the instructions, when executed, cause at least one processor of a central processing unit (CPU) communicatively coupled to a plurality of graphics processing units (GPUs) to, while the plurality of GPUs is training the machine learning model:A) acquire a first state, X, containing values for one or more metrics related to the training of the machine learning model from the plurality of GPUs spanning iteration steps i through i+m-1 of the training wherein z and m are positive integer values;B) acquire a second state, X containing values for the one or more metrics spanning iteration steps z+z through i+m-l+z of the training wherein z is a positive integer value;C) determine a best fit linear operator, A, that relates X to X’ according to a first relation:X’ ~AXD) determine a future state, X”, usings, the future state containing values for the one or more metrics at an iteration step i+n wherein n is a positive integer value greater than m- l+z andE) evaluate the future state against an early stop criterion.

38. A method for evaluating an early stop criterion during training of a large language model (LLM), the method comprising: acquiring a first state, X, associated with the training of the LLM, the first state being a matrix containing values for a set of metrics associated with the training of the LLM spanning iteration steps z through z+zzz-1 of the training wherein z and m are integer values; acquiring a second state, X’, associated with the training of the LLM, the second state being a matrix containing values for the set of metrics spanning iteration steps i+1 through i+m of the training; determining a best fit linear operator, A, that relates X to X’ according to a first relation:X’ ~AXdetermining a future state, X’ associated with the training of the LLM using A, the future state being a matrix containing values for the set of metrics at an iteration step i+n wherein n is an integer value greater than and evaluating the future state against the early stop criterion.

39. The method of claim 38, wherein determining the best fit linear operator, A, comprises: determining^ according to a second relation:A = X’Xl- wherein Jf' is a Moore-Penrose generalized inverse of X.

40. The method of claim 38, wherein determining the best fit linear operator, A, comprises: determining a singular value decomposition of X according to a second relation: X^ UX wherein CZis a matrix having columns corresponding to left-singular vectors, X is a matrix containing singular values, and V!is a matrix having columns corresponding to rightsingular vectors; determining a projection, A, of A according to a third relation:A = IXX’VX1wherein IX is a conjugate transpose of U, Eis a conjugate transpose of W, and X"1is an inverse of 27; executing an eigendecomposition of A according to a fourth relation:AW = WA wherein Wis a matrix having columns corresponding to eigenvectors and / I is a matrix containing eigenvalues; and reconstructing an eigenvector, , of A according to a fifth relation: d^XAXXW41. The method of claim 38, wherein the set of metrics comprises at least one of a cross entropy, a gradient norm, a perplexity, a learning rate, a validation loss, a BiLingual Evaluation Understudy (BLEU) score, a Recall-Oriented Understudy for Gisting Evaluation (ROUGE) score, a Fl score, precision, recall, an Area Under the Curve Receiver OperatingCharacteristics (AUC-ROC) score, or an Area Under the Curve Precision Recall Curve (AUC-PRC) score.

42. The method of claim 38, wherein evaluating the future state against the early stop criterion comprises: determining, for a metric of the set of metrics, a running mean,and a standard deviation, cr, associated with the metric based on the values of the metric in at least one of the first state or the second state; and determining whether a value of the metric in the future state is one of greater than or equal to / / +3cr or less than or equal to u- o.

43. The method of claim 42, wherein the metric is a gradient norm.

44. The method of claim 38, wherein evaluating the future state against the early stop criterion comprises: determining, for a metric of the set of metrics, a running mean,and a standard deviation, cr, associated with the metric based on the values of the metric in the future state, the first state, and the second state; determining a CPK value of the metric according to a second relation:wherein UL is a predetermined upper limit and LL is a predetermined lower limit; and determining whether the CPK value is below a predetermined threshold.

45. The method of claim 44, wherein the predetermined threshold is equal to 1.33.

46. The method of claim 38, further comprising: in response to determining the future state does not satisfy the early stop criterion, continuing training of the LLM.

47. The method of claim 38, further comprising: in response to determining the future state satisfies the early stop criterion, terminating training of the LLM.

48. A non-transitory computer-readable storage medium storing instructions for evaluating an early stop criterion during training of a large language model (LLM), the instructions, when executed, cause at least one processor to: acquire a first state, A, associated with the training of the LLM, the first state being a matrix containing values for a set of metrics associated with the training of the LLM spanning iteration steps i through i+m-1 of the training wherein z and m are integer values; acquire a second state, X’, associated with the training of the LLM, the second state being a matrix containing values for the set of metrics spanning iteration steps z+7 through i+m of the training; determine a best fit linear operator, A, that relates X to X’ according to a first relation:X’ ~AX determine a future state, X”, associated with the training of the LLM using A, the future state being a matrix containing values for the set of metrics at an iteration step i+n wherein n is an integer value greater than m; and evaluate the future state against the early stop criterion.

Citation Information

Patent Citations

  • Photovoltaic power generation model prediction control system based on Koopman operator

    CN114944669A

  • Method and System for Predicting Dynamical Flows from Control Inputs and Limited Observations

    US20200311439A1

  • System and method to integrate a dynamic model for agents in a simulation environment using a deep koopman model

    US20210081808A1

Cited By

  • Sound model training method and sound model training device

    CN120544545A