Performance-guided parameter initialization for neural network training

WO2026207472A1PCT designated stage Publication Date: 2026-10-01ENTANGLEMENT INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/021332
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2026-03-27
Filing Date
2026-03-27
Publication Date
2026-10-01

Smart Images

  • Figure US2026021332_01102026_PF_FP_ABST
    Figure US2026021332_01102026_PF_FP_ABST
Patent Text Reader

Abstract

A method of training a neural network includes generating an initial set of weights which provide a starting point for subsequent gradient-based optimization, the generating comprising applying multiple sets of weights to the neural network, evaluating performance of the multiple sets of weights using a loss function associated with the neural network, iteratively adjusting at least one set of weights based on the evaluated performance without computing gradients of the loss function with respect to the weights, and selecting a set of weights as the initial set of weights based on the iterative adjustments. The gradient-based optimization is then performed using the selected set of weights as the starting point, benefiting from a more advantageous initialization than conventional weight initialization techniques.
Need to check novelty before this filing date? Find Prior Art

Description

PERFORMANCE-GUIDED PARAMETER INITIALIZATION FOR NEURAL NETWORK TRAININGCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Application No. 63 / 778,459, filed March 27, 2025, entitled "WEIGHT INITIALIZATION HEURISTIC FOR NEURAL NETWORKS," the contents of which are incorporated herein by reference in their entirety.TECHNICAL FIELD

[0002] The present disclosure relates to neural network training, and more particularly to performance-guided parameter initialization techniques that provide advantageous starting points for subsequent optimization.BACKGROUND

[0003] Neural networks are computational models designed to recognize patterns, leam from data, make predictions, and solve complex problems. They are composed of layers of interconnected artificial neurons, each processing and transmitting information through weighted connections. These weights determine the strength of signals passing between neurons. Upon receiving an input, a neuron multiplies it by its corresponding weight, sums the weighted inputs, and processes the result through an activation function. As training progresses, the network refines these weights, improving its ability to generalize and make accurate predictions.

[0004] The initialization of these weights significantly impacts how effectively a neural network learns. Most neural networks begin with randomly assigned weight values, a simple and widely used approach that eliminates prior assumptions about tire optimal weight distribution. However, random initialization introduces inefficiencies, particularly in large and deep networks. Arbitrary starting points can slow convergence, increase computational costs, and cause models to become trapped in suboptimal local minima. Additionally, the inherent randomness complicates reproducibility, making it difficult to achieve consistent results across different training runs.

[0005] To address the limitations of random initialization, various initialization techniques have been developed. Xavier initialization scales initial weights based on the number of input and output neurons to maintain stable variance of activations and gradients across layers. He initialization draws weights from aGaussian distribution optimized for networks employing Rectified Linear Unit (ReLU) activation functions. While these techniques improve upon random initialization, they remain limited in flexibility. Their weight values are determined using predefined formulas before training begins, making them less adaptable to specific datasets or problem structures. Furthermore, these methods rely on assumptions about network architecture and activation functions that may not generalize across all neural network configurations.

[0006] Orthogonal initialization aims to improve upon random initialization by ensuring weight matrices are orthogonal, helping to maintain stable gradient propagation across layers. However, generating orthogonal matrices requires computationally expensive operations such as QR decomposition or singular value decomposition (SVD), which can impose significant overhead particularly in large-scale networks. Accordingly, there is a need for improved weight initialization techniques that are adaptive to specific datasets and network architectures, computationally efficient, and capable of providing consistent and advantageous starting points for subsequent gradient-based optimization.SUMMARY

[0007] Embodiments herein relate to a performance-guided approach for initializing parameters in neural networks. The heuristic iteratively explores weight configurations to determine weight values that serve as an advantageous starting point for subsequent gradient-based optimization. By providing a more informed starting point than conventional initialization techniques, this approach may accelerate convergence, reduce computational overhead during the gradient-based optimization phase, and improve final model performance.

[0008] The heuristic evaluates multiple weight configurations using the loss function associated with training the network. By directly assessing weight performance and iteratively refining weight values based on loss function evaluations rather than gradient computations, the heuristic adapts its search strategy to the specific characteristics of the dataset and network architecture. This performance-based feedback enables producing more informed and consistent starting points across different training runs, in contrast to the variability of purely random initialization and the fixed formulas of analytical initialization methods.

[0009] Tire heuristic is not constrained by predetermined formulas or assumptions about optimal weight distributions or network architecture. It can flexibly adapt to various types of neural networks, activation functions, and optimizers, and may be applied to any subset of network layers, including a plurality of hidden layers. These features collectively address several challenges in neural network training, including slow convergence, inconsistent results across training runs, and scalability limitations, while maintaining flexibility across a broad range of applications and network architectures. In someembodiments, the approach may also be applied to initialization of parameters in other machine learning contexts beyond neural network weight initialization.

[0010] In one embodiment, a method includes generating, for a neural network, an initial set of weights which provide a starting point for subsequent optimization of the neural network, the generating comprising applying multiple sets of weights to the neural network, evaluating performance of the multiple sets of weights using a loss function associated with the neural network, iteratively adjusting at least one set of weights from the multiple sets of weights based on the evaluated performance, and selecting a set of weights as the initial set of weights based on the iterative adjustments.

[0011] This summary is provided to introduce a selection of concepts in a simplified fonn that are further described below in the detailed description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter. The detailed description offers further modifications and variations.DESCRIPTION OF THE DRAWINGS

[0012] Some embodiments of the present disclosure are now described, by way of example only, and with reference to the accompanying drawings.

[0013] FIG. 1 is a diagram illustrating a neural network training environment according to aspects of the disclosure.

[0014] FIG. 2 is a flowchart illustrating a method of training a neural network according to aspects of the present disclosure.

[0015] FIG. 3 is a flowchart illustrating a method of training a neural network using heuristic-based weight initialization according to aspects of the present disclosure.

[0016] FIG. 4 is a flowchart illustrating a method of heuristic-based weight initialization according to aspects of the present disclosure.

[0017] FIG. 5 illustrates an example computing environment suitable for implementing aspects of the systems and methods described herein according to aspects of the present disclosure.DETAILED DESCRIPTION

[0018] The figures and the following description illustrate specific example embodiments. It will thus be appreciated that those skilled in the art will be able to devise various arrangements that, although notexplicitly described or shown herein, embody the principles of the embodiments and are included within the scope of the embodiments. Furthermore, any examples described herein are intended to aid in understanding the principles of the embodiments, and are to be construed as being without limitation to such specifically recited examples and conditions. As a result, the inventive concept(s) is not limited to the specific embodiments or examples described below, but by the claims and their equivalents.

[0019] FIG. 1 is a diagram illustrating a neural network training environment 100 according to aspects of the disclosure. The neural network training environment 100 illustrates a two-phase training approach in which a heuristic-based weight initialization phase precedes a gradient-based optimization phase, providing an advantageous starting point for the optimizer and improving overall training efficiency and model performance. The neural network training environment 100 includes training data 102, a weight initializer 120. an optimizer 140. and a neural network 150.

[0020] The training data 102 comprises a dataset suitable for training the neural network 150 and may include labeled or unlabeled examples depending on the training objective. The training data 102 is provided to both the weight initializer 120 and the optimizer 140, enabling both phases of training to operate on the same dataset.

[0021] The weight initializer 120 is configured to generate a set of initial weights 130 for the neural network 150 through initializer-based training 122. The initializer-based training 122 comprises a heuristicbased approach that iteratively explores weight configurations by applying multiple sets of weights to the neural network 150, evaluating their performance using a loss function, and adjusting weight values based on the evaluated performance, as described in further detail with respect to FIGS. 3 and 4. The weight initializer 120 produces initial weights 130 that serve as an advantageous starting point for subsequent optimization by the optimizer 140, reflecting the characteristics of the specific training data 102 and network architecture rather than relying on predetermined mathematical formulas or random assignment.

[0022] Tire optimizer 140 is configured to perform optimizer-based training 142 as the second phase of training, using the initial weights 130 generated by the weight initializer 120 as the starting point. Hie optimizer 140 performs gradient-based optimization to further update and refine the weights of the neural network 150, minimizing the loss function and improving model perfonnance. By starting from the initial weights 130 rather than randomly initialized weights, the optimizer-based training 142 may benefit from faster convergence, reduced sensitivity to learning rate selection, and improved final model perfonnance. Tire optimizer 140 may comprise any optimizer configured to update weights based on calculated gradients, including a stochastic gradient descent (SGD) optimizer, a root mean square propagation (RMSprop) optimizer, an adaptive gradient (Adagrad) optimizer, or an adaptive moment estimation (Adam) optimizer.

[0023] The neural network 150 may comprise any suitable neural network architecture, including a feedforward neural network, a convolutional neural network, a recurrent neural network, or a large language model. The initial weights 130 generated by the weight initializer 120 may be applied to any subset of layers of the neural network 150, including a plurality of hidden layers, enabling the heuristic-based initialization to target the most influential weight parameters for a given architecture and task.

[0024] FIG. 2 is a flowchart illustrating a method 200 of training a neural network according to aspects of the present disclosure. The steps of the method 200 will be described with respect to the neural network training environment 100 of FIG. 1, though it will be appreciated that the steps may be performed in other enviromnents and systems, and may include alternative steps, fewer steps, or additional other steps not shown.

[0025] In step 202, the weight initializer 120 generates the initial set of weights 130 using a heuristic that iteratively explores weight configurations and moves weight values toward better-performing configurations. The heuristic evaluates weight performance directly using the loss function associated with training the neural network 150, without computing gradients of the loss function with respect to the weights. This gradient-free approach reduces the computational overhead of the initialization phase compared to gradient-based methods, and may utilize only two loss function evaluations per iteration rather than the foil computation graph traversal required by gradient-based optimizers. The iterative nature of the heuristic allows it to adapt to the characteristics of a particular neural network architecture and dataset, producing initial weights that reflect the specific optimization landscape rather than relying on predetermined formulas. The heuristic is described in further detail with respect to FIGS. 3 and 4.

[0026] In step 204. the weight initializer 120 initializes the neural network 150 with the generated initial weights 130. The initial weights 130 may be applied to all layers of the neural network 150 or to a selected subset of layers, such as a plurality of hidden layers. In some embodiments, initial weights for layers not covered by the heuristic initialization may be set using conventional initialization techniques such as random initialization, Xavier initialization, or He initialization. In some embodiments, weights from a pre-trained model may be incorporated alongside the heuristically generated initial weights, with the pretrained weights applied to selected layers and the heuristically generated weights applied to remaining layers.

[0027] In step 206, the optimizer 140 continues training the neural network 150 using gradient-based optimization, starting from the initial weights 130 established in step 204. The optimizer-based training benefits from the advantageous starting point provided by the heuristic initialization, which may reduce tirenumber of gradient-based training epochs required to reach a target performance level and improve the stability and consistency of the training process across different training runs.

[0028] FIG. 3 is a flowchart illustrating a method 300 of training a neural network using heuristicbased weight initialization according to aspects of the present disclosure. The method 300 provides a high-level overview of the heuristic-based weight initialization process described with respect to step 202 of FIG. 2. The steps of the method 300 will be described with respect to the weight initializer 120 of FIG. 1, though it will be appreciated that the steps may be performed in other environments and systems, and may include alternative steps, fewer steps, or additional other steps not shown.

[0029] In step 302, multiple sets of weights are applied to the neural network. In some embodiments, the multiple sets of weights are initialized by randomly sampling weight values from a predetermined range, such as the range [-1, 1]. Each set of weights represents a distinct configuration of the neural network's parameters, and applying each set to the neural network allows the heuristic to evaluate the loss function under different weight configurations. In some embodiments, the multiple sets of weights may include weights derived from a pre-trained model, providing an alternative starting point for the heuristic exploration.

[0030] In step 304, performance of the multiple sets of weights is evaluated using the loss function associated with the neural network. The loss function measures the discrepancy between the neural network's predictions and the target values in the training data, providing a scalar performance measure for each set of weights. In some embodiments, all sets of weights are evaluated using tire same batch of training data to ensure a fair and consistent comparison of weight performance.

[0031] In step 306, at least one set of weights is iteratively adjusted based on the evaluated performance. The adjustment refines weight values based on the relative performance of better-performing and worse-performing configurations, as described in further detail with respect to FIG. 4. The iterative nature of the adjustment allows the heuristic to progressively refine the weight configurations toward regions of the weight space associated with lower loss values.

[0032] In step 308, a set of weights is selected as the initial set of weights based on the iterative adjustments. The selected set of weights typically corresponds to the lowest evaluated loss among all weight configurations explored during the heuristic process, providing the most advantageous starting point for subsequent gradient-based optimization.

[0033] FIG. 4 is a flowchart illustrating a method 400 of heuristic-based weight initialization according to aspects of the present disclosure. The method 400 describes the detailed iterative process by which theweight initializer 120 explores and refines weight configurations to generate the initial set of weights 130. The steps of the method 400 will be described with respect to the weight initializer 120 of FIG. 1, though it will be appreciated that the steps may be performed in other environments and systems, and may include alternative steps, fewer steps, or additional other steps not shown.

[0034] In step 402, a first set of weights and a second set of weights are applied to the neural network to obtain a first loss value and a second loss value, respectively. In some embodiments, the first set of weights and the second set of weights are initialized by randomly sampling weight values from a predetermined range, such as the range [-1, 1], providing an unbiased starting point for the heuristic exploration that does not depend on assumptions about the optimal weight distribution for a given network architecture or dataset. In some embodiments, each set of weights is evaluated using the same batch of training data, ensuring a fair and consistent comparison of weight performance across all iterations of the heuristic.

[0035] In step 404, a better-performing set of weights and a worse -performing set of weights are determined by comparing the first loss value and the second loss value. Tire set of weights associated with the lower loss value is designated as tire better-performing set, and the set associated with the higher loss value is designated as the worse -performing set. This comparison establishes the initial directional signal for the heuristic, indicating which region of the weight space is associated with better performance on the training data.

[0036] In step 406, values of the worse -performing set of weights are adjusted toward values of the better-performing set of weights to create an adjusted set of weights. The adjustment moves weight values in the direction of the better-performing configuration, exploring the weight space between the two sets and potentially identifying weight configurations that outperform both. In some embodiments, the adjustment may be expressed as:w_adjusted = w worse + a(w_better - w worse)

[0037] where w adjusted is the adjusted weight value, w worse is the weight value from the worseperforming set, w better is the corresponding weight value from the better-performing set, and a is a step size parameter where 0 < a < 1. The step size a may be fixed or dynamically adjusted based on factors such as the current iteration or the relative performance difference between the two sets of weights. This formulation ensures that the adjustment always moves in the direction of the better-performing weights regardless of the sign of the weight values, providing a principled exploration of the weight space between the two configurations. For example, if a weight in the worse -performing set is 0.3 and the correspondingweight in the better-performing set is 0.7, with a = 0.25, the adjusted weight is 0.3 + 0.25 x (0.7 - 0.3) = 0.4, moving 25% of the distance toward the better-performing weight.

[0038] In other embodiments, the adjustment may incorporate the magnitude of the worse-performing weights:new weights = worst_weights + leaming_rate x step_size x (best_weights - |worst_weights|)

[0039] In some implementations, the use of the absolute value of the worse-performing weights introduces a non-dircctional perturbation component that, when combined with subsequent normalization, promotes exploration of the weight space while maintaining bounded parameter magnitudes. In this embodiment, the adjustment is followed by a normalization step that constrains weight values to a predetermined range, such as [-1, 1], ensuring that weight values remain bounded throughout the iterative process. In some embodiments, the adjustment may be applied globally across all trainable parameters of the neural network simultaneously. In other embodiments, the adjustment may be applied layerwise, with separate better-performing and worse-performing weight sets maintained for each layer or group of layers.

[0040] In step 408, the adjusted set of weights is applied to the neural network to obtain an additional loss value. This step evaluates the performance of the adjusted weights on the same batch of training data used in step 402, enabling a direct comparison between the adjusted weights and the original betterperforming and worse-performing sets.

[0041] In step 410, the additional loss value is compared to respective loss values of the betterperforming set of weights and the worse-performing set of weights. This three-way comparison determines the relative performance of the adjusted weights and informs the redesignation decision in step 414.

[0042] In step 412, a determination is made as to whether a crossover condition is satisfied. The crossover condition defines the transition point between the heuristic-based initialization phase and the gradient-based optimization phase. The crossover condition may be based on various criteria, including a pre-defined maximum number of iterations, a target loss value, or a threshold loss improvement between iterations. If the crossover condition is not satisfied, the method proceeds to step 414.

[0043] In step 414, the better-performing set of weights and the worse -performing set of weights are selectively redesignated based on the comparison of step 410. In some embodiments, the redesignation is performed as follows: if the additional loss value of the adjusted set of weights is lower than the loss value of the better-performing set of weights, the adjusted set of weights is designated as the better-performing set and the previous better-performing set is designated as the worse -performing set for the subsequentiteration; if the additional loss value is between the loss values of the better-performing and worseperforming sets, the adjusted set of weights is designated as the worse-performing set and the betterperforming set is maintained for the subsequent iteration; and if the additional loss value exceeds the loss values of both the better-performing and worse-performing sets of weights, the better-performing and worse-performing designations may be maintained unchanged for tire subsequent iteration. In other embodiments, the redesignation may be performed by ranking all three sets of weights based on their respective loss values and designating the set with the lowest loss value as the better-performing set and the set with the next lowest loss value as the worse-performing set for the subsequent iteration. The process then returns to step 406 to continue iterating and refining the weights.

[0044] If the crossover condition is satisfied, the method proceeds to step 416, where the initial set of weight values is selected to provide the starting point for subsequent gradient-based optimization of the neural network. The set of weights that produced the lowest loss value during the iterative process is typically selected as the initial set. Step 416 represents the transition point at which the heuristic-based initialization phase concludes and the gradient-based training phase begins, with the selected initial weights providing an advantageous starting configuration that reflects the optimization landscape of the specific training data and network architecture.

[0045] FIG. 5 illustrates an example computing environment 500 suitable for implementing aspects of the systems and methods described herein according to aspects of the present disclosure. The computing environment 500 includes a computing device 510 that may represent any suitable physical or virtual machine capable of executing program instructions to perform the operations described herein. The computing device 510 includes one or more processors 512, such as general-purpose central processing units (CPUs), that execute software instructions stored in a system memory 514 or a storage medium 516. The storage medium 516 may include non-volatile memory devices such as solid-state drives, magnetic disks, or other persistent computer-readable media, and may store program instructions 518 corresponding to the modules and operations described with respect to FIGS. 1 through 4.

[0046] The computing device 510 may further include one or more hardware accelerators 520, such as graphics processing units (GPUs), tensor processing units (TPUs), or field-programmable gate arrays (FPGAs), configured to evaluate the loss function and perform the iterative weight adjustments described herein. In some embodiments, the weight initializer 120 and optimizer 140 of FIG. 1 may execute on different hardware components, for example the loss function evaluations of the heuristic-based initialization phase may be perfonned on one or more hardware accelerators 520 while the input and outputlayers of tire neural network operate on the processors 512, enabling efficient parallel computation during both the initialization and optimization phases.

[0047] A bus or interconnect 550 couples the components of the computing device 510, including one or more input / output devices 532, a display device interface 534, and a network interface 536. The network interface 536 enables communication with external systems or remote computing resources 542 via a network 540, which may include local area networks (LANs), wide area networks (WANs), the Internet, or cloud-based infrastructures.

[0048] Tire computing environment 500 provides a flexible, hardware-agnostic platfonn capable of executing the heuristic -based weight initialization and gradient-based optimization techniques described herein, supporting a range of neural network architectures and training configurations.

[0049] The heuristic-based weight initialization method described herein provides technical improvements to neural network training by transforming an arbitrary or formulaically determined starting weight configuration into a pcrfonnancc -guided starting configuration that reflects the specific optimization landscape of the training data and network architecture. This process improves the operation of the machine learning system by producing an initialization that reduces sensitivity to random initialization and improves convergence behavior during subsequent training. This transformation addresses a key challenge in neural network training in which gradient-based optimizers are sensitive to their starting configuration and may converge slowly, produce inconsistent results across training runs, or become trapped in suboptimal local minima when initialized with random or analytically derived weights. By providing a more informed starting point, the described approach may reduce the number of gradient-based training iterations required to reach a target performance level and improve convergence stability across different training runs.

[0050] The heuristic-based weight initialization approach described herein differs from prior approaches in several respects. Unlike analytical initialization methods such as Xavier and He initialization, the described approach is dataset- and architecture-responsive, adapting its weight search to the specific characteristics of the training data and network rather than applying predetermined formulas. Unlike pure random restart methods, the described approach uses directional performance-guided movement, specifically the iterative designation of better-performing and worse-performing parameter sets and directional updating of the worse-performing set toward the better-performing set based on loss function feedback, rather than independently sampling new weight configurations at each iteration. Unlike gradientbased warm start approaches, the described approach performs the initialization phase using only loss function evaluations, without computing gradients of the loss function with respect to the weights, reducing computational overhead during initialization. Unlike population-based global search methods such asevolutionary algorithms, the described approach operates on a minimal two-set paired comparison framework comprising only two candidate configurations per iteration, in which a better-performing set and a worse -performing set are iteratively designated and the worse-performing set is updated toward the better-performing set, rather than maintaining and evolving a large population of candidate solutions.

[0051] The gradient-free nature of the heuristic -based initialization phase provides a key computational advantage. Unlike gradient-based methods that require traversal of the full computation graph to compute parameter gradients, the heuristic may utilize only two loss function evaluations per iteration. During the initialization phase, the heuristic maintains two sets of w eights, the better-performing set and the worse-performing set, without maintaining first- or second-moment gradient estimates. This may reduce optimizer-state memory overhead during the initialization phase relative to gradient-based optimization methods such as the Adam optimizer, which maintains additional memory for first and second moment estimates of the gradients. The reduced memory requirement of the heuristic-based initialization phase may enable more efficient use of GPU memory, particularly during the early stages of training when gradient-based optimizers are most memory -intensive.

[0052] The heuristic described herein may be referred to as tire Best Direction Heuristic (BDH). Hie BDH may begin with two sets of weights and iteratively explore the weight space by moving the worseperforming set tow ard the better-performing set. In one embodiment, pseudocode for the BDH is as follows:INPUT: Stopping criteria expressed in terms of number of epochs x and training data weight 1 = generate random weight in range [-1, 1]weight2 = generate random weight in range [-1, 1]lossl = training loss based on batch and weight!Ioss2 = training loss based on batch and weight2best weights = weights corresponding to min(lossl, loss2)worst w eights = w eights corresponding to max(lossl, loss2)for i = ] 1. 2, ..., x}:ncw_wcights = worst weights + leaming rate x step_size x (best_weights - |worst_weights|)normalize new weights to range [-1, 1]new loss = training loss based on batch and new weightsif new_loss < best loss:best weights = new weightsworst weights = previous best weightselse if new loss < w orst_loss:worst weights = new weightsreturn best weights

[0053] In some embodiments, the overall training process incorporating the BDH comprises two phases: a first phase in which the BDH is run for x epochs to produce an initial set of best weights, and asecond phase in which a gradient-based optimizer such as the Adam optimizer is initialized with the best weights and run for the remaining N-x epochs, where N is the total number of training epochs. The crossover point x defines the transition between the two phases and may be determined based on a pre-defined maximum number of iterations, a target loss value, or a threshold loss improvement between iterations.

[0054] The weight adjustment in the BDH may alternatively be expressed as:w_adjusted = w worse + a(w_better - w worse)

[0055] where w adjusted is the adjusted weight value, w worse is the weight from the worseperforming set, w better is the corresponding weight from the better-performing set, and a is a step size parameter where 0 < a < 1. This fonnulation ensures movement toward the better-performing weights regardless of the sign of the weight values. For example, if a weight in the w'orse-performing set is 0.3 and the corresponding weight in the better-performing set is 0.7, with a = 0.25, the adjusted weight is 0.3 + 0.25 x (0.7 - 0.3) = 0.4, moving 25% of the distance tow ard the better-performing w eight. The step size a may be fixed or dynamically adjusted based on factors such as the current iteration or the relative performance difference between the two sets of w eights.

[0056] The paired-comparison directional update embodiment described herein represents one instance of a broader class of gradient-free performance-guided initialization heuristics in which multiple weight configurations are evaluated using the loss function and at least one configuration is iteratively adjusted based on relative performance comparisons betw een evaluated configurations. Other embodiments within this class may vary in the number of weight configurations maintained, the specific update rule applied to move a lower-performing configuration toward a higher-performing configuration, or the criteria used to determine when to transition to subsequent gradient-based optimization, while retaining the core approach of iterative adjustment guided by loss function evaluation.

[0057] In some embodiments, the hcuristic-bascd initialization may be applied to a plurality of hidden layers of the neural network, with the generated initial weights applied to those layers prior to the gradientbased optimization phase. In some embodiments, weights from a pre-trained model may be incorporated alongside the heuristically generated initial weights, with pre-trained weights applied to selected layers and heuristically? generated weights applied to remaining layers. This hybrid initialization approach may be particularly useful for transfer learning applications where some layers benefit from pre-trained weights while others require initialization adapted to the specific training data.

[0058] In some embodiments, the systems and methods described herein may be implemented using the computing environment described with respect to FIG. 5. A system for implementing the performance-guided parameter initialization methods described herein may comprise one or more processors and memory storing instructions that, when executed by the one or more processors, cause the system to perform operations comprising any of the methods described herein. A non-transitory computer-readable medium may store instructions that, when executed by one or more processors, cause the processors to perfomr operations comprising any of the methods described herein.

[0059] Although described herein in the context of neural network weight initialization, the performance-guided approach described herein may be applied more broadly to initialization of parameters for other machine learning models, hyperparameter optimization, and other optimization problems in which a set of numerical parameters can be evaluated using an objective function and a more advantageous starting point improves the efficiency of subsequent optimization. In all such applications, the core approach of maintaining a better-performing and worse-performing parameter set and iteratively updating the worseperforming set toward the better-performing set based on objective function evaluations remains applicable.

[0060] Although specific embodiments were described herein, the scope of the invention is not limited to those specific embodiments. The scope of the invention is defined by the following claims and any equivalents thereof.

Claims

CLAIMS1. A method of training a neural network, comprising:generating, for the neural network, an initial set of weights which provide a starting point for subsequent optimization of the neural network, the generating comprising:applying multiple sets of weights to the neural network;evaluating performance of the multiple sets of weights using a loss function associated with the neural network;iteratively adjusting at least one set of weights from the multiple sets of weights based on the evaluated performance; andselecting a set of weights as the initial set of weights based on the iterative adjustments.

2. The method of claim 1, wherein the generating comprises:(i) applying a first set of weights and a second set of weights to the neural network to obtain a first loss value and a second loss value, respectively;(ii) determining a better-performing set of weights and a worse-performing set of weights by comparing the first loss value and the second loss value;(iii) adjusting values of the worse-performing set of weights toward values of the betterperforming set of weights to create an adjusted set of weights;(iv) applying the adjusted set of weights to the neural network to obtain an additional loss value;(v) comparing the additional loss value to respective loss values of the better-performing set of weights and the worse-performing set of weights;(vi) selectively redesignating the better-performing set of weights and the worseperforming set of weights based on the comparison; and(vii) repeating steps (iii)-(vi) until a crossover condition is met to obtain the initial set of weights.

3. The method of claim 2, wherein the selectively redesignating comprises:if the additional loss value of the adjusted set of weights is lower than a loss value of the better-performing set of weights:designating the adjusted set of weights as a better-performing set of weights for a subsequent iteration; anddesignating the better-performing set of weights as a worse-performing set of weights for the subsequent iteration; andif the additional loss value of the adjusted set of weights is between the loss values of the better-performing set of weights and the worse-performing set of weights:designating the adjusted set of weights as a worse-performing set of weights for the subsequent iteration; andmaintaining the better-performing set of weights as a better-performing set of weights for the subsequent iteration.

4. The method of claim 2, wherein the selectively redesignating comprises:ranking the adjusted set of weights, the better-performing set of weights, and the worseperforming set of weights based on respective loss values;designating a set of weights having a lowest loss value as a better-performing set of weights for a subsequent iteration; anddesignating a set of weights having a next lowest loss value as a worse-performing set of weights for the subsequent iteration.

5. The method of claim 2, wherein the adjusting comprises:adjusting a value of each individual weight in the worse-performing set of weights in a direction toward a value of each corresponding weight in the better-performing set of weights;wherein a magnitude of each adjustment is based on a learning rate and a step size.

6. The method of claim 2, wherein the crossover condition defines a first number of epochs for performing steps (iii)-(vi), and further comprising training the neural network using the subsequent optimization for a second number of epochs, wherein a sum of the first number of epochs and the second number of epochs equals a total number of epochs for training the neural network.

7. The method of claim 2, further comprising:determining whether the crossover condition is met based on one or more of: a pre-defined maximum number of iterations, a target loss value, and a threshold loss improvement between iterations.

8. The method of claim 2, wherein the first set of weights and the second set of weights are each initialized by randomly sampling weight values from a predetermined range.

9. The method of claim 2, wherein the first set of weights, the second set of weights, and the adjusted set of weights are each evaluated using a same batch of training data.

10. The method of claim 1, wherein evaluating performance of the multiple sets of weights using the loss function is performed on one or more hardware accelerators and neural network layer computations are performed on one or more processors.

11. The method of claim 1, wherein applying the multiple sets of weights comprises incorporating weights derived from a pre-trained model as at least one of the multiple sets of weights.

12. The method of claim 1, wherein the subsequent optimization comprises gradient-based optimization, and further comprising performing the gradient-based optimization on the neural network using the selected set of weights as the starting point.

13. The method of claim 12, wherein the gradient-based optimization comprises using an adaptive moment estimation optimizer to update the weights of the neural network.

14. The method of claim 1, wherein the iteratively adjusting is performed without computing gradients of the loss function with respect to the weights.

15. The method of claim 1, wherein the initial set of weights is applied to a plurality of hidden layers of the neural network.

16. A system comprising:a weight initializer configured to apply multiple sets of weights to a neural network, evaluate performance of the multiple sets of weights using a loss function associated with the neural network, and iteratively adjust at least one set of weights based on the evaluated performance to generate an initial set of weights for the neural network; andan optimizer configured to perform subsequent optimization of the neural network using the initial set of weights as a starting point.

17. The system of claim 16, wherein the weight initializer is further configured to:apply a first set of weights and a second set of weights to the neural network to obtain a first loss value and a second loss value, respectively;determine a better-performing set of weights and a worse-performing set of weights by comparing the first loss value and the second loss value;adjust values of the worse-performing set of weights toward values of the betterperforming set of weights to create an adjusted set of weights;apply the adjusted set of weights to the neural network to obtain an additional loss value; compare the additional loss value to respective loss values of the better-performing set of weights and the worse-performing set of weights;selectively redesignate the better-performing set of weights and the worse-performing set of weights based on the comparison; andrepeat the adjusting, applying, comparing, and selectively redesignating until a crossover condition is met.

18. The system of claim 16, further comprising one or more hardware accelerators configured to evaluate the loss function and perform the iterative adjusting.

19. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:generating, for a neural network, an initial set of weights which provide a starting point for subsequent optimization of the neural network, the generating comprising:applying multiple sets of weights to the neural network;evaluating performance of the multiple sets of weights using a loss function associated with the neural network;iteratively adjusting at least one set of weights from the multiple sets of weights based on the evaluated performance; andselecting a set of weights as the initial set of weights based on the iterative adjustments.

20. The non-transitory computer-readable medium of claim 19, wherein the iteratively adjusting is performed without computing gradients of the loss function with respect to the weights.