System and method for progressive learning for machine-learned model to optimize training speed
By incrementally adjusting regularization and training complexity, the method optimizes the training speed and resource usage of machine learning models, addressing the challenges of high computational costs in training large models.
Patent Information
- Application Number
- JP2025031977
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-02-04
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-12-29
AI Technical Summary
The high computational costs and resource intensiveness of training large and complex machine learning models, particularly deep learning models, due to their size and the data required for training.
The method involves incrementally adjusting the regularization scale during training, starting with a weak regularization and gradually increasing it, along with incrementally increasing the complexity of training samples, to optimize training speed and reduce computational resources.
This approach significantly reduces the computational resources needed for training, increases training speed, and improves model accuracy by efficiently managing regularization and data complexity.
Smart Images

Figure 2025087772000001_ABST
Abstract
Description
Technical Field
[0001] Related Applications This application claims priority and the benefit of U.S. Provisional Patent Application No. 63 / 145,830, which is hereby incorporated by reference in its entirety.
[0002] This disclosure generally relates to incremental learning of machine learning models. More particularly, this disclosure relates to incremental adjustment of regularization during training of machine learning models to optimize training speed.
Background Art
[0003] Recent advances in machine learning have substantially increased the size and complexity of both machine learning models (e.g., neural networks, etc.) and the data used to train them. As an example, training of state-of-the-art deep learning models sometimes requires the use of thousands of graphics processing units for weeks at a time, and thus exhibits prohibitively high computational costs. Other networks can be trained quickly, but may involve costly overhead for a large number of parameters. Thus, any method of increasing training speed and parameter efficiency would substantially increase the availability of computing resources for other tasks.
Prior Art Documents
Non-Patent Documents
[0004]
Non-Patent Document 1
Non-Patent Document 2
[0005] Aspects and advantages of embodiments of the present disclosure are described in part in the following description, or can be learned from the description, or can be learned through the implementation of the embodiments.
[0006] One exemplary aspect of the present disclosure is directed to a computer-implemented method for efficient machine learning model training. The method may include obtaining, by a computing system comprising one or more computing devices, a plurality of training samples for a machine learning model. The method may include training, by the computing system for one or more first training iterations, a machine learning model using one or more respective first training samples of the plurality of training samples, at least partially based on a first regularization scale configured to control a relative effect of one or more regularization techniques. The method may include training, by the computing system for one or more second training iterations, a machine learning model using one or more respective second training samples of the plurality of training samples, at least partially based on a second regularization scale greater than the first regularization scale.
[0007] Another exemplary aspect of the present disclosure is directed to a computing system for determining a model with an optimized training speed. The computing system may include one or more processors. The computing system may include one or more tangible non-transitory computer-readable media storing computer-readable instructions that, when executed by the one or more processors, cause the one or more processors to perform operations. The operations may include generating a first machine learning model from a defined model search space, the defined model search space including one or more searchable parameters, and the first machine learning model including one or more first values for the one or more searchable parameters. The operations may include performing a model training process on the first machine learning model to obtain first training data that describes a first training speed. The operations may include generating a second machine learning model from the defined model search space at least partially based on the first training data, the second machine learning model including one or more second values for the one or more searchable parameters, and at least one of the one or more second values being different from the one or more first values. The operations may include performing a model training process on the second machine learning model to obtain second training data that describes a second training speed, the second training speed being faster than the first training speed.
[0008] Another exemplary aspect of the present disclosure is directed to one or more tangible non-transitory computer-readable media storing computer-readable instructions that, when executed by one or more processors, cause the one or more processors to perform operations. The operations may include generating a first machine learning model from a defined model search space, the defined model search space including one or more searchable parameters, the first machine learning model including one or more first values for the one or more searchable parameters. The operations may include performing a model training process on the first machine learning model to obtain first training data that describes a first training speed. The operations may include generating a second machine learning model from the defined model search space at least partially based on the first training data, the second machine learning model including one or more second values for the one or more searchable parameters, at least one of the one or more second values being different from the one or more first values, the second machine learning model including a plurality of sequential model stages, each model stage including one or more model layers, the first model stage including fewer model layers than a second model stage of the plurality of model stages. The operations may include performing a model training process on the second machine learning model to obtain second training data that describes a second training speed, the second training speed being faster than the first training speed.
[0009] Another exemplary aspect is directed to one or more tangible non-transitory computer-readable media that store a machine learning model including a first sequence of a plurality of Fused-MBConv stages and a second sequence of a plurality of MBConv stages, where the second sequence of the plurality of MBConv stages follows the first sequence of the plurality of Fused-MBConv stages, and computer-readable instructions that, when executed by one or more processors, cause the one or more processors to perform operations including obtaining a model input and processing the model input with the machine learning model to generate a model output. In some implementations, the plurality of Fused-MBConv stages consists of three Fused-MBConv stages. In some implementations, the three Fused-MBConv stages include a first, a second, and a third Fused-MBConv stage having 2, 4, and 4 layers respectively. In some implementations, the three Fused-MBConv stages include a first, a second, and a third Fused-MBConv stage having 24, 48, and 64 channels respectively. In some implementations, the three Fused-MBConv stages include a first, a second, and a third Fused-MBConv stage each having a 3×3 kernel.
[0010] Other aspects of the present disclosure are directed to various systems, devices, non-transitory computer-readable media, user interfaces, and electronic devices.
[0011] These and other features, aspects, and advantages of the various embodiments of the present disclosure will be better understood with reference to the following description and the appended claims. The accompanying drawings, which are incorporated herein and constitute a part of this specification, illustrate exemplary embodiments of the present disclosure and, together with the description, serve to explain the relevant principles.
[0012] A detailed description of embodiments directed to those of ordinary skill in the art is set forth herein with reference to the accompanying figures.
Brief Description of the Drawings
[0013]
Figure 1A
Figure 1B
Figure 1C
Figure 2A
Figure 2B
Figure 3
Figure 4
Figure 5
DETAILED DESCRIPTION OF THE INVENTION
[0014] Reference numbers repeated across multiple figures identify the same features in various implementations.
[0015] Overview Generally, the present disclosure is directed to incremental learning of machine learning models. More particularly, the present disclosure relates to incremental adjustment of regularization during training of a machine learning model to optimize training speed. By way of example, a plurality of training samples (e.g., training images, training datasets, etc.) can be obtained for a machine learning model (e.g., a convolutional neural network, a deep learning network, etc.). For one or more training iterations, the machine learning model can be trained using one or more of these training samples based on a first regularization scale. The first regularization scale can be configured to control the relative effect of one or more regularization techniques (e.g., model dropout, training data augmentation, etc.). For one or more second training iterations, the model can be trained based at least in part on a second regularization scale that is greater than the first regularization scale. Further, in some implementations, the complexity of the training samples (e.g., image size, dataset size, etc.) can be incrementally increased substantially similarly. By initially training the model with a relatively weak level of regularization and data complexity and then incrementally increasing both parameters, the systems and methods of the present disclosure substantially reduce the computational resources required during relatively early training iterations, thus increasing the accuracy of the model while simultaneously increasing the overall speed at which the model is trained.
[0016] To further increase the training efficiency of the machine learning model, the architecture of the model can be selected based on architecture search for improving training efficiency (e.g., neural architecture search, etc.). As an example, a first machine learning model including one or more parameters can be generated from a defined model search space. More specifically, the first machine learning model can include one or more values for one or more parameters. To obtain training data describing a first training speed, a model training process (e.g., progressive regularization described above, etc.) can be applied to the first machine learning model. Based at least in part on the first training data, a second machine learning model including one or more second values different from the one or more first values can be generated from the defined model search space. To obtain second training data, a model training process can be performed on the second machine learning model. The second training data can describe a second training speed that is faster than the first training speed. In that way, architecture search techniques can be used to optimize with respect to the training speed, and thus significantly increase the overall training speed of the machine learning model.
[0017] In some implementations, the complexity of the training data can be incrementally increased during training according to a regularization scale. As an example, a first complexity (e.g., image size, etc.) can be determined for a first set of training images (e.g., 280×280 pixels, etc.). The first set of training images can be used in the training process to train the machine learning model at a first regularization scale. A second complexity greater than the first can be determined for a second set of training images according to a second regularization scale greater than the first (e.g., 1080×720 pixels, etc.). The second set of training images can be used in the training process to train the machine learning model based on the second regularization scale. In this way, both accuracy and training efficiency (e.g., speed, number of training epochs, etc.) can be substantially increased.
[0018] In some implementations, the training complexity can represent or otherwise describe issues associated with processing the training samples. As an example, the training samples can be image data. Increasing the training complexity of the image data can include augmenting one or more characteristics of the image data (e.g., increasing the resolution, increasing the color of the image data, increasing the number of image features, adding noise to the image, rotating the image, augmenting the image with a second image, etc.). As another example, if the training samples are polygonal meshes, increasing the training complexity of the polygonal meshes can include increasing the number of polygons included. Thus, it should be widely understood that adjusting the training complexity of the training samples can be accomplished in any way that adjusts the issues associated with processing the training samples by each machine learning model.
[0019] In some implementations, the models of the present disclosure can be attention-based models or, otherwise, can include attention-based models. By way of example, the models of the present disclosure can include one or more self-attention mechanisms and / or one or more attention level layers. For instance, the models of the present disclosure can be Transformer models or, otherwise, can include Transformer models. In some implementations, regularization techniques can include augmenting the attention scope in these models. For example, the regularization techniques can augment the full self-attention layer in some manner (e.g., adjusting the weights of the attention, modifying the attention architecture, modifying the scope in which attention is determined between the attention heads of the layer, etc.).
[0020] The systems and methods described herein provide several technical effects and benefits. As an example, by gradually increasing the scale of regularization during training, the systems and methods of the present disclosure can substantially reduce the overall complexity and the computational resources required to train the model, for example, compared to previous training techniques that use a constant scale of regularization (e.g., less processing power, less memory usage, less power consumption, etc.). Thus, the proposed techniques can enable more efficient training of machine learning models using fewer computational resources.
[0021] Furthermore, by generating models with architectures optimized for training speed and efficiency, the systems and methods of the present disclosure can further reduce the overall time and computational resources required for model training. For example, by leveraging architecture search techniques to generate segmented machine learning models where the complexity (e.g., number of layers, etc.) increases for each segment, the systems and methods of the present disclosure can further enhance the effectiveness of progressive regularization during training and thus further reduce the computational resource requirements.
[0022] As another exemplary technical effect and benefit, the gradual adjustment of training complexity (e.g., the size of training data, the relative "problems" of correct processing of training data, etc.) is generally known to increase the training speed. However, this can also lead to a reduction in model accuracy. By gradually increasing the regularization of training in a corresponding manner, the increase in training speed from the gradual training complexity can be enhanced, and any reduction in model accuracy can be mitigated. Therefore, the gradual regularization of the model during training can enhance other methods for increasing the training speed and substantially reduce or eliminate any reduction in accuracy from those methods.
[0023] As another exemplary technical effect and benefit, meta-learning (e.g., neural architecture search, etc.) techniques often use "early stopping" techniques when generating machine learning models. For example, meta-learning techniques can implement early stopping for the cycle of training iterations based on the initial results of the training iterations. By using the gradual regularization technique, relatively early training iterations in the training cycle have a relatively low computational cost. By reducing the overall computational cost associated with the early training iterations, the systems and methods of the present disclosure can substantially reduce the adverse effects (e.g., wasted computational resources, etc.) associated with early stopping in meta-learning techniques.
[0024] Accordingly, the exemplary embodiments of the present disclosure are directed to a new family of convolutional networks that have a faster training speed and better parameter efficiency than previous models. To develop these models, the exemplary systems described herein can use a combination of training recognition neural architecture search and scaling to optimize both the training speed and parameter efficiency. The models are searched from a search space augmented with new options such as Fused-MBConv. Experiments show that the models proposed herein train much faster than state-of-the-art models and are up to 6.8 times smaller.
[0025] Furthermore, the training of the model can be further accelerated by gradually increasing the image size during training. However, this often causes a decrease in accuracy. To compensate for this decrease in accuracy, the present disclosure proposes an improved method of progressive learning that adaptively adjusts regularization (e.g., data augmentation) in addition to the image size.
[0026] More specifically, training efficiency has recently received significant attention. For example, NFNet aims to improve training efficiency by eliminating costly batch normalization, some recent studies emphasize improving the training speed by adding attention layers to convolutional networks (ConvNets), and vision transformers improve training efficiency in large-scale datasets by using transformer blocks.
[0027] However, these methods often involve costly overheads for large parameter sizes. In contrast, the present disclosure uses a combination of training recognition neural architecture search (NAS) and scaling to improve both training speed and parameter efficiency. Assuming the parameter efficiency of several models known as EfficientNet (see Tan, M. and Le, Q. V. EfficientNet: Rethinking model scaling for convolutional neural networks. ICML, 2019a), the present disclosure systematically examines the training bottlenecks in EfficientNet. These examinations show that in EfficientNet, (1) training with very large image sizes is slow, (2) depthwise unit convolutions are slow in the early layers, and (3) it is sub-optimal to scale up all stages equally.
[0028] Based on these observations, the present disclosure provides a search space augmented with additional operations such as Fused-MBConv. The present disclosure also applies training recognition NAS and scaling to optimize the model accuracy, training speed, and parameter size together. The resulting network may be called EfficientNetV2, which trains up to 4 times faster than conventional models and has a parameter size up to 6.8 times smaller.
[0029] In some exemplary implementations, training can be further accelerated by gradually increasing the image size during training. Many previous studies have used smaller image sizes for training while maintaining the same regularization for all image sizes, causing a decrease in accuracy. Therefore, it is not ideal to maintain the same regularization for different image sizes. For the same network, smaller image sizes lead to smaller network capacity and thus require weaker regularization, and vice versa, larger image sizes require stronger regularization to combat overfitting.
[0030] Based on this insight, the present disclosure proposes an improved method of progressive learning. In the early training epochs, the network may be trained with a small image size and weak regularization (e.g., dropout and data augmentation), and then the image size may be gradually increased and stronger regularization may be added. This approach can speed up training without causing a decrease in accuracy.
[0031] Using improved progressive learning, an exemplary implementation of the proposed EfficientNetV2 achieves strong results on the ImageNet, CIFAR-10, CIFAR-100, Cars, and Flowers datasets. On ImageNet, EfficientNet V2 achieves a top accuracy of 85.7%, trains 3 to 9 times faster than previous models, and is up to 6.8 times smaller. EfficientNetV2 and progressive learning also make it easier to train models on larger datasets. For example, ImageNet21k is approximately 10 times larger than ImageNet ILSVRC2012, but EfficientNetV2 can finish training within 2 days using a modest amount of computing resources of 32 TPUv3 cores. By pre-training on the publicly available ImageNet21k, EfficientNetV2 achieves a top accuracy of 87.3% on ImageNet ILSVRC2012, outperforms the latest ViT-L / 16 by 2.0% in accuracy, and trains 5 to 11 times faster.
[0032] Therefore, the present disclosure provides the advancement of EfficientNetV2, a new family of smaller and faster models. As we have found through our training-aware NAS and scaling, EfficientNetV2 outperforms previous models in both training speed and parameter efficiency. An improved method of progressive learning adaptively adjusts regularization along with the image size. This method speeds up training while improving accuracy at the same time. The proposed approach demonstrates up to 11x faster training speed and up to 6.8x better parameter efficiency than the prior art on ImageNet, CIFAR, Cars, and Flowers datasets.
[0033] Exemplary EfficientNetV2 Architecture Design In this section, we examine the training bottlenecks of EfficientNet, introduce the proposed training-aware NAS and scaling, and the EfficientNetV2 model.
[0034] Overview of EfficientNet EfficientNet is a family of models optimized for FLOPs and parameter efficiency. It utilizes NAS to search for a baseline EfficientNet-B0 with a better trade-off in accuracy and FLOPs. The baseline model is then scaled up with a compound scaling strategy to obtain model families B1 - B7. Recent research claims large gains in training or inference speed, but often has worse parameter and FLOPs efficiency than EfficientNet. The present disclosure improves the training speed while maintaining parameter efficiency.
[0035] Understanding Training Efficiency This section describes the training bottlenecks of EfficientNet (hereinafter referred to as EfficientNetV1) and several simple techniques for improving training speed. Training at very large image sizes is slow, and as pointed out by previous research, the large image sizes of EfficientNetV1 result in a large amount of memory usage. Since the total memory in GPUs / TPUs is fixed, these models are generally trained with smaller batch sizes, which significantly slows down training. A simple improvement is to apply FixRes by using a smaller image size for training than for inference. Smaller image sizes lead to less computation, enable larger batch sizes, and thus improve the training speed by up to 2.2 times. Using a smaller image size for training also leads to slightly better accuracy. In some implementations, no layer is fine-tuned after training. Depthwise separable convolutions are slow in the early layers but effective at later stages, and another training bottleneck of EfficientNetV1 comes from a wide range of depthwise separable convolutions. Depthwise separable convolutions have fewer parameters and FLOPs than normal convolutions, but often cannot fully utilize advanced accelerators. Recently, Fused-MBConv has been used to better utilize mobile or server accelerators. Fused-MBConv is described in Gupta, S. and Tan, M. EfficientNet-EdgeTPU: Creating accelerator-optimized neural networks with automl. https: / / ai.googleblog.com / 2019 / 08 / efficientnetedgetpu-creating.html, 2019. Fused-MBConv replaces the depthwise conv3×3 and expansion conv1×1 in MBConv with a single normal conv3×3.MBConv is described in Sandler et al., Mobilenetv2: Inverted residuals and linear bottlenecks. CVPR, 2018 as well as Tan, M. and Le, Q. V. EfficientNet: Rethinking model scaling for convolutional neural networks. ICML, 2019a. The Fused-MBConv and MBConv architectures are shown in FIG. 2A.
[0036] To systematically compare these two building blocks, in the exemplary experiments, the original MBConv in EfficientNet-B4 was gradually replaced with Fused-MBConv. When applied in the early stages 1-3, Fused-MBConv can improve the training speed with a small overhead in parameters and FLOPs, but when all blocks are replaced with Fused-MBConv (stages 1-7), it significantly increases the parameters and FLOPs and slows down the training as well. Finding the correct combination of these two building blocks, namely MBConv and Fused-MBConv, is not straightforward and is the problem solved by the present disclosure through the use of neural architecture search to automatically explore the best combination. Scaling up all stages equally is sub-optimal, and EfficientNetV1 scales up all stages equally using a simple compound scaling rule. For example, when the depth coefficient is 2, all stages in the network double the number of layers. However, these stages do not contribute equally to the training speed and parameter efficiency. The exemplary implementations of the present disclosure use a non-uniform scaling strategy to gradually add more layers to the later stages. Further, v1 EfficientNet aggressively scales up the image size, leading to large memory consumption and slow training. To address this problem, the exemplary implementations of the present disclosure slightly modify the scaling rule to limit the maximum image size to a smaller value.
[0037] Exemplary Training-Aware NAS and Scaling Exemplary implementations of the present disclosure provide multiple design choices for improving training speed. To explore the best combination of those choices, this section proposes training-aware NAS. NAS search, i.e., the exemplary training-aware NAS framework proposed by the present disclosure, aims to optimize accuracy, parameter efficiency, and training efficiency together in progressive accelerators. Specifically, NAS uses EfficientNet as its backbone. The search space may be a stage-based factorization space similar to Tan et al., Mnasnet: Platform-aware neural architecture search for mobile. CVPR, 2019, which consists of design choices for convolutional operation types {MBConv, Fused-MBConv}, number of layers, kernel sizes {3×3, 5×5}, and expansion ratios {1, 4, 6}. On the other hand, the search space size is optional and can be reduced by (1) removing unnecessary search options such as pooling skip options that are not used in the original EfficientNet, and (2) reusing the same channel sizes from the backbone that have already been explored in the original EfficientNet. Since the search space is smaller, the search process can apply reinforcement learning or simply random search in much larger networks with sizes comparable to EfficientNet-B4. Specifically, the exemplary search method can sample up to 1000 models and train each model for about 10 epochs with a reduced image size for training. The exemplary search reward can combine the model accuracy A, the normalized training step time S, and the parameter size P using a simple weighted product A·S w ·P v where w = -0.07 and v = -0.05 are empirically determined to balance the trade-off.
[0038] The following table shows an exemplary architecture for the model EfficientNetV2-S to be explored. Compared with the EfficientNet backbone, the proposed explored EfficientNetV2 has several major differences. (1) The first difference is that EfficientNetV2 widely uses both MBConv and the newly added Fused-MBConv in the early layers. (2) Second, since smaller expansion ratios tend to have less memory access overhead, EfficientNetV2 prefers smaller expansion ratios for MBConv. (3) Third, EfficientNetV2 prefers a smaller 3×3 kernel size, but adds more layers to compensate for the reduced receptive field resulting from the smaller kernel size. (4) Finally, EfficientNetV2 completely removes the last stride-1 stage in the original EfficientNet, probably due to the large parameter size and memory access overhead.
[0039] [Table 1]
[0040] Another exemplary architecture is shown in the following table.
[0041] [Table 2]
[0042] EfficientNetV2 Scaling: Some exemplary implementations can scale up EfficientNetV2-S to obtain EfficientNetV2-M / L using the same compound scaling as described for EfficientNetV1, with some additional optimizations, namely, (1) since very large images often lead to costly memory and training speed overheads, the maximum inference image size can be limited to 480, and (2) as a heuristic, more layers can be gradually added at later stages (e.g., stage 5 and 6) to increase network capacity without adding significant runtime overhead.
[0043] Exemplary Progressive Learning Approach Exemplary Induction As discussed in the previous section, image size plays an important role in training efficiency. In addition to FixRes, many other studies vary the image size dynamically during training, but often cause a decrease in accuracy. This decrease in accuracy may be due to unbalanced regularization, and it is also optimal to adjust the regularization strength accordingly (rather than using fixed regularization as in previous studies) when training with different image sizes. In fact, larger models generally require stronger regularization to combat overfitting. For example, EfficientNet-B7 uses a larger dropout and stronger data augmentation than B0. The present disclosure suggests that for the same network, smaller image sizes lead to smaller network capacity and thus require weaker regularization, and vice versa, larger image sizes lead to more computations at larger capacities and are thus more vulnerable to overfitting.
[0044] Exemplary Progressive Learning Using Adaptive Regularization An exemplary training process using improved progressive learning is as follows. That is, in the early training epochs, the network is trained with smaller images and weaker regularization, by which the network can easily and quickly learn simple representations. Then, the image size can be gradually increased, but by adding stronger regularization, the learning is also made more difficult.
[0045] Formally, the overall training has a total of N steps, and the target image size is the regularization scale
[0046]
Number
[0047] with a list of S e is assumed, where k represents the type of regularization such as the dropout rate or the mixup rate value. Some exemplary implementations divide the training into M stages, and for each stage 1 ≤ i ≤ M, the model is trained with the image size S i and the regularization scale
[0048]
Number
[0049] The last stage M will use the target image size S e and the regularization Φ e For simplicity, some exemplary implementations heuristically select the initial image size S 0 and the regularization Φ 0 and then use linear interpolation to determine the values for each stage. An exemplary algorithm is given below. At the beginning of each stage, the network inherits all the weights from the previous stage. Unlike transformers where weights (e.g., positional embeddings) can depend on the input length, ConvNet weights do not depend on the image size and can thus be easily inherited.
[0050] Exemplary Algorithm for Progressive Learning Using Adaptive Regularization: Input: Initial image size S 0 and regularization
[0051]
Number
[0052] 。 Input: Final image size S e and regularization
[0053]
Number
[0054] 。 Input: Total number of training steps N and stage M. for i = 0 to M - 1 do Image size:
[0055]
Number
[0056] Regularization:
[0057]
Number
[0058] Train the model
[0059]
Number
[0060] for only i steps of S i and R end for
[0061] The exemplary implementations of the proposed progressive learning are generally compatible with any existing regularization. As an example, the following types of regularization can be progressively adapted as described herein.
[0062] Dropout: Network-level regularization that reduces co-adaptation by randomly dropping channels. Progressive learning can be applied to adjust the dropout rate γ. Dropout is described in Srivastava et al., Dropout: a simple way to prevent neural networks from overfitting, The Journal of Machine Learning Research, 15(1):1929~1958, 2014.
[0063] RandAugment: Image-level data augmentation with an adjustable magnitude ε. Progressive learning can be applied to adjust the magnitude. RandAugment is described in Cubuk et al., Randaugment: Practical automated data augmentation with a reduced search space. ECCV, 2020.
[0064] Mixup: Cross-image data augmentation. Given two images with labels (x i , y i ) and (x j , y j ), they are combined with a mixup ratio λ, i.e.,
[0065]
Number
[0066] and
[0067]
Number
[0068] This results. Progressive learning can be applied to adjust the mixup ratio λ during training. Mixup is described in Zhang et al., Mixup: Beyond empirical risk minimization. ICLR, 2018.
[0069] Exemplary Devices and Systems FIG. 1A shows a block diagram of an exemplary computing system 100 for performing model training using progressive regularization, according to an exemplary embodiment of the present disclosure. The system 100 includes a user computing device 102, a server computing system 130, and a training computing system 150 that are communicatively coupled via a network 180.
[0070] The user computing device 102 can be any type of computing device, such as, for example, a personal computing device (e.g., a laptop or desktop), a mobile computing device (e.g., a smartphone or tablet), a gaming console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.
[0071] The user computing device 102 includes one or more processors 112 and a memory 114. The one or more processors 112 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.), and can be one processor or multiple processors operably connected. The memory 114 can include one or more non-transitory computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 114 can store data 116 and instructions 118 that are executed by the processor 112 to cause the user computing device 102 to perform operations.
[0072] In some implementations, the user computing device 102 can store or include one or more machine learning models 120. For example, the machine learning model 120 can be various machine learning models such as a neural network (e.g., a deep neural network) or other types of machine learning models including non-linear models and / or linear models, or can otherwise include those machine learning models. The neural network can include a feed-forward neural network, a recurrent neural network (e.g., a long short-term memory recurrent neural network), a convolutional neural network, or other forms of neural networks. Some exemplary machine learning models can utilize attention mechanisms such as self-attention. For example, some exemplary machine learning models can include a multi-head self-attention model (e.g., a transformer model).
[0073] In some implementations, one or more machine learning models 120 are received from a server computing system 130 via a network 180, stored in a user computing device memory 114, and then used by, or otherwise implemented by, one or more processors 112. In some implementations, the user computing device 102 can implement multiple parallel instances of a single machine learning model 120 (e.g., for parallel progressive regularization training across multiple instances of the model).
[0074] Additionally or alternatively, one or more machine learning models 140 can be included in a server computing system 130 that communicates with the user computing device 102 according to a client - server relationship, or otherwise stored and implemented by the server computing system 130. For example, the machine learning model 140 can be implemented by the server computing system 140 as part of a web service (e.g., a classification service, a prediction service, etc.). Thus, one or more models 120 can be stored and implemented on the user computing device 102, and / or one or more models 140 can be stored and implemented on the server computing system 130.
[0075] The user computing device 102 can also include one or more user input components 122 that receive user input. For example, the user input component 122 can be a touch - sensitive component (e.g., a touch - sensitive display screen or a touch pad) that is sensitive to the touch of a user input object (e.g., a finger or a stylus). The touch - sensitive component can be useful for implementing a virtual keyboard. Other exemplary user input components include a microphone, a conventional keyboard, or other means by which a user can provide user input.
[0076] Server computing system 130 includes one or more processors 132 and a memory 134. The one or more processors 132 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.), and can be one processor or multiple processors operably connected. The memory 134 can include one or more non-transitory computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc. and combinations thereof. The memory 134 can store data 136 and instructions 138 that are executed by the processor 132 to cause the server computing system 130 to perform operations.
[0077] In some implementations, the server computing system 130 includes one or more server computing devices or, alternatively, is implemented by a server computing device. In cases where the server computing system 130 includes multiple server computing devices, such server computing devices can operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.
[0078] As described above, the server computing system 130 can store one or more machine learning models 140, or alternatively, can include the machine learning models 140. For example, the models 140 can be, or alternatively can include, various machine learning models. Exemplary machine learning models include neural networks or other multi-layer non-linear models. Exemplary neural networks include feed-forward neural networks, deep neural networks, regression neural networks, and convolutional neural networks. Some exemplary machine learning models can utilize attention mechanisms such as self-attention. For example, some exemplary machine learning models can include multi-head self-attention models (e.g., transformer models).
[0079] The user computing device 102 and / or the server computing system 130 can train the models 120 and / or 140 through interaction with a training computing system 150 communicatively coupled via the network 180. The training computing system 150 can be separate from the server computing system 130 or can be a part of the server computing system 130.
[0080] The training computing system 150 includes one or more processors 152 and a memory 154. The one or more processors 152 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.), and can be one processor or multiple processors operably connected. The memory 154 can include one or more non-transitory computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc. and combinations thereof. The memory 154 can store data 156 and instructions 158 that are executed by the processor 152 to cause the training computing system 150 to perform operations. In some implementations, the training computing system 150 includes one or more server computing devices or, alternatively, is implemented by one or more server computing devices.
[0081] The training computing system 150 can include a model trainer 160 that trains a machine learning model 120 and / or 140 stored in the user computing device 102 and / or the server computing system 130 using various training or learning techniques such as, for example, error backpropagation. For example, a loss function can be backpropagated through the model to update one or more parameters of the model (e.g., based on the gradient of the loss function). Various loss functions can be used, such as mean squared error, likelihood loss, cross-entropy loss, hinge loss, and / or various other loss functions. Gradient descent techniques can be used to iteratively update the parameters for a number of training iterations.
[0082] In some implementations, performing backpropagation may include performing reduced backpropagation over time. The model trainer 160 can perform several generalization techniques (e.g., weight decay, dropout, etc.) to improve the generalization ability of the model being trained.
[0083] In particular, the model trainer 160 can train the machine learning model 120 and / or 140 based on a set of training data 162. By way of example, the training data 162 can include a plurality of training samples (e.g., training images, training data sets, etc.) for a machine learning model (e.g., model 120, model 140, etc.). For one or more training iterations, the machine learning model 120 / 140 can be trained (e.g., using the model trainer 160) using one or more of these training samples 162 based on a first regularization scale. The first regularization scale can be configured for the model trainer 160 to control the relative effect of one or more regularization techniques (e.g., model dropout, training data augmentation, etc.). For one or more second training iterations, the model 120 / 140 can be trained by the model trainer 160 based at least in part on a second regularization scale that is greater than the first regularization scale. Further, in some implementations, the complexity of the training samples 162 (e.g., image size, data set size, etc.) can be incrementally increased substantially similarly (e.g., using the model trainer 160). By incrementally increasing the regularization scale over the number of training iterations, the accuracy of the model 120 / 140 can be increased while simultaneously increasing the overall training efficiency of the model 120 / 140.
[0084] In some implementations, when the user has given consent, the training examples may be provided by the user computing device 102. Thus, in such implementations, the model 120 provided to the user computing device 102 may be trained by the training computing system 150 with respect to user-specific data received from the user computing device 102. In some cases, this process may be referred to as model customization.
[0085] The model trainer 160 includes computer logic used to provide the desired functionality. The model trainer 160 can be implemented in hardware, firmware, and / or software that controls a general-purpose processor. For example, in some implementations, the model trainer 160 includes a program file stored on a storage device, loaded into memory, and executed by one or more processors. In other implementations, the model trainer 160 includes one or more sets of computer-executable instructions stored on a tangible computer-readable storage medium such as RAM, a hard disk, or an optical or magnetic medium.
[0086] The network 180 can be any type of communication network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or some combination thereof, and can include any number of wired or wireless links. Generally, communication via the network 180 can be carried over any type of wired and / or wireless connection using a variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or security schemes (e.g., VPN, secure HTTP, SSL).
[0087] In some implementations, the input to the machine learning model of the present disclosure may be image data. The machine learning model may process the image data to generate an output. By way of example, the machine learning model may process the image data to generate an image recognition output (e.g., recognition of the image data, latent embedding of the image data, encoded representation of the image data, hash of the image data, etc.). As another example, the machine learning model may process the image data to generate an image segmentation output. As another example, the machine learning model may process the image data to generate an image classification output. As another example, the machine learning model may process the image data to generate an image data modification output (e.g., modification of the image data, etc.). As another example, the machine learning model may process the image data to generate an encoded image data output (e.g., an encoded and / or compressed representation of the image data, etc.). As another example, the machine learning model may process the image data to generate an upscaled image data output. As another example, the machine learning model may process the image data to generate a prediction output.
[0088] In some implementations, the input to the machine learning model of the present disclosure may be text or natural language data. The machine learning model may process the text or natural language data to generate an output. As an example, the machine learning model may process natural language data to generate a language encoding output. As another example, the machine learning model may process text or natural language data to generate a latent text embedding output. As another example, the machine learning model may process text or natural language data to generate a conversion output. As another example, the machine learning model may process text or natural language data to generate a classification output. As another example, the machine learning model may process text or natural language data to generate a text segmentation output. As another example, the machine learning model may process text or natural language data to generate a semantic intent output. As another example, the machine learning model may process text or natural language data to generate an upscaled text or natural language output (e.g., text or natural language data that is of higher quality than the input text or natural language, etc.). As another example, the machine learning model may process text or natural language data to generate a prediction output.
[0089] In some implementations, the input to the machine learning model of the present disclosure may be audio data. The machine learning model may process the audio data to generate an output. As an example, the machine learning model may process the audio data to generate an audio recognition output. As another example, the machine learning model may process the audio data to generate an audio conversion output. As another example, the machine learning model may process the audio data to generate a latent embedding output. As another example, the machine learning model may process the audio data to generate an encoded audio output (e.g., an encoded and / or compressed representation of the audio data, etc.). As another example, the machine learning model may process the audio data to generate an upscaled audio output (e.g., audio data of higher quality than the input audio data, etc.). As another example, the machine learning model may process the audio data to generate a text representation output (e.g., a text representation of the input audio data, etc.). As another example, the machine learning model may process the audio data to generate a prediction output.
[0090] In some implementations, the input to the machine learning model of the present disclosure may be latent encoded data (e.g., a latent space representation of the input, etc.). The machine learning model may process the latent encoded data to generate an output. As an example, the machine learning model may process the latent encoded data to generate a recognition output. As another example, the machine learning model may process the latent encoded data to generate a reconstruction output. As another example, the machine learning model may process the latent encoded data to generate a search output. As another example, the machine learning model may process the latent encoded data to generate a reclustering output. As another example, the machine learning model may process the latent encoded data to generate a prediction output.
[0091] In some implementations, the input to the machine learning model of the present disclosure may be statistical data. Statistical data can be data that is calculated and / or derived from some other data source, represents this, or otherwise includes this. The machine learning model can process the statistical data to generate an output. As an example, the machine learning model can process the statistical data to generate a recognition output. As another example, the machine learning model can process the statistical data to generate a prediction output. As another example, the machine learning model can process the statistical data to generate a classification output. As another example, the machine learned model can process the statistical data to generate a segmentation output. As another example, the machine learning model can process the statistical data to generate a visualization output. As another example, the machine learning model can process the statistical data to generate a diagnostic output.
[0092] In some implementations, the input to the machine learning model of the present disclosure may be sensor data. The machine learning model can process the sensor data to generate an output. As an example, the machine learning model can process the sensor data to generate a recognition output. As another example, the machine learning model can process the sensor data to generate a prediction output. As another example, the machine learning model can process the sensor data to generate a classification output. As another example, the machine learning model can process the sensor data to generate a segmentation output. As another example, the machine learning model can process the sensor data to generate a visualization output. As another example, the machine learning model can process the sensor data to generate a diagnostic output. As another example, the machine learning model can process the sensor data to generate a detection output.
[0093] In some cases, the machine learning model can be configured to perform tasks that include encoding input data for reliable and / or efficient transmission or storage (and / or corresponding decoding). For example, the task may be an audio compression task. The input may include audio data and the output may include compressed audio data. In another example, the input includes visual data (e.g., one or more images or videos), the output includes compressed visual data, and the task is a visual data compression task. In another example, the task may include generating an embedding for the input data (e.g., input audio or visual data).
[0094] In some cases, the input includes visual data and the task is a computer vision task. In some cases, the input includes pixel data for one or more images and the task is an image processing task. For example, the image processing task may be image classification, where the output is a set of scores, each score corresponding to a different object class and representing the likelihood that one or more images show an object belonging to the object class. The image processing task may be object detection, where the image processing output identifies one or more regions in one or more images and, for each region, the likelihood that the region shows a target object. As another example, the image processing task may be image segmentation, where the image processing output defines, for each pixel in one or more images, a respective likelihood for each category in a predetermined set of categories. For example, the set of categories may be foreground and background. As another example, the set of categories may be object classes. As another example, the image processing task may be depth estimation, where the image processing output defines a respective depth value for each pixel in one or more images. As another example, the image processing task may be motion estimation, where the network input includes a plurality of images and the image processing output defines, for each pixel in one of the input images, the motion of the scene shown at the pixel between the images in the network input.
[0095] In some cases, the input includes audio data representing speech, and the task is a speech recognition task. The output may include a text output mapped to the speech. In some cases, the task includes encrypting or decrypting the input data. In some cases, the task includes microprocessor-implemented tasks such as branch prediction or memory address translation.
[0096] FIG. 1A shows one exemplary computing system that can be used to implement the present disclosure. Other computing systems may be used. For example, in some implementations, the user computing device 102 may include a model trainer 160 and a training dataset 162. In such implementations, the model 120 can be both trained locally and used on the user computing device 102. In some of such implementations, the user computing device 102 can implement the model trainer 160 to customize the model 120 based on user-specific data.
[0097] FIG. 1B shows a block diagram of an exemplary computing device 10 that performs model training using progressive regularization according to an exemplary embodiment of the present disclosure. The computing device 10 may be a user computing device or a server computing device.
[0098] The computing device 10 includes several applications (e.g., applications 1 to N). Each application includes its own machine learning library and machine learning-based model. For example, each application may include a machine learning-based model. Exemplary applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, and the like.
[0099] As shown in FIG. 1B, each application can communicate with some other components of the computing device, such as one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, each application can communicate with each device component using an API (e.g., a public API). In some implementations, the API used by each application is specific to that application.
[0100] FIG. 1C shows a block diagram of an exemplary computing device 50 that implements the generation of a machine learning model with an optimized training speed, according to an exemplary embodiment of the present disclosure. The computing device 50 may be a user computing device or a server computing device.
[0101] The computing device 50 includes several applications (e.g., applications 1 to N). Each application communicates with a central intelligence layer. Exemplary applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc. In some implementations, each application can communicate with the central intelligence layer (and the models stored therein) using an API (e.g., a common API across all applications).
[0102] The central intelligence layer includes several machine learning models. For example, as shown in FIG. 1C, each machine learning model can be provided to each application and managed by the central intelligence layer. In other implementations, two or more applications can share a single machine learning model. For example, in some implementations, the central intelligence layer can provide a single model to all applications. In some implementations, the central intelligence layer is included in or otherwise implemented by the operating system of the computing device 50.
[0103] The central intelligence layer can communicate with the central device data layer. The central device data layer can be a centralized repository of data for the computing device 50. As shown in FIG. 1C, the central device data layer can communicate with some other components of the computing device, such as one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).
[0104] Exemplary Model Arrangement FIG. 2B shows a block diagram of an exemplary machine learning model 202 generated through an architecture exploration technique for increasing training speed according to an exemplary embodiment of the present disclosure. In some implementations, the machine learning model 202 receives a set of input data 204 (e.g., image data, statistical data, video data, etc.) and is trained to provide output data 210 as a result of receiving the input data 204. More specifically, the machine learning model 202 can include multiple stages. The machine learning model 202 can include a first model stage 206 and a second model stage 208.
[0105] More specifically, the machine learning model 202 can be generated using a neural architecture search method configured to increase the training speed. As an example, the model can include two stages (e.g., a first model stage 206, a second model stage 208, etc.). The first model stage 206 can include two model layers 206A and 206B (e.g., a convolutional layer, a fused convolutional layer, etc.). The second model stage 208 can include three layers 208A-208C.
[0106] More specifically, the architecture of the machine learning model 202 can be configured to facilitate model training using a progressively increasing regularization scale (e.g., the degree to which a regularization technique affects the training of the model 202, etc.). As an example, the machine learning model 202 can be trained using an initial regularization scale corresponding to the number of layers (e.g., 206A, 206B, etc.) included in the first model stage 206. The regularization scale can increase to a second scale that is larger than the initial scale and corresponds to the number of layers (e.g., 208A, 208B, 208C, etc.) included in the second model stage 208. In that way, the architecture of the model 202 can include multiple stages each including a number of layers corresponding to a particular scale of regularization, and thus increase the effectiveness of progressive regularization during training.
[0107] FIG. 3 shows a graphical diagram 300 of an exemplary neural architecture search technique for increasing training speed and accuracy according to an exemplary embodiment of the present disclosure. More specifically, a training controller 302 can determine the scale of regularization to be implemented in a trainer 304. For example, the controller 302 can progressively regularize a training process (e.g., changing training data, model dropout, etc.) implemented by the trainer 304 over a number of iterations. The trainer 304 can train a machine learning model 306 for a number of iterations with the regularization scale indicated by the controller 302.
[0108] After an initial number of training iterations, the multi-reward objective function 308 can evaluate the accuracy and training time of the machine learning model 306 (e.g., as reported by the trainer 304, etc.). Based on this evaluation, the multi-reward objective function 308 may provide feedback to the controller 302, and the controller 302 can subsequently adjust the scale of regularization for additional training iterations of the machine learning model 306.
[0109] As an example, the controller 302 can obtain a plurality of training samples and provide the training samples to the trainer 304. Next, the controller 302 can determine an initial scale of regularization for implementation in the trainer 304 that is configured to control the relative effectiveness of one or more regularization techniques (e.g., model dropout, training data augmentation, etc.). For example, the controller can determine a relatively weak scale of regularization for training and then gradually increase the regularization over time. In some implementations, the controller 302 can also determine the initial complexity of the training samples corresponding to the scale of regularization. As an example, the controller 302 can determine a relatively weak level of training sample complexity corresponding to a relatively weak level of regularization. For example, if the training samples include image data, the controller can downscale the images (e.g., from 800×600 to 80×60, etc.) to reduce the overall level of complexity.
[0110] Trainer 304 can train machine learning model 306 based at least in part on a regularization scale determined by controller 302 for one or more initial training iterations. After the one or more initial training iterations, trainer 304 can use multi-reward objective function 308 to provide model accuracy data and training speed data for evaluation. In some implementations, multi-reward objective function 308 can be configured to reward both training speed and accuracy, thereby increasing the training speed while maintaining a particular, threshold level of accuracy.
[0111] Based on this evaluation, feedback can be provided to controller 302. Depending on the feedback, controller 302 can then incrementally increase the regularization scale in trainer 304 and the complexity of the training samples. Controller 302 can then incrementally increase the regularization and sample complexity implemented by trainer 304 during model training. Thus, by starting with relatively weak regularization and complexity, trainer 304 can train machine learning model 306 substantially faster than would be possible using a static level of regularization and sample complexity, thereby increasing the training speed and substantially reducing the associated computational resource cost.
[0112] Figure 4 shows a dataflow diagram 400 of an exemplary method for generating a machine learning model with an optimized training speed. More specifically, a model search architecture 402 (e.g., a neural search architecture, etc.) can include or otherwise access a defined model search space 402A. The defined model search space 402A can be or include one or more searchable parameters (e.g., number of layers, type of layer (e.g., convolutional layer, fused convolutional layer, fused MB-CONV layer, etc.), learning rate, hyperparameters, layer size, channel size, kernel size, etc.). It should be widely understood that one or more searchable parameters of the defined model search space 402A can define or otherwise control any particular aspect or implementation of a machine learning model (e.g., 404, 410, etc.).
[0113] Based on the defined model search space 402A, the model search architecture 402 can generate a first machine learning model 404. The first machine learning model 404 can include one or more values for one or more searchable parameters. The first machine learning model 404 can then be trained using a model training process 406 to obtain first training data. The first training data can describe a first training speed (e.g., the amount of time required for training, etc.) for the training process 406 implemented on the first machine learning model 404.
[0114] The first training data 408 may be provided to the model search architecture 402. Based on the first training data 408, the model search architecture can generate a second machine learning model 410 from the defined model search space 402A. The second machine learning model 410 may include one or more values for one or more searchable parameters that are different from one or more values of the first machine learning model 404. As an example, one value of the first machine learning model 404 for a searchable parameter may define that a first type of layer (e.g., a standard convolutional layer, etc.) is used in the first machine learning model 404. The second machine learning model may, instead, include a value for the same parameter that defines that a different type of layer (e.g., a fused convolutional layer, etc.) is used in the second machine learning model 410.
[0115] The second machine learning model 410 may be trained using the same model training process 406 to obtain second training data 412. The second training data 412 may describe a second training speed that is faster than the first training speed described by the first training data 408. In that way, the model search architecture can repeatedly generate machine learning models with increasingly faster training speeds by sampling the defined model search space 402A, and thus generate a machine learning model (e.g., the second machine learning model 410, etc.) with optimal training speed characteristics.
[0116] FIG. 5 shows a flowchart diagram of an exemplary method implemented in accordance with an exemplary embodiment of the present disclosure. FIG. 5 shows steps implemented in a specific order for purposes of explanation and discussion, but the methods of the present disclosure are not limited to the specific order or sequence shown. The various steps of method 500 may be omitted, reordered, combined, and / or adapted in various ways without departing from the scope of the present disclosure.
[0117] In 502, the computing system can obtain a plurality of training samples for a machine learning model. More specifically, the computing system can obtain a plurality of training samples (e.g., image data, dataset data, etc.) of a certain complexity that can be used to train a machine learning model for one or more tasks.
[0118] In 504, the computing system can train a machine learning model for one or more first training iterations based at least in part on a first regularization scale. More specifically, for one or more first training iterations, the computing system can train the machine learning model using one or more respective first training samples of the plurality of training samples based at least in part on a first regularization scale configured to control the relative effectiveness of one or more regularization techniques.
[0119] In 506, the computing system can train a machine learning model for one or more second training iterations based at least in part on a second regularization scale. More specifically, for one or more second training iterations, the computing system can train the machine learning model using one or more respective second training samples of the plurality of training samples based at least in part on a second regularization scale that is larger than the first regularization scale.
[0120] Additional Disclosure The technology described in this specification refers to servers, databases, software applications, and other computer-based systems, as well as actions taken and information sent between such systems. The inherent flexibility of computer-based systems allows for a wide variety of possible configurations, combinations, and divisions of tasks and functions among components. For example, the processes described in this specification can be implemented using a single device or component, or multiple devices or components operating in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.
[0121] This subject matter has been described in detail with respect to various specific exemplary embodiments, but each example is provided for illustration rather than limitation of the disclosure. Those skilled in the art, upon understanding the above, can readily create modifications, variations, and equivalents of such embodiments. Accordingly, the disclosure does not exclude the inclusion of such changes, variations, and / or additions to the subject matter as will be readily apparent to those skilled in the art. For example, features illustrated or described as part of one embodiment can also be used with another embodiment to yield further embodiments. Accordingly, the disclosure is intended to cover such modifications, variations, and equivalents.
Explanation of Reference Numerals
[0122] 10 Computing Device 50 Computing Device 100 Computing System 102 User Computing Device 112 Processor 114 Memory, Computing Device Memory 120 Model, Machine Learning Model 122 User Input Component 130 Server Computing System 132 Processor 134 Memory 140 Model, Machine Learning Model 150 Training Computing System 152 Processor 154 Memory 160 Model Trainer 180 Network 202 Machine Learning Model 302 Training Controller 304 Trainer 306 Machine Learning Model 402 Model Search Architecture 404 Machine Learning Model 410 Machine Learning Model
Claims
1. 1. A computer-implemented method for efficient machine learning model training, comprising: obtaining, by a computing system comprising one or more computing devices, a plurality of training samples for a machine learning based model; For one or more first training iterations, training, by the computing system, the machine learning based model using a first training sample of one or more of the plurality of training samples based at least in part on a first regularization magnitude configured to control a relative effect of one or more regularization techniques; For one or more second training iterations, and training, by the computing system, the machine learning model using one or more respective second training samples of the plurality of training samples based at least in part on a second regularization scale that is greater than the first regularization scale.
2. obtaining the plurality of training samples for the machine learning model further includes determining, by the computing system, a first sample complexity for the one or more first training samples; 2. The computer-implemented method of claim 1, wherein prior to training the machine learning model with the one or more respective second training samples, the method includes determining, by the computing system, a second sample complexity for the one or more second training samples, the second sample complexity being greater than the first sample complexity.
3. The plurality of training samples includes a respective plurality of training images; 3. The computer-implemented method of claim 2, wherein determining the second sample complexity for the one or more second training samples comprises adjusting, by the computing system, a size of one or more second training images, wherein the size of the one or more second training images is greater than a size of one or more first training images.
4. Prior to obtaining the plurality of training samples for the machine learning model, generating, by the computing system using a machine learning model search architecture, an initial machine learning model including one or more first values for one or more respective parameters; determining, by the computing system, a first training speed of the initial machine learning based model; generating, by the computing system using the machine learning model exploration architecture, the machine learning model including one or more second values for the one or more respective parameters, at least one of the one or more second values being different from the one or more first values.
5. 5. The computer-implemented method of claim 4, further comprising determining, by the computing system, a second training rate of the machine learning based model, the second training rate being greater than the first training rate.
6. 6. The computer-implemented method of claim 4 or 5, wherein the machine learning model includes a plurality of sequential model stages, each model stage including one or more layers, and a first model stage including fewer layers than a second model stage of the plurality of model stages.
7. The one or more regularization techniques include: adjusting, by the computing system, a number of model channels of at least one layer of the machine learning model; or 7. The computer-implemented method of claim 1, further comprising at least one of adjusting, by the computing system, at least one characteristic of one or more training samples of the plurality of training samples.
8. 2. The computer-implemented method of claim 1, wherein the second regularization magnitude is based at least in part on one or more respective training outputs from the one or more first training iterations.
9. 1. A computing system for determining a model having an optimized training rate, comprising: one or more processors; and one or more tangible, non-transitory computer-readable media storing computer-readable instructions that, when executed by the one or more processors, cause the one or more processors to perform operations, including: generating a first machine learning model from a defined model search space, the defined model search space including one or more searchable parameters, the first machine learning model including one or more first values for the one or more searchable parameters; performing a model training process on the first machine learning model to obtain first training data describing a first training speed; generating a second machine learning based model from the defined model search space based at least in part on the first training data, the second machine learning based model including one or more second values for the one or more explorable parameters, at least one of the one or more second values being different from the one or more first values; and performing the model training process on the second machine learning model to obtain second training data describing a second training rate, the second training rate being faster than the first training rate.
10. The model layers of the defined model search space include: Convolutional layers, or The computing system of claim 9 , further comprising at least one of a fused convolutional layer.
11. 11. The computing system of claim 9 or 10, wherein the second machine learning model includes a plurality of sequential model stages, each model stage including one or more model layers, and a first model stage includes fewer model layers than a second model stage of the plurality of model stages.
12. 12. The computing system of claim 9, wherein the first training data further describes a first training accuracy, and the second training data further describes a second training accuracy that is greater than the first training accuracy.
13. 13. The computing system of claim 12, wherein generating the second machine learning based model from the defined model search space is further based at least in part on the first training accuracy.
14. Performing a model training process on the first machine learning model includes: Obtaining a plurality of training samples for the first machine learning model; For one or more first training iterations, training the first machine learning based model using one or more respective first training samples of the plurality of training samples based at least in part on a first regularization magnitude configured to control a relative effect of one or more regularization techniques; For one or more second training iterations, and training the first machine learning based model using one or more respective second training samples of the plurality of training samples based at least in part on a second regularization scale that is greater than the first regularization scale.
15. The plurality of training samples includes a respective plurality of training images; 15. The computing system of claim 14, wherein determining a second sample complexity for the one or more second training samples comprises adjusting a size of one or more second training images, the size of the one or more second training images being greater than a size of the one or more first training images.
16. The plurality of training samples includes a respective plurality of training images; 16. The computing system of claim 15, wherein determining the second sample complexity for the one or more second training samples comprises adjusting a size of one or more second training images, the size of the one or more second training images being greater than a size of one or more first training images.
17. 17. The computing system of claim 9, wherein at least one of the one or more parameters is configured to control one or more of a type of model layer or a number of model layers included in the machine learning based model.
18. 18. The computing system of claim 9, wherein the operations further include providing the second machine learning based model as an output.
19. 19. The computing system of claim 9, wherein generating the second machine learning model from the defined model search space further comprises determining a plurality of sequential processing stages for the second machine learning model, each of the plurality of sequential processing stages being associated with one or more model layers, and a number of model layers associated with a second processing stage of the plurality of processing stages being greater than a number of model layers associated with a first processing stage of the plurality of processing stages.
20. One or more tangible, non-transitory computer-readable media, The first sequence of multiple Fused-MBConv stages, and a machine learning model including a second sequence of multiple MBConv stages, the second sequence of multiple MBConv stages following the first sequence of the multiple Fused-MBConv stages; and computer readable instructions that, when executed by one or more processors, cause the one or more processors to perform operations, the operations including: Obtaining model input; and processing the model inputs with the machine learning based model to generate model outputs.
21. 21. The one or more tangible, non-transitory computer-readable media of claim 20, wherein the plurality of Fused-MBConv stages consists of three Fused-MBConv stages.
22. 22. The one or more tangible, non-transitory computer-readable media of claim 21, wherein the three Fused-MBConv stages include first, second, and third Fused-MBConv stages having 2, 4, and 4 layers, respectively.
23. 23. The one or more tangible, non-transitory computer-readable media of claim 21 or 22, wherein the three Fused-MBConv stages include first, second, and third Fused-MBConv stages having 24, 48, and 64 channels, respectively.
24. 24. The one or more tangible, non-transitory computer-readable media of claim 21, wherein the three Fused-MBConv stages include first, second, and third Fused-MBConv stages each having a 3x3 kernel.
Citation Information
Patent Citations
Compound model scaling for neural networks
US20200234132A1
CLR、2018