Systems and methods for pre-training artificial intelligence networks

CN122840145APending Publication Date: 2026-09-29CENT FOR INTELLIGENT MULTIDIMENSIONAL DATA ANALYSIS LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510499861.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-03-26
Filing Date
2025-04-21
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

这种方法并未显著提高训练时间和准确性

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122840145A_ABST
    Figure CN122840145A_ABST
Patent Text Reader

Abstract

A system for training an AI model on a training dataset includes: a processor for executing instructions; and a computer-readable medium including instructions for execution on the processor. The instructions cause the processor to perform the following steps: dividing the training dataset into multiple layers according to category labels; wherein each layer represents a single label; creating subset data by selecting samples from each layer; training the AI ​​model using the subset data until a first stopping condition is met; and training the AI ​​model using the training dataset until a second stopping condition is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of Artificial Intelligence (AI) and provides a method and system for pre-training AI networks. In particular, this disclosure provides an efficient subset pre-training system and method that reduces the training time of deep network models. Background Technology

[0002] Inspired by the structure of the human brain, neural networks can be traced back to the late 1950s and early 1960s, with the perceptron, or neuron, proposed by Frank Rosenblatt forming its prototype. With continuous technological advancements, neural networks have made significant progress, becoming a crucial technology in modern AI, widely used in pattern recognition, data classification, and prediction, achieved through interconnected layers of neurons.

[0003] Neural networks typically have a hierarchical structure, including an input layer, multiple hidden layers, and an output layer. Each neuron in each layer processes the input signal through an activation function, such as the Rectified Linear Unit (ReLU) or the sigmoid function, and outputs the result to the next layer. The connections between neurons are characterized by parameters (i.e., weights and biases), which are iteratively fine-tuned during training using optimization algorithms such as gradient descent. This iterative process aims to minimize the prediction error of the neural network, thereby continuously improving the model's performance.

[0004] The goal of the training process is to enable the neural network to have good generalization ability, thereby accurately identifying patterns and making reliable predictions on previously unseen data. This ensures the robustness and adaptability of the model in practical applications. Optimizing the above parameters is crucial for the neural network to learn the potential distribution patterns of the data and achieve high-precision predictions.

[0005] From early perceptrons or basic neuron structures to advanced deep learning frameworks, artificial neural networks have undergone significant technological advancements in computational methods. The advent of the backpropagation algorithm enabled multi-layered networks to optimize complex functions, driving breakthroughs in techniques such as Convolutional Neural Networks (CNNs) for image processing and Recurrent Neural Networks (RNNs) for sequential data analysis. Although the development of deep learning has made it possible to build models with a large number of layers (i.e., deep networks) and has surpassed traditional machine learning techniques in performance, training these deep learning models still incurs a high computational burden when the training data is large, typically requiring high-performance computing resources such as Graphics Processing Units (GPUs) or Tensor Processing Units (TPUs).

[0006] Limited by high-performance computing resources, deep learning training is often time-consuming, which may hinder the widespread application of deep learning technology, especially in small institutions or individual researchers with limited computing power, or in scenarios requiring instant response, such as real-time systems. Furthermore, many challenges remain in model training, such as overfitting, underfitting, and vanishing / exploding gradients. For example, traditional CNNs demonstrate excellent performance in many image classification tasks, but as the number of network layers increases, the vanishing gradient problem easily occurs, making effective training of deep models more difficult.

[0007] With the continuous development of deep network technology, a variety of innovative methods have been proposed to address the difficulties encountered in deep learning training. One existing technique is the Residual Network (ResNet) architecture with residual connections (also known as skip connections or shortcut connections). This architecture, by introducing residual connections, allows gradients to propagate more easily throughout the network during training, thus mitigating the performance degradation problem in deep networks, enabling efficient training of very deep networks, and improving the network's prediction accuracy. Furthermore, residual connections can also be viewed as a form of regularization, helping to reduce overfitting and thus improving the network's robustness. Another advanced deep network architecture is the Densely Connected Network (DenseNet), which connects each layer to all the layers before it in a cascaded manner, forming a set of cascaded feature maps. This structure enhances feature reuse and improves the transmission of information and gradients within the network.

[0008] In recent years, various techniques have been developed to shorten the training time of deep neural networks (such as ResNet and DenseNet) to address the increasing complexity and scale of modern models. These methods can be broadly categorized into three types: pre-training optimization (e.g., transfer learning, knowledge distillation), in-training optimization (e.g., batch normalization, network model pruning, learning rate adjustment, parallel and distributed training, early stopping), and alternative learning paradigms (e.g., curriculum learning). Pre-training and in-training optimizations are typically performed on the basis of processing the complete training data as a whole, while curriculum learning employs a method of training the model first with simple samples and then gradually transitioning to complex samples.

[0009] In one existing technique, the determination and selection of "simpler examples" and "more complex examples" are correlated with the network's test error criteria. This approach has not significantly improved training time and accuracy. Therefore, a better method for training AI models is needed to improve the efficiency of machine learning. Summary of the Invention

[0010] According to a first aspect of this disclosure, a subset pre-training method is provided for reducing the overall training time of deep neural networks.

[0011] According to one aspect of this disclosure, a system for training an AI model on a training dataset is provided. The system includes:

[0012] A processor is used to execute instructions;

[0013] A computer-readable medium comprising instructions for execution on the processor, the instructions causing the processor to perform the following steps:

[0014] The training dataset is divided into multiple layers based on the category labels; each layer represents a single label.

[0015] Subset data is created by selecting samples from each of the layers;

[0016] The AI ​​model is trained using the subset of data until a first stopping condition is met; and

[0017] The AI ​​model is trained using the training dataset until the second stopping condition is met.

[0018] In a preferred embodiment, the subset data is created by selecting a preset percentage α of samples from each layer.

[0019] In a preferred embodiment, the subset data is created by selecting random samples from each layer.

[0020] In a preferred embodiment, the subset data is created by selecting samples based on the information content of samples in each layer.

[0021] In a preferred embodiment, the subset data is created by selecting samples based on ascending information content.

[0022] In a preferred embodiment, the subset data is created by selecting samples based on descending information content.

[0023] In a preferred embodiment, the information content ranking of samples in each layer is determined by the following steps:

[0024] The samples in each layer are converted into vectors, and then organized into matrix X. i ;

[0025] Perform a QR decomposition method with column pivoting to rearrange X i Columns to obtain the matrix Make XP = QR or X :,p =QR, where P represents a square binary permutation matrix, Q represents an orthogonal matrix, R represents an upper triangular matrix, and X... :,p The first column X :,p(1) The last column, X, contains the most information. :,p(|D|) This is the column with the least amount of information.

[0026] In a preferred embodiment, a permutation matrix P is chosen such that the diagonal elements of R are non-increasing: r 11 ≥r 22 ≥…,≥r nn P is a square binary matrix.

[0027] In a preferred embodiment, the data matrix X used for image recognition is associated with Sobel edge detection.

[0028] In a preferred embodiment, the data matrix X used for image recognition is associated with a Gabor filter.

[0029] In a preferred embodiment, the first stopping condition is considered to have been met when the AI ​​model has reached the maximum number of epochs.

[0030] In a preferred embodiment, the second stopping condition is considered to have been met when the AI ​​model has reached the convergence threshold.

[0031] In a preferred embodiment, the percentage α is between 10% and 50%.

[0032] In a preferred embodiment, the percentage α is between 20% and 30%.

[0033] In a preferred embodiment, the AI ​​model is a deep learning artificial neural network model.

[0034] In a preferred embodiment, the deep learning artificial neural network model is a ResNet model.

[0035] In a preferred embodiment, the deep learning artificial neural network model is the DenseNet model.

[0036] In a preferred embodiment, the step of training the AI ​​model includes: adjusting the loss function on the subset. The minimized steps make

[0037]

[0038] Where f(x; θ) represents the model's prediction, and y represents the true label. Let θ represent the loss function being optimized, and let θ represent the model parameters.

[0039] In a preferred embodiment, the step of training the AI ​​model includes updating the model parameters.

[0040] In a preferred embodiment, the model parameters include the weights or biases of the neuron nodes.

[0041] In a preferred embodiment, the step of training the AI ​​model includes the following steps: selecting samples of the subset data based on QR decomposition with column principal components, and then using the entire complete dataset as the training dataset.

[0042] In a preferred embodiment, the subset data is adaptively and progressively selected from 20% to 100% of the size of the entire complete dataset (used for training). Attached Figure Description

[0043] Specific embodiments of this disclosure will now be described by way of example with reference to the accompanying drawings.

[0044] Figure 1 This is a schematic diagram of a computer system for implementing a method for pre-training an AI network according to an embodiment of the present disclosure.

[0045] Figure 2 A processing flowchart of a method for pre-training an AI network according to an embodiment of this disclosure is shown.

[0046] Figure 3 For pre-training AI networks according to embodiments of this disclosure to perform Figure 2 A schematic diagram of the process. Detailed Implementation

[0047] refer to Figure 1 and Figure 2 The following illustration shows embodiments of the present disclosure. These embodiments provide a system and method for pre-training AI networks (artificial intelligence networks). The present disclosure provides an efficient subset pre-training method and system for reducing the training time of deep network models.

[0048] In this exemplary embodiment, the interface and processor are implemented by a computer with a suitable user interface. The computer can be implemented by any computing architecture, including: a portable computer, a tablet computer, a personal computer (PC), a smart device, an Internet of Things (IoT) device, an edge computing device, a client / server architecture, a "dumb" terminal / host architecture, a cloud-based architecture, or any other suitable architecture. The computing device can be appropriately programmed to implement this disclosure.

[0049] refer to Figure 1 A schematic diagram of a computer system or server 100 for implementing a system for effective subset pre-training to reduce the training time of deep network models in AI systems, according to embodiments of the present disclosure, is provided.

[0050] In this embodiment, a server 100 is included. Server 400 includes suitable components required to receive, store, and execute appropriate computer instructions. These components may include: a processing unit 102, which includes a Central Processing Unit (CPU), a Math Co-processing Unit (Math Processor), a Graphics Processing Unit (GPU), a Neural Processing Unit (NPU), or a Tensor Processing Unit (TPU) for tensor or multidimensional array computation or manipulation operations; read-only memory (ROM) 104; random-access memory (RAM) 106; input / output (I / O) devices (e.g., disk drive 108); input devices 110 (e.g., Ethernet port, USB port, etc.); a display 112, such as a liquid crystal display, a light-emitting display, or any other suitable display; and a communication link 114. Server 100 may include instructions that can be stored in ROM 104, RAM 106, or disk drive 108 and executed by processing unit 102. Multiple communication links 114 may be provided, which may be connected differently to one or more computing devices, such as servers, personal computers, terminals, wireless or handheld computing devices, IoT devices, smart devices, and edge computing devices. At least one of the multiple communication links may be connected to an external computing network via a telephone line or other types of communication link.

[0051] Server 100 may also include storage devices such as disk drives 108, which may include solid-state drives, hard disk drives, optical drives, tape drives, or remote or cloud-based storage devices. Server 100 may use a single disk drive or multiple disk drives, or remote storage services. Server 100 may also have a suitable operating system 116 residing on the disk drives or in the ROM of server 100.

[0052] In one embodiment of this disclosure, a novel subset pre-training method is provided to reduce the total training time of deep neural networks (classic ResNet and DenseNet architectures) while maintaining the accuracy of the test dataset.

[0053] The method disclosed herein focuses only on the selection of a subset of training data, and can be seamlessly integrated with various machine learning techniques, including pre-training strategies and training time optimization, to further enhance model efficiency and performance.

[0054] According to one embodiment of this disclosure, a system for training an AI model on a training dataset is provided, comprising:

[0055] A processor is used to execute instructions;

[0056] A computer-readable medium comprising instructions for execution on a processor, the instructions causing the processor to perform the following steps:

[0057] The training dataset is divided into multiple layers based on the category labels; each layer represents a single label.

[0058] Subset data is created by selecting samples from each layer;

[0059] Train the AI ​​model using a subset of data until the first stopping condition is met; and

[0060] The AI ​​model is trained using the training dataset until the second stopping condition is met.

[0061] A subset of data is created by selecting random samples from each layer. This subset is created by selecting samples based on descending information content sorting. The information content sorting of samples in each layer is determined by first converting each layer's samples into vectors, and then organizing them into a matrix using a permutation matrix. The permutation matrix P can be chosen such that the diagonal elements of R are non-increasing: r 11 ≥r 22 ≥…,≥r nn P and P are square binary matrices.

[0062] When the preset convergence threshold or the maximum number of epochs is reached, it is considered that the stopping condition or criterion has been met.

[0063] In a preferred embodiment, a subset of data is created by selecting a predetermined percentage α of samples from each layer. The percentage α is between 10% and 50%, or more specifically between 20% and 30%. The AI ​​model is a deep learning artificial neural network model, such as ResNet or DenseNet. The AI ​​model includes the step of minimizing a loss function on the training data (the entire set of data or a subset of the data). The model parameters of the AI ​​model, such as the weights and biases of the neurons in the AI ​​neural network, are updated based on the loss function or other learning algorithms.

[0064] In a preferred embodiment, the method involves selecting samples from a subset of data using QR decomposition with column pivoting (QRCP), and then using the entire dataset as training data. This method enables system 100 to use QRCP for pre-training sample selection and then fine-tuning using the complete dataset.

[0065] In another embodiment, system 100 is adapted to perform incremental pre-training. For example, it selects data samples incrementally—starting with 20% for initial training, then 40%, 60%, 80%, and finally 100% of the dataset for further training. While this approach may not offer significant efficiency advantages when dealing with small datasets, it can be more effective for large datasets (e.g., datasets with more than 100 parameters or millions of data points). In one embodiment, a subset of data is adaptively selected incrementally from 20% to 100% for training, based on the size of the entire dataset.

[0066] Compared to other publicly disclosed training methods, the methods of this disclosure require significant computational and time resources, especially when using large datasets. This disclosure novelly adds a stage for training only a subset of the entire training dataset, followed by a traditional full training stage using the entire training dataset. Experimental results show that this subset pre-training strategy significantly reduces the total training time required to reach a preset accuracy threshold by 50%. A detailed analysis of the methodology, experimental setup, and results highlights the advantages of this invention in deep learning model training.

[0067] The subset pre-training method disclosed herein comprises two main stages:

[0068] Initial training on a selected subset of the training dataset, and

[0069] Full training on the complete dataset.

[0070] The goal of this method is to reduce overall training time while maintaining the accuracy of standard details regarding subset selection, training phases, and hyperparameter tuning, as well as the expected theoretical benefits, such as faster convergence, as described below.

[0071] A. Subset selection

[0072] Choosing a suitable subset of the training dataset is crucial for the success of the proposed training method. In one embodiment, stratified sampling is used as a useful technique for selecting this subset. This involves dividing the training dataset into layers according to class labels, ensuring that samples are drawn proportionally from each layer to maintain the class distribution. Specifically, this involves dividing the entire dataset D into k (for labeled datasets, k can be the number of labels) non-overlapping layers (D1, D2, ..., D...). k ), making Then, each layer D identified by label i is independently processed using the following different methods. i Sampling is performed to select subsets, and each method aims to improve the training efficiency of ResNet / DenseNet models.

[0073] Random sampling: A simple method in which a random subset of the training data is selected. This method ensures the diversity and representativeness of the entire dataset. The size of the subset can be defined as:

[0074] |S i |=α·|D i | (1)

[0075] Where D i S represents the entire training dataset with label i. i From D i The selected subset, α, is a percentage metric that still needs adjustment. Random sampling is beneficial for quickly obtaining diverse sample sets, which can help the model learn the general characteristics of the data.

[0076] 2. QR decomposition for data sampling: QR decomposition is a matrix decomposition.

[0077] X = QR, (2)

[0078] Where Q is an orthogonal matrix (i.e., Q... T Q = I, where I is the identity matrix, and E is an upper triangular matrix. Orthogonal Q ensures that each subsequent column vector represents a new direction unrelated to the previous one. Therefore, any noisy or less important components are naturally separated from the main signal, focusing on more information dimensions. Thus, the upper triangular R encapsulates how much of the original data X can be represented along orthogonal Q. The diagonal lines of R aim to indicate directions with higher variance and lower importance. Therefore, the diagonal entries of R can generally reflect the importance of each principal component or direction. Geometrically, Q can be viewed as a rotation or reflection of the dataset in multidimensional space. By extracting the essential features of the prominent data structure, the transformation helps to adapt the problem to a more solvable format. This is the basis for QR decomposition to select important columns.

[0079] QR decomposition with column pivoting utilizes a permutation matrix P, which indicates the reordering of the columns of X:

[0080] XP = QR. (3)

[0081] P is usually chosen such that the diagonal elements of E are non-increasing: r 11 ≥r 22 ≥…,≥r nnTherefore, this order helps to prioritize the columns of X according to their importance or "informativeness". To improve representation and computational efficiency, a more compact permutation array p can be used instead of the permutation matrix P to achieve the same effect of rearranging the columns. Details on how the permutation array p corresponds to the permutation matrix P are described below. The permutation matrix P is a square binary matrix with exactly one entry "1" in each row and column, and all other entries "0", for example:

[0082]

[0083] For P1, since the "1" in the first column is in the second row, the "1" in the second column is in the third row, and the "1" in the third column is in the first row, the corresponding permutation array / vector is defined as p1 = [2 3 1]. Similarly, for P2, since the "1" in the first column is in the third row, the "1" in the second column is in the first row, and the "1" in the third column is in the second row, the corresponding p2 = [3 1 2].

[0084] Therefore, the ordered training data XP in (3) can be replaced with:

[0085] X :,p =QR (4)

[0086] Among them, X :,p The first column X :,p(1) The last column, X, contains the most information. :,p(|D|) This is the column with the least amount of information.

[0087] Therefore, in order to extract the D representing label i i The α·|D that yields the most information i |Sample, server 100 should follow these steps: 1) D i All sample data are directly converted into vectors, and then organized to form matrix X. i 2) Apply the QR decomposition method with column pivoting to rearrange X. i The columns thus produce an ordered matrix based on equation (4). Then, α·|D, which contains the most information i The sample subset is:

[0088]

[0089] in, This represents the j-th sample with the most information. Alternatively, to obtain the minimum "value" or information about the underlying data structure, where the focus is on maximizing the relevance of the data, samples are selected as follows:

[0090]

[0091] It is conceivable that other features could be included (attached) to the columns of X to make the data more discriminative and to make the selection of samples for pre-training more effective. The feature extraction method will depend on the application.

[0092] In one embodiment, Sobel edge detection and Gabor filtering are used to augment the data matrix X, resulting in better principal component sorting based on QR decomposition. These affect the permutation matrix P, but there is no direct relationship between P and the Sobel edge detection and Gabor filtering.

[0093] During QR decomposition, samples are processed in a vectorized format, thus converting the original training data into column features. Standard methods involve directly vectorizing the data to form columns, thereby X... i This is how it is derived. Alternatively, Sobel edge detection, Gabor filtering, and other methods can be used to derive features, which are then vectorized. The vectorized data is then organized into a matrix Y. i or Z i Sort using QR decomposition with column pivoting:

[0094]

[0095] Among them, the obtained permutation vector and These are associated with Sobel edge detection and Gabor filters, respectively. Therefore, a selected subset of the training dataset based on Sobel edge detection (SE) and Gabor filters (G) can be any one or more of the following:

[0096]

[0097]

[0098] Another embodiment of this disclosure may include exploring more feature extraction techniques. However, since only permutation vectors are available... and Instead of Y in (7) i or Z i As a valid index in (8), the permutation vector obtained by the alternative feature extraction method is similarly used as X in (8). i Valid indexes.

[0099] B. Training Phase

[0100] The training process consists of two main phases: initial training on a selected subset and full training on the complete dataset. This structure improves efficiency while ensuring high model accuracy.

[0101] Initial training of the subset: The initial phase involves training the ResNet / DenseNet model using a selected subset S of training data. Random sampling, such as in (1), can be used. ), data sampling with large information content ((5) and (8) and or data sampling with less information ((6)) and (8) and The subset is derived using methods such as [methods described above]. The training objective can be defined as minimizing the loss function on the subset.

[0102]

[0103] Where f(x; θ) represents the model's prediction, and y represents the true label. Let θ represent the loss function being optimized, and let θ represent the model parameters. Key hyperparameters for this stage may include the subset proportion α, the number of epochs E, and the learning rate L from (1). r Training continues until preset criteria are met, such as a convergence threshold or a maximum number of epochs.

[0104] In one embodiment of this disclosure, when the selected subset S is significantly small, the entire selected subset S is loaded into local memory for initial training.

[0105] 2. Full Training on the Complete Dataset: After initial training, the model undergoes full training using the complete training dataset D. This stage refines the parameters learned from the initial training, allowing for better generalization. The loss function remains unchanged, but the model is now evaluated on the entire dataset:

[0106]

[0107] Importantly, this stage involves setting hyperparameters (number of epochs E and learning rate L). r Adjustments are made to ensure the model is optimally configured for the complete dataset.

[0108] Using this two-stage approach, several advantages are expected, including faster convergence and less training time. The initial focus on a smaller subset allows for rapid model iteration and adaptation, while subsequent full training on the complete dataset ensures robustness and accuracy. Furthermore, training on a subset allows the model more freedom to explore the parameter space, reducing the risk of getting trapped in local minima. Additionally, the subset selection strategy has the opportunity to improve training efficiency by allowing the model to learn from diverse or highly informative samples. Algorithm 1 presents the proposed method step-by-step. Figure 2The flowcharts of the subset pre-training method and the original ResNet / DenseNet of System 300 are shown, as follows. Figure 3 As shown.

[0109]

[0110] refer to Figure 2 and Figure 3 The AI ​​system 300 disclosed herein is suitable for performing the following steps:

[0111] Step 202: Begin at the point where system 300 is adapted to load or retrieve a dataset from database 310;

[0112] Step 204: System 300 is adapted to divide the dataset into layers 312 according to class labels, so that each layer retains the class distribution;

[0113] Step 206: For each layer D identified by label i i Independent sampling was performed to form a training subset 314;

[0114] Step 208: System 300 passes training subset 314 to AI model 316 for initial subset pre-training;

[0115] Step 210: System 300 passes the entire dataset to AI model 316 for full dataset training;

[0116] Step 214: For each epoch in step 208 or step 210, system 300 evaluates the AI ​​model by comparing the prediction result 318 with the model prediction f(x; θ) with the true label or target output 322 in model validation module 320. The module stops training when a stopping condition is met. The stopping condition can be that the loss function is less than a threshold or that a certain number of epochs has been reached. Learning algorithm module 324 calculates the loss function in (9) based on the training data. The AI ​​model 326 is updated by adjusting the model's parameters θ (such as the weights or biases of neuron nodes) to minimize the loss function during the training phase.

[0117] C. Hyperparameter Adjustment

[0118] Hyperparameter tuning within this training framework is crucial because it directly affects the model's ability to generalize from the training data.

[0119] During subset pre-training, the hyperparameter α needs to be experimentally validated to select an appropriate value. The parameter value should not be too small, as a small value will not effectively achieve the purpose of introducing subset pre-training. Conversely, the parameter value should not be too large, as the runtime per epoch is insufficient to reduce training time. The number of epochs E can also be preset; it should not be too small, nor too large, as a large number of epochs will cause the parameters to overfit to the subset, thus reducing its effectiveness in training the entire dataset. In one embodiment, the learning rate L can be allocated according to the rules of the OneCycleLR function. r Because it has a preset epoch number E.

[0120] In the second stage (for complete training data) of the embodiments of this disclosure, a test accuracy threshold can be used instead of a preset epoch number as the iteration stopping condition. This adjustment is made to preferentially reduce the running time when a specified test accuracy is reached. Once the threshold is reached, the iteration stops, and the current epoch number is recorded as the running time. Therefore, E′ here is merely an upper limit. Furthermore, the learning rate L... r It is configured with a step learning rate, consistent with the traditional training process used for time-reduced comparative evaluation.

[0121] D. Faster convergence

[0122] Reducing the amount of data processed can facilitate faster convergence. The primary focus is on hyperparameters related to a subset α. With fewer samples, the model can iterate faster and identify optimal parameter configurations or settings earlier during training. This is the core principle behind the subset pre-training approach introduced in this publication.

[0123] Appropriate adjustments to other hyperparameters can also achieve faster convergence speeds. For example, the learning rate L... r However, to emphasize the importance of subset pre-training, the impact of the learning rate is deliberately minimized. Therefore, these learning rates are adjusted only for subset pre-training and training on the entire dataset, based on basic requirements. The guidelines for setting these rates are detailed in the paragraphs above.

[0124] In summary, the structured training method disclosed herein can maximize efficiency and effectiveness, enabling high-quality AI model training without compromising resource expenditure. However, if the tuning of relevant variables / hyperparameters is not carefully considered, these factors may become limitations or obstacles to the seamless implementation of this methodology.

[0125] experiment

[0126] In one embodiment of this disclosure, a comprehensive experimental setup has been implemented to evaluate the performance of popular deep learning frameworks, including ResNet and DenseNet. These experiments were conducted using established datasets, including the CIFAR10 and CIFAR100 datasets, and executed on standard desktop PCs and GPU P100s, ensuring adequate hardware conditions. Deep learning AI models were trained using various versions of TensorFlow and PyTorch to evaluate the impact on accuracy and performance of the proposed model frameworks. Extensive performance evaluation methods were employed, with a focus on accuracy metrics and training duration. The results were analyzed using numerical metrics of performance differences between different model architectures to improve the clarity and understanding of the results.

[0127] A. Hyperparameter Selection

[0128] In one embodiment, system 300 is used to identify optimal hyperparameter values ​​for a well-known ResNet / DenseNet model, with particular attention to the score (subset size) of the training dataset, represented by variable α, and the number of epochs E in the initial subset pre-training. Specifically, the following section uses training a well-known ResNet on the CIFAR10 dataset as an example to explain how α and E are selected. To achieve the purposes of this disclosure, a hyperparameter grid search strategy for examining various configurations of the AI ​​model will be implemented. Candidate values ​​for α are selected from the set α = {0.1, 0.2, 0.3, 0.4, 0.5}, representing different proportions of training data to be used in subset selection, while the number of epochs E in the subset pre-training process will be obtained from the selection {10, 20, 30, 50}.

[0129] This method begins by selecting a subset for a pre-training phase on an initial subset, followed by a full-scale training phase on the entire dataset. Two key performance metrics are then monitored at the end of the full-scale training phase: training time (epochs αE+E′ throughout the training process) and test accuracy. This systematic approach allows us to characterize efficient hyperparameter settings that best perform on the ResNet model on the CIFAR10 dataset, enhancing our understanding of how changing these hyperparameters during the initial pre-training phase affects the performance of the model disclosed herein. Results for α and E are shown in Tables 1, 2, and 3, respectively. According to the results in the tables, α = 0.2 / E = 20 and α = 0.3 / E = 10 perform well for training the ResNet model on the CIFAR10 dataset.

[0130] Table 1: Initial Training Configuration: The Influence of Random Subset Proportion α and Training Number E on Subset Training

[0131]

[0132] Table 2: Initial Training Configuration: The Impact of the Proportion of Subsets with High Information Content α and the Number of Training Sets E on Subset Training

[0133]

[0134]

[0135] Table 3: Initial Training Configuration: The Impact of the Proportion of Less Information-Containing Subsets α and the Number of Episodes E on Subset Pre-training

[0136]

[0137] B. ResNet pre-training on a subset of CIFAR10

[0138] One embodiment of this disclosure implements a rigorous evaluation of the proposed two-stage subset pre-training framework for image classification on the CIFAR10 benchmark dataset using ResNet 18. The experimental method begins with an initial training phase, in which a subset is selected using stratified sampling, followed by the identification of relevant training subsets included in step 2 of Algorithm 1 using random, informative, or less informative sampling methods.

[0139] For training data initialization, system 300 is configured with parameters batch_size = 128, shuffle = True, and num_workers = 2, and for test data initialization, the system has a configuration of batch_size = 100, shuffle = False, and num_workers = 2. System 300 in this embodiment focuses on several performance metrics to measure effectiveness, including the total training time (epochs αE+E′) of the combined two phases and the test accuracy achieved at the end of the full training phase. These multiple selections are inspired by the idea in the curriculum learning framework that "the definition of 'simple samples' is not yet clear." The results are then summarized in Table 4. Once the target test accuracy of 92% is achieved, the time savings for subset pre-training ResNet using various subset selection methods are typically between 50% and 55%. The subset with less information based on Sobel edge extraction features results in the most significant time reduction in Table IV, achieving a reduction of 50.2%. In contrast, other subset pre-training techniques with less information generally perform poorly. On the other hand, subset pre-training methods with high information content consistently achieved time reductions between 51% and 52%.

[0140] Table 4: Performance of ResNet18 pre-trained using subsets on CIFAR10

[0141]

[0142] C. DenseNet pre-training on a subset of CIFAR100

[0143] One embodiment of this disclosure is suitable for evaluating the effectiveness of a subset pre-training framework applied to the CIFAR100 dataset using a lighter model architecture, particularly DenseNet. The AI ​​model training in this embodiment of the disclosure will employ a two-stage approach, during which hyperparameters, including α and E, will be tuned based on insights gained from previous experiments. Compared to conventional training methods that do not employ subset selection, the performance metrics of this embodiment will include test accuracy and training efficiency. For ease of data presentation, Table 5 illustrates the test accuracy and training time results for various training settings, providing the model's performance and efficiency.

[0144] Using this method, as can be observed in Table 5, the subset pre-training framework significantly reduces the training time of DenseNet on the more complex CIFAR100 dataset by approximately 49% to 53% compared to traditional training techniques. The most significant time reduction is 49.4% achieved through the information-rich subset pre-training framework.

[0145] Table 5: Performance of DenseNet pre-trained using subsets on CIFAR100

[0146]

[0147] in conclusion

[0148] This disclosure introduces a novel subset pre-training method designed to significantly reduce the training time of deep neural networks (specifically ResNet / DenseNet) while maintaining model accuracy on the test dataset. Traditional methods for training neural networks typically require substantial computational resources and time, especially when dealing with large datasets. By implementing the two-stage training process of this disclosure—training with a smaller subset of training data before transitioning to training on the full dataset—the proposed method demonstrates its effectiveness in accelerating convergence and more efficiently achieving established accuracy benchmarks.

[0149] The results of this disclosure confirm the effectiveness of the subset pre-training method, demonstrating a significant reduction in training time while maintaining expected accuracy on unseen data. This analysis provides a comprehensive overview of the methodological framework, including the initial subset selection process, the transition to training on the full dataset, and evaluation metrics for assessing model performance.

[0150] In summary, the novel subset pre-training method and system disclosed herein not only help reduce training time but also maintain the integrity and performance of neural networks. Future work may explore the implementation of this method in various architectures and domains, potentially setting a new standard for computational efficiency in deep learning practice. By emphasizing the importance of efficient training procedures, we hope to pave the way for further innovation in this field.

[0151] Those skilled in the art will understand that various modifications and substitutions can be made to the above embodiments without departing from the overall scope of this disclosure. These variations may include, but are not limited to, changes to specific implementations, configurations, or methods, as long as they remain within the general principles and purposes of this disclosure. Therefore, the provided embodiments should be considered exemplary rather than restrictive, and any changes that do not alter the core functionality and purpose of this disclosure are considered to be within its scope.

Claims

1. A system for training an AI model on a training dataset, characterized in that, include: A processor is used to execute instructions; A computer-readable medium comprising instructions for execution on the processor, the instructions causing the processor to perform the following steps: The training dataset is divided into multiple layers based on the category labels; each layer represents a single label. Subset data is created by selecting samples from each of the layers; The AI ​​model is trained using the subset of data until a first stopping condition is met; and The AI ​​model is trained using the training dataset until the second stopping condition is met.

2. The system according to claim 1, characterized in that, in, The subset data is created by selecting a preset percentage α of samples from each layer.

3. The system according to claim 2, characterized in that, in, The subset data is created by selecting random samples from each layer.

4. The system according to claim 2, characterized in that, in, The subset data is created by selecting samples based on the information content of samples in each layer.

5. The system according to claim 4, characterized in that, in, The subset data is created by selecting samples based on ascending information content.

6. The system according to claim 4, characterized in that, in, The subset data is created by selecting samples based on descending information content.

7. The system according to claim 4, characterized in that, in, The following steps are used to determine the information content ranking of samples in each layer: The samples in each layer are converted into vectors, and then organized into matrix X. i ; Perform a QR decomposition method with column pivoting to rearrange X i Columns to obtain the matrix Make XP = QR or X :,p =QR, where P represents a square binary permutation matrix, Q represents an orthogonal matrix, R represents an upper triangular matrix, and X... :,p The first column X :,p(1) The last column, X, contains the most information. :,p(|D|) This is the column with the least amount of information.

8. The system according to claim 7, characterized in that, in, Choose P such that the diagonal elements of R are non-increasing: r 11 ≥r 22 ≥…,≥r nn .

9. The system according to claim 7, characterized in that, in, The data matrix X includes Sobel edge detection results to enhance input features.

10. The system according to claim 7, characterized in that, in, The data matrix X includes Gabor filtering results to enhance input features.

11. The system according to claim 1, characterized in that, in, The first stopping condition is met when the AI ​​model has reached the maximum convergence epoch.

12. The system according to claim 1, characterized in that, in, When the AI ​​model has reached the convergence threshold, the second stopping condition is met.

13. The system according to claim 1, characterized in that, in, The AI ​​model is a deep learning artificial neural network model.

14. The system according to claim 1, characterized in that, in, The steps of training the AI ​​model include adjusting the loss function on the subset. The minimized steps make Where f(x; θ) represents the model's prediction, and y represents the true label. Let θ represent the loss function being optimized, and let θ represent the model parameters.

15. The system according to claim 1, characterized in that, in, The steps for training the AI ​​model include: selecting samples of the subset data based on the QR decomposition with column principal components, and then using the entire complete dataset as the training dataset.

16. The system according to claim 1, characterized in that, in, The subset data is adaptively and progressively selected from 20% to 100% of the size of the entire complete dataset.