A method for training an AI large model based on deep learning

CN122389943BActive Publication Date: 2026-09-11BEIJING LIYANG ZHIGUANG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610680791.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-18
Publication Date
2026-09-11
Estimated Expiration
2046-05-18

AI Technical Summary

Technical Problem

[0003]然而,上述训练加速与调度策略多为静态配置或仅依赖单一损失指标,缺乏对损失变化特征、梯度变化特征与注意力分布特征的建模,未对训练稳定性进行刻画,也未利用频域震荡特征识别高频不稳定模式

Benefits of technology

本发明通过在多阶段稳态训练调度机制中联合考虑损失变化特征、梯度变化特征与注意力分布特征,并引入基于频域震荡特征修正的训练稳定性指标与稳态评分序列,使训练状态由单一尺度的“是否收敛”提升为对时域趋势与频域震荡同时刻画的综合度量,在此基础上将稳态评分与注意力加速机制紧密耦合,实现对注意力稀疏度配置、标记合并配置和数值精度配置的动态联动控制,避免了传统静态稀疏或固定剪裁导致的训练发散或精度损失。通过在阶段化的训练模型结构生成过程中引入分层目标注意力稀疏度配置以及预热阶段、过渡阶段、稳态阶段的阶段切换与回退逻辑,本发明能够在训练早期保持较高注意力保留比例与较高数值精度,在训练进入稳态后逐步提升稀疏度与标记合并比例,并在检测到训练不稳定时自动回退至保守配置,从而在保证大规模深度学习模型收敛稳定性的前提下显著提升训练效率、缩短有效收敛时间,并降低由于注意力加速策略不当带来的性能退化风险。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122389943B_ABST
    Figure CN122389943B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on deep learning AI big model training method of model, including the following steps: obtaining training data and model to be trained, construct attention acceleration mechanism and multi-stage steady training scheduling mechanism, execute first forward propagation;Stability monitoring is executed, and steady score sequence is obtained by frequency domain analysis;The attention acceleration configuration corresponding to current training stage is generated, and attention structure is updated;Forward propagation and back propagation are executed, and stage training result is obtained;Steady score sequence is recalculated to complete stage switching determination, and the scheduling instruction of next training cycle is generated, and stage training configuration is updated;Training cycle is repeatedly executed, and training completion model is obtained to meet termination condition.The application utilizes frequency domain steady score and layered attention acceleration training mechanism, constructs dynamic and can back multi-stage steady scheduling method, realizes that big model training is more stable, faster, more efficient.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of large model training, and more particularly to a method for training large AI models based on deep learning. Background Technology

[0002] With the widespread application of deep learning, especially large AI models based on attention structures, in natural language processing, image generation, and other scenarios, the industry typically improves expressive power by increasing the number of model layers and parameter size. During training, techniques such as learning rate warm-up and piecewise decay, mixed precision training, sparse attention, or label truncation are used to reduce computational and memory overhead. Existing technologies also involve simple monitoring of metrics such as loss curves and gradient norms, using human experience to determine whether training has converged or gradient explosion has occurred, thereby assisting in selecting training epochs, saving checkpoints, or adjusting optimizer parameters.

[0003] However, the aforementioned training acceleration and scheduling strategies are mostly statically configured or rely solely on a single loss metric, lacking modeling of loss variation characteristics, gradient variation characteristics, and attention distribution characteristics. They fail to characterize training stability or utilize frequency domain oscillation characteristics to identify high-frequency unstable patterns. Furthermore, existing attention acceleration methods typically set sparsity and label merging ratios globally without considering the training phase and the sensitivity of the attention layer for hierarchical attention sparsity configuration. They also lack closed-loop control linked to phase switching and rollback operations, making it difficult to optimize training efficiency and model accuracy while ensuring convergence stability.

[0004] Therefore, how to provide a method for training large AI models based on deep learning is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] One objective of this invention is to propose a method for training large AI models based on deep learning. This invention utilizes frequency domain steady-state scoring and hierarchical attention to accelerate training mechanisms and constructs a dynamic and regressible multi-stage steady-state scheduling method to achieve more stable, faster, and more efficient training of large models.

[0006] A method for training a large AI model based on deep learning according to an embodiment of the present invention includes the following steps: Acquire training data and a large-scale deep learning model to be trained, construct an attention acceleration mechanism and a multi-stage steady-state training scheduling mechanism, input the training data into the large-scale deep learning model and perform the first forward propagation to obtain the initial training state information; Based on the initial training state information, stability monitoring is performed using a multi-stage steady-state training scheduling mechanism to obtain loss change characteristics, gradient change characteristics, and attention distribution characteristics, generating training stability indicators, and calculating steady-state score sequences through frequency domain analysis. The steady-state score sequence is input into the attention acceleration mechanism to generate the attention acceleration configuration corresponding to the current training stage. The attention structure of the large-scale deep learning model is updated based on the attention acceleration configuration to obtain the staged training model structure. Based on the phased training model structure, forward and backward propagation are performed to obtain the phased training results, which are then fed back to the multi-phase steady-state training scheduling mechanism. Based on the stage training results, the steady-state score sequence is recalculated, the stage switching determination is completed, the scheduling instruction for the next training cycle is generated, and when training instability is detected, a rollback operation is performed to update the attention acceleration configuration, resulting in the updated stage training configuration. The updated stage training configuration is input into the large-scale deep learning model and the training cycle is repeated until the termination condition is met, finally resulting in a trained large-scale deep learning model.

[0007] Optionally, the generation of the initial training state information specifically includes: Acquire large-scale, multi-source training data for training, distinguish them according to data type, unify the format of the original training data, handle missing values ​​and remove outlier samples to obtain a structured training sample set. Deep learning preprocessing operations are performed on the structured training sample set, and the training sample set is divided according to preset rules to obtain the batch training data set; Construct the network structure of the large-scale deep learning model to be trained, initialize the parameter set of the large-scale deep learning model, associate the parameter set with the network structure, and obtain the initialized large-scale deep learning model. Based on the initialization of the large-scale deep learning model, the internal configuration items of the attention acceleration mechanism and the multi-stage steady-state training scheduling mechanism are predefined and set to the initial state to obtain the initial mechanism configuration information. The first batch of training data is selected from the batch training dataset, associated with the initialized large-scale deep learning model and the initial mechanism configuration information, and the first forward propagation is performed to obtain the output representation and initial loss value of the large-scale deep learning model on the first batch of training data. During the first forward propagation, the attention weight distribution corresponding to each attention layer of the large-scale deep learning model is calculated to obtain the attention distribution matrix set corresponding to the first batch of training data. The initial training state information is constructed based on the initial loss value and the attention distribution matrix set.

[0008] Optionally, the generation of the steady-state score sequence specifically includes: Read the initial loss value and initial attention distribution matrix set from the initial training state information. In the multi-stage steady-state training scheduling mechanism, establish a loss sequence cache and attention distribution sequence cache corresponding to the training rounds, and write the initial loss value and initial attention distribution matrix set into them respectively to obtain the basic training state sequence. Establish a training stability index value sequence cache that stores the training stability index values ​​of each training cycle. After completing several subsequent training cycles, the loss value of the current training cycle and the loss value of the previous training cycle are extracted from the basic training state sequence, and the difference is calculated as the loss change of the current training cycle. Using an exponentially weighted moving average, the loss value of the current training cycle and the smoothed loss value of the previous training cycle are weighted and combined according to a preset smoothing coefficient to obtain the smoothed loss value of the current training cycle. This smoothed loss value is then combined with the smoothed loss value to form a loss change feature, which is written into the loss change feature cache. During the backpropagation process of the current training cycle, the gradients of the large-scale deep learning model on the parameters of each layer are obtained and combined into the overall gradient representation of the current training cycle. The L2 norm of the overall gradient representation of the current training cycle is calculated as the gradient norm of the current training cycle and compared with the gradient norm of the previous training cycle to obtain the gradient change magnitude. This is combined with the gradient change magnitude of the current training cycle to form the gradient change feature and written into the gradient change feature cache. Read the set of attention distribution matrices corresponding to the current training period from the attention distribution sequence cache, calculate the information entropy corresponding to the attention distribution probability at each position, and calculate the attention sparsity of each attention layer based on the non-zero proportion of the attention distribution probability at each position. Normalize and aggregate to obtain the attention distribution features of the current training period and add them to the attention distribution feature cache. The loss change features, gradient change features, and attention distribution features of the current training cycle are read from the loss change feature cache, gradient change feature cache, and attention distribution feature cache, respectively. After normalization, a linear weighted sum is performed based on the preset weighting coefficients to obtain the initial training stability index value of the current training cycle. The initial training stability index value of the current training cycle and the initial training stability index values ​​corresponding to several historical training cycles are written together into the training stability index value sequence cache. Frequency domain oscillation analysis is performed to extract the frequency domain oscillation features that reflect the degree of high-frequency oscillation. The energy ratio of high-frequency components is calculated and a correction coefficient is generated and applied to the initial training stability index value of the current training cycle to obtain the training stability index value of the current training cycle. The difference between the training stability index value of the current training cycle and the preset upper bound value is used as the steady-state score of the current training cycle, thus obtaining the steady-state score sequence.

[0009] Optionally, the generation of the staged training model structure specifically includes: Read the steady-state score corresponding to the current training cycle from the steady-state score sequence and compare it with the pre-set thresholds for the warm-up stage, the transition stage, and the steady-state stage in the multi-stage steady-state training scheduling mechanism to determine the training stage label to which the current training cycle belongs and obtain the current training stage status information. Based on the current training phase state information, the steady-state score is normalized, and the target attention sparsity corresponding to the current training cycle is calculated according to the normalized steady-state score and attention sparsity configuration. Based on the preset hierarchical sensitivity parameters, hierarchical scaling is performed, and the hierarchical sensitivity parameters of each attention layer are adjusted to obtain the hierarchical target attention sparsity for each attention layer. Based on the linear mapping relationship between the normalized steady-state score and the minimum and maximum label merging ratio thresholds corresponding to the label merging configuration in the current training phase, calculate the target label merging ratio for the current training period. Based on the threshold division relationship between the normalized steady-state score and the numerical accuracy configuration in the current training phase, the target numerical accuracy level corresponding to the current training cycle is determined. The hierarchical target attention sparsity, target label merging ratio and target numerical accuracy level are combined to form the attention acceleration configuration corresponding to the current training cycle, and the updated attention acceleration configuration is obtained. The updated attention acceleration configuration is applied to each attention layer of the large-scale deep learning model. The attention connection retention mask, redundant label merging rules and attention calculation numerical precision mode of each attention layer are uniformly updated. The updated large-scale deep learning model is output as a staged training model structure.

[0010] Optionally, the generation of the stage training results specifically includes: Select the batch training data corresponding to the current training period from the batch training data set, input the staged training model structure and perform forward propagation to obtain the output representation of the current batch training data and the set of loss values ​​and attention distribution matrices corresponding to the current training period, and combine them to form the forward propagation result of the current training period; The supervision labels and training objectives corresponding to the current batch of training data are used together with the output representation to construct the loss function input for the current training cycle. Backpropagation is performed based on the staged training model structure. The gradients of the trainable parameters in each attention layer, feedforward transformation layer and normalization layer are calculated in sequence. The gradients are aggregated according to the network structure to form the overall gradient representation of the current training cycle and obtain the gradient norm of the current training cycle. Based on the overall gradient representation and gradient norm of the current training cycle, gradient updates are performed on each trainable parameter in the staged training model structure to obtain the set of parameters of the large-scale deep learning model after the current training cycle. The loss values, overall gradient representation, gradient norm and attention distribution matrix before and after the update are organized to form stage training statistics. The stage training statistics are written into the loss sequence cache and attention distribution sequence cache in the multi-stage steady-state training scheduling mechanism. The overall gradient representation and gradient norm corresponding to the current training cycle are written into the intermediate data cache used for training stability index calculation in the multi-stage steady-state training scheduling mechanism to obtain the stage training results used for subsequent stability monitoring.

[0011] Optionally, the generation of the updated stage training configuration specifically includes: The training results of each stage are read from the multi-stage steady-state training scheduling mechanism. The loss change characteristics, gradient change characteristics and attention distribution characteristics, including the current training cycle and several historical training cycles, are updated and calculated to obtain the training stability index value and steady-state score corresponding to the current training cycle, forming an updated steady-state score sequence. The steady-state score corresponding to the current training cycle is compared with the pre-set thresholds for the warm-up stage, transition stage, and steady-state stage in the multi-stage steady-state training scheduling mechanism to determine the target training stage label for the next training cycle and obtain the stage switching judgment result. Based on the stage switching determination result, the attention sparsity configuration boundary, label merging configuration boundary and numerical accuracy configuration level set corresponding to the target training stage label are found in the multi-stage steady-state training scheduling mechanism. The steady-state score and the target training stage label are written into the scheduling control logic to generate a scheduling instruction for the next training cycle. Based on the training stability index value and stage training statistics of the current training cycle, training instability detection is performed in the multi-stage steady-state training scheduling mechanism. The training stability index value is compared with the preset instability threshold, and the loss change and gradient norm change magnitude of several adjacent training cycles are statistically analyzed. When the training stability index value exceeds the instability threshold, the current training cycle is marked as a training unstable state, and a training instability label is generated. When the training instability flag is true, a rollback operation is performed based on the current training stage state information and the target training stage label to update the attention acceleration configuration. The target training stage label is rolled back from the steady state stage to the transition stage, or from the transition stage to the warm-up stage. In the corresponding training stage after rollback, the attention acceleration configuration is readjusted according to the preset rollback ratio to obtain the updated stage training configuration. When the training instability is marked as false, the target training stage label and attention acceleration configuration corresponding to the stage switching judgment result are output as the updated stage training configuration.

[0012] Optionally, the generation of the trained large-scale deep learning model specifically includes: Read the updated stage training configuration, associate it with the parameter set of the large-scale deep learning model in the current training cycle, obtain the stage training initialization information for the next training cycle, and write the target training stage label and the corresponding attention sparsity configuration, label merging configuration and numerical precision configuration into the stage training initialization information. Based on the initialization information of the stage training, the attention sparsity configuration, label merging configuration and numerical precision configuration in the attention acceleration mechanism are loaded synchronously and applied to each attention layer of the large-scale deep learning model. They are then associated with the next batch of training data to complete forward propagation and backward propagation and obtain the stage training results. The training results of each stage are written into the loss sequence cache, attention distribution sequence cache and intermediate data cache in the multi-stage steady-state training scheduling mechanism, the training stability index value and steady-state score sequence are updated, and stability evolution information for termination condition determination is generated. Based on stability evolution information, the training stability index value and validation loss value of the current and historical training cycles are read from the multi-stage steady-state training scheduling mechanism, their change range is calculated, and the termination condition is compared with the maximum training round, the minimum training stability change threshold and the minimum validation loss change threshold to obtain the termination condition judgment result. When the termination condition is not met, the updated stage training configuration is written into the multi-stage steady-state training scheduling mechanism, the training round count is updated, and the staged training model structure and stage training initialization information are retained to form a cross-cycle training closed loop. When the termination condition is met, the parameter set of the large-scale deep learning model in the current training cycle is frozen, and the staged training model structure is regarded as the large-scale deep learning model that has been trained.

[0013] The beneficial effects of this invention are: This invention, by jointly considering loss variation characteristics, gradient variation characteristics, and attention distribution characteristics in a multi-stage steady-state training scheduling mechanism, and introducing a training stability index and steady-state scoring sequence based on frequency domain oscillation characteristics correction, elevates the training state from a single-scale "whether it has converged" to a comprehensive measure that simultaneously characterizes time-domain trends and frequency-domain oscillations. Based on this, the steady-state scoring is tightly coupled with the attention acceleration mechanism to achieve dynamic linkage control of attention sparsity configuration, label merging configuration, and numerical accuracy configuration, avoiding training divergence or accuracy loss caused by traditional static sparsity or fixed pruning. By introducing hierarchical target attention sparsity configuration and stage switching and rollback logic for the warm-up, transition, and steady-state stages during the staged training model structure generation process, this invention can maintain a high attention retention ratio and high numerical accuracy in the early stages of training. After training enters a steady state, it gradually increases the sparsity and label merging ratio, and automatically rolls back to a conservative configuration when training instability is detected. This significantly improves training efficiency, shortens effective convergence time, and reduces the risk of performance degradation due to inappropriate attention acceleration strategies while ensuring the convergence stability of large-scale deep learning models. Attached Figure Description

[0014] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of a method for training a large AI model based on deep learning, as proposed in this invention. Figure 2 This is a schematic diagram of the multi-stage steady-state training scheduling mechanism of a deep learning-based AI large model training method proposed in this invention. Detailed Implementation

[0015] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0016] refer to Figures 1-2 A method for training large AI models based on deep learning includes the following steps: Acquire training data and a large-scale deep learning model to be trained, construct an attention acceleration mechanism and a multi-stage steady-state training scheduling mechanism, input the training data into the large-scale deep learning model and perform the first forward propagation to obtain the initial training state information; Based on the initial training state information, stability monitoring is performed using a multi-stage steady-state training scheduling mechanism to obtain loss change characteristics, gradient change characteristics, and attention distribution characteristics, generating training stability indicators, and calculating steady-state score sequences through frequency domain analysis. The steady-state score sequence is input into the attention acceleration mechanism to generate the attention acceleration configuration corresponding to the current training stage. The attention structure of the large-scale deep learning model is updated based on the attention acceleration configuration to obtain the staged training model structure. Based on the phased training model structure, forward and backward propagation are performed to obtain the phased training results, which are then fed back to the multi-phase steady-state training scheduling mechanism. Based on the stage training results, the steady-state score sequence is recalculated, the stage switching determination is completed, the scheduling instruction for the next training cycle is generated, and when training instability is detected, a rollback operation is performed to update the attention acceleration configuration, resulting in the updated stage training configuration. The updated stage training configuration is input into the large-scale deep learning model and the training cycle is repeated until the termination condition is met, finally resulting in a trained large-scale deep learning model.

[0017] In this embodiment, the generation of the initial training state information specifically includes: Acquire large-scale, multi-source training data for training, distinguish them according to data type, unify the format of the original training data, handle missing values ​​and remove outlier samples to obtain a structured training sample set. Deep learning preprocessing operations are performed on the structured training sample set, and the training sample set is divided according to preset rules to obtain the batch training data set; The deep learning preprocessing operations include word segmentation and tag encoding of text samples, size normalization and pixel normalization of image samples, and feature normalization of other samples. Construct the network structure of the large-scale deep learning model to be trained, initialize the parameter set of the large-scale deep learning model, associate the parameter set with the network structure, and obtain the initialized large-scale deep learning model. The network structure includes multiple attention layers, feedforward transformation layers, and normalization layers; Based on the initialization of the large-scale deep learning model, the internal configuration items of the attention acceleration mechanism and the multi-stage steady-state training scheduling mechanism are predefined and set to the initial state to obtain the initial mechanism configuration information. The attention acceleration mechanism consists of an attention sparsity configuration that controls the proportion of attention connection retention, a label merging configuration that controls the proportion of redundant label merging, and a numerical precision configuration that controls the numerical precision level of attention calculation. The multi-stage steady-state training scheduling mechanism consists of loss change characteristics that monitor loss changes, gradient change characteristics that monitor gradient changes, attention distribution characteristics that monitor attention distribution changes, and scheduling control logic that generates training stability indicators and stage switching thresholds. The first batch of training data is selected from the batch training dataset, associated with the initialized large-scale deep learning model and the initial mechanism configuration information, and the first forward propagation is performed to obtain the output representation and initial loss value of the large-scale deep learning model on the first batch of training data. During the first forward propagation, the corresponding attention weight distribution of each attention layer of the large-scale deep learning model is calculated to obtain the set of attention distribution matrices corresponding to the first batch of training data. The initial training state information is constructed based on the initial loss value and the set of attention distribution matrices. In this process, each attention layer calculates the similarity between the query representation matrix and the key representation matrix of the first batch of training data in the attention layer, scales them according to the constant corresponding to the feature dimension, and then performs row-wise normalization to obtain the attention distribution probability corresponding to each position, thus forming the attention distribution matrix.

[0018] In this embodiment, the generation of the steady-state scoring sequence specifically includes: Read the initial loss value and initial attention distribution matrix set from the initial training state information. In the multi-stage steady-state training scheduling mechanism, establish a loss sequence cache and attention distribution sequence cache corresponding to the training rounds, and write the initial loss value and initial attention distribution matrix set into them respectively to obtain the basic training state sequence. Establish a training stability index value sequence cache that stores the training stability index values ​​of each training cycle. After completing several subsequent training cycles, the loss value of the current training cycle and the loss value of the previous training cycle are extracted from the basic training state sequence, and the difference is calculated as the loss change of the current training cycle. Using an exponentially weighted moving average, the loss value of the current training cycle and the smoothed loss value of the previous training cycle are weighted and combined according to a preset smoothing coefficient to obtain the smoothed loss value of the current training cycle. This smoothed loss value is then combined with the smoothed loss value to form a loss change feature, which is written into the loss change feature cache. During the backpropagation process of the current training cycle, the gradients of the large-scale deep learning model on the parameters of each layer are obtained and combined into the overall gradient representation of the current training cycle. The L2 norm of the overall gradient representation of the current training cycle is calculated as the gradient norm of the current training cycle and compared with the gradient norm of the previous training cycle to obtain the gradient change magnitude. This is combined with the gradient change magnitude of the current training cycle to form the gradient change feature and written into the gradient change feature cache. Read the set of attention distribution matrices corresponding to the current training period from the attention distribution sequence cache, calculate the information entropy corresponding to the attention distribution probability at each position, and calculate the attention sparsity of each attention layer based on the non-zero proportion of the attention distribution probability at each position. Normalize and aggregate to obtain the attention distribution features of the current training period and add them to the attention distribution feature cache. The loss change features, gradient change features, and attention distribution features of the current training cycle are read from the loss change feature cache, gradient change feature cache, and attention distribution feature cache, respectively. After normalization, a linear weighted sum is performed based on the preset weighting coefficients to obtain the initial training stability index value of the current training cycle. The initial training stability index value of the current training cycle and the initial training stability index values ​​corresponding to several historical training cycles are written together into the training stability index value sequence cache. Frequency domain oscillation analysis is performed to extract the frequency domain oscillation features that reflect the degree of high-frequency oscillation. The energy ratio of high-frequency components is calculated and a correction coefficient is generated and applied to the initial training stability index value of the current training cycle to obtain the training stability index value of the current training cycle. The difference between the training stability index value of the current training cycle and the preset upper bound value is used as the steady-state score of the current training cycle, thus obtaining the steady-state score sequence.

[0019] In this embodiment, the generation of the staged training model structure specifically includes: Read the steady-state score corresponding to the current training cycle from the steady-state score sequence and compare it with the pre-set thresholds for the warm-up stage, the transition stage, and the steady-state stage in the multi-stage steady-state training scheduling mechanism to determine the training stage label to which the current training cycle belongs and obtain the current training stage status information. The training phase labels include warm-up phase, transition phase, and steady-state phase; Based on the current training phase state information, the steady-state score is normalized, and the target attention sparsity corresponding to the current training cycle is calculated according to the normalized steady-state score and attention sparsity configuration. Based on the preset hierarchical sensitivity parameters, hierarchical scaling is performed, and the hierarchical sensitivity parameters of each attention layer are adjusted to obtain the hierarchical target attention sparsity for each attention layer. The target attention sparsity is mapped to a specific attention connection retention ratio by linearly interpolating between the minimum attention sparsity threshold and the maximum attention sparsity threshold in the current training phase. Based on the linear mapping relationship between the normalized steady-state score and the minimum and maximum label merging ratio thresholds corresponding to the label merging configuration in the current training phase, calculate the target label merging ratio for the current training period. Based on the threshold division relationship between the normalized steady-state score and the numerical accuracy configuration in the current training phase, the target numerical accuracy level corresponding to the current training cycle is determined. The hierarchical target attention sparsity, target label merging ratio and target numerical accuracy level are combined to form the attention acceleration configuration corresponding to the current training cycle, and the updated attention acceleration configuration is obtained. The updated attention acceleration configuration is applied to each attention layer of the large-scale deep learning model. The attention connection retention mask, redundant label merging rules and attention calculation numerical precision mode of each attention layer are uniformly updated. The updated large-scale deep learning model is output as a staged training model structure.

[0020] In this embodiment, the generation of the stage training results specifically includes: Select the batch training data corresponding to the current training period from the batch training data set, input the staged training model structure and perform forward propagation to obtain the output representation of the current batch training data and the set of loss values ​​and attention distribution matrices corresponding to the current training period, and combine them to form the forward propagation result of the current training period; The supervision labels and training objectives corresponding to the current batch of training data are used together with the output representation to construct the loss function input for the current training cycle. Backpropagation is performed based on the staged training model structure. The gradients of the trainable parameters in each attention layer, feedforward transformation layer and normalization layer are calculated in sequence. The gradients are aggregated according to the network structure to form the overall gradient representation of the current training cycle and obtain the gradient norm of the current training cycle. Based on the overall gradient representation and gradient norm of the current training cycle, gradient updates are performed on each trainable parameter in the staged training model structure to obtain the set of parameters of the large-scale deep learning model after the current training cycle. The loss values, overall gradient representation, gradient norm and attention distribution matrix before and after the update are organized to form stage training statistics. The stage training statistics are written into the loss sequence cache and attention distribution sequence cache in the multi-stage steady-state training scheduling mechanism. The overall gradient representation and gradient norm corresponding to the current training cycle are written into the intermediate data cache used for training stability index calculation in the multi-stage steady-state training scheduling mechanism to obtain the stage training results used for subsequent stability monitoring.

[0021] In this embodiment, the generation of the updated stage training configuration specifically includes: The training results of each stage are read from the multi-stage steady-state training scheduling mechanism. The loss change characteristics, gradient change characteristics and attention distribution characteristics, including the current training cycle and several historical training cycles, are updated and calculated to obtain the training stability index value and steady-state score corresponding to the current training cycle, forming an updated steady-state score sequence. The steady-state score corresponding to the current training cycle is compared with the pre-set thresholds for the warm-up stage, transition stage, and steady-state stage in the multi-stage steady-state training scheduling mechanism to determine the target training stage label for the next training cycle and obtain the stage switching judgment result. Based on the stage switching determination result, the attention sparsity configuration boundary, label merging configuration boundary and numerical accuracy configuration level set corresponding to the target training stage label are found in the multi-stage steady-state training scheduling mechanism. The steady-state score and the target training stage label are written into the scheduling control logic to generate a scheduling instruction for the next training cycle. The scheduling instruction is used to instruct the attention acceleration mechanism to update the attention sparsity configuration, label merging configuration, and numerical precision configuration according to the target training stage label in the next training cycle; Based on the training stability index value and stage training statistics of the current training cycle, training instability detection is performed in the multi-stage steady-state training scheduling mechanism. The training stability index value is compared with the preset instability threshold, and the loss change and gradient norm change magnitude of several adjacent training cycles are statistically analyzed. When the training stability index value exceeds the instability threshold, the current training cycle is marked as a training unstable state, and a training instability label is generated. Among them, the training instability detection obtains the instability metric of the current training cycle by weighting the normalized loss change, the normalized gradient-related features, and the normalized attention distribution-related features, and compares the instability metric with the instability threshold to determine the training instability label. When the training instability flag is true, a rollback operation is performed based on the current training stage state information and the target training stage label to update the attention acceleration configuration. The target training stage label is rolled back from the steady state stage to the transition stage, or from the transition stage to the warm-up stage. In the corresponding training stage after rollback, the attention acceleration configuration is readjusted according to the preset rollback ratio to obtain the updated stage training configuration. When the training instability is marked as false, the target training stage label and attention acceleration configuration corresponding to the stage switching judgment result are output as the updated stage training configuration.

[0022] In this embodiment, the generation of the trained large-scale deep learning model specifically includes: Read the updated stage training configuration, associate it with the parameter set of the large-scale deep learning model in the current training cycle, obtain the stage training initialization information for the next training cycle, and write the target training stage label and the corresponding attention sparsity configuration, label merging configuration and numerical precision configuration into the stage training initialization information. Based on the initialization information of the stage training, the attention sparsity configuration, label merging configuration and numerical precision configuration in the attention acceleration mechanism are loaded synchronously and applied to each attention layer of the large-scale deep learning model. They are then associated with the next batch of training data to complete forward propagation and backward propagation and obtain the stage training results. The training results of each stage are written into the loss sequence cache, attention distribution sequence cache and intermediate data cache in the multi-stage steady-state training scheduling mechanism, the training stability index value and steady-state score sequence are updated, and stability evolution information for termination condition determination is generated. Based on stability evolution information, the training stability index value and validation loss value of the current and historical training cycles are read from the multi-stage steady-state training scheduling mechanism, their change range is calculated, and the termination condition is compared with the maximum training round, the minimum training stability change threshold and the minimum validation loss change threshold to obtain the termination condition judgment result. When the termination condition is not met, the updated stage training configuration is written into the multi-stage steady-state training scheduling mechanism, the training round count is updated, and the staged training model structure and stage training initialization information are retained to form a cross-cycle training closed loop. When the termination condition is met, the parameter set of the large-scale deep learning model in the current training cycle is frozen, and the staged training model structure is regarded as the large-scale deep learning model that has been trained.

[0023] Example 1: To verify the feasibility of this invention in practice, it was applied to a large-scale language model training task conducted by an artificial intelligence research institution in an eastern coastal province from March to June 2025. This institution consistently faced problems such as training instability, frequent oscillations, high resource consumption, and excessively long training cycles when jointly pre-training multi-source text, images, and structured data. These issues were particularly prevalent in the early stages of the model and during transitions between stages, leading to sudden increases in loss, abnormal gradient amplification, and extreme attention distribution, resulting in slow training progress and difficulty in convergence. The frequency domain steady-state scoring mechanism and hierarchical attention acceleration strategy proposed in this invention are precisely designed to address these pain points.

[0024] In this example, the research institution first inputs multi-source training data into a large-scale deep learning model to be trained after cleaning, alignment, and unified encoding. Following the steps of this invention, an attention acceleration mechanism and a multi-stage steady-state training scheduling mechanism are constructed. In the early stages of training, due to the complexity of data sources and the large range of features, the model loss curve exhibits significant fluctuations. This invention obtains the initial training state through the first round of forward propagation and writes the loss sequence and attention distribution sequence into the steady-state scheduling structure. Subsequently, in subsequent training cycles, the scheduling mechanism continuously tracks the loss change characteristics and gradient change characteristics, and then automatically extracts the degree of high-frequency oscillations by combining frequency domain oscillation analysis to identify unstable training periods. By correcting the training stability index in the frequency domain, this invention can accurately determine whether the current training cycle is in an excessive oscillation state, making the changes in the steady-state score more consistent with the actual training process.

[0025] In actual training, as the model gradually enters the transition phase, the hierarchical sparse configuration of this invention begins to play a role. The training team dynamically adjusts the target attention sparsity of each layer based on the sensitivity coefficients of different attention layers during gradient updates, enabling higher-level attention to converge faster and lower-level attention to maintain stable computation. This hierarchical adjustment not only improves the computational efficiency of attention but also avoids the information loss caused by past coarse-grained sparse strategies, laying the foundation for subsequent stable training. Simultaneously, the label merging strategy and dynamic numerical precision switching strategy of this invention are automatically activated at different training stages, ensuring further improvement in computational throughput during stable training phases and maintaining robustness during fluctuating training phases.

[0026] As training progresses across multiple phases, the phase switching logic of this invention begins to function frequently. The training team observed that phase switching determined by frequency domain steady-state scoring is more accurate than traditional loss trend-based methods, effectively avoiding loss oscillations caused by erroneous switching. During periods of training instability, the rollback operation of this invention takes effect promptly, reverting the training phase from steady state to a transition or warm-up phase, allowing attention configuration to converge back to a safe range, thereby avoiding gradient collapse, a problem that was prone to occur in the past.

[0027] Throughout the approximately three-month training period, the organization continuously monitored loss changes, gradient evolution, and frequency domain steady-state conditions using the method of this invention. Finally, after meeting the termination condition logic, the model parameters were frozen, resulting in a fully trained large-scale deep learning model. Training process records show that this method exhibits significant advantages in training resource utilization, training time, convergence stability, and stepwise optimization of the attention layer. The training team reported that the introduction of this invention makes large-scale training tasks more controllable and stable, with theoretical and practical results consistent, providing a sustainable architectural foundation for subsequent large-scale model training processes.

[0028] Table 1. Performance comparison between deep learning-based AI large model training methods and traditional methods.

[0029] As can be seen from the table, traditional training methods exhibit significant instability across several key metrics. For example, the loss oscillation amplitude is as high as 0.42, indicating a large fluctuation range during training. Furthermore, the gradient norm peak is close to 10, representing a highly unstable gradient during training, which can easily lead the model in the wrong direction in the optimization space. In contrast, the training framework of this invention, because the frequency domain steady-state score can actively identify high-frequency oscillations and dynamically correct the calculation path, significantly reduces the oscillation amplitude to 0.17, and the gradient peak also decreases to 5.1, falling within a significantly more stable gradient range.

[0030] The standard deviation of the gradient norm decreased from 2.31 to 0.94, demonstrating more continuous gradient updates without significant abrupt changes. This directly reflects the differentiated control applied to different attention layers by the hierarchical sparse configuration. When lower-level attention still needs to maintain high accuracy, this invention automatically allocates a lower sparsity ratio to avoid information loss; while increasing sparsity in higher-level attention layers ensures reduced computational costs. This hierarchical sensitivity-based differentiated strategy reduces the variation in attention sparsity from 0.36 to 0.12, resulting in a more balanced training structure.

[0031] In the frequency domain analysis results, the proportion of high-frequency oscillation energy decreased from 0.31 to 0.09, indicating that the present invention effectively suppressed the fluctuation noise in the later stage of training, making the switching of training phases more accurate. Therefore, in the entire training cycle, the number of misjudgments in phase switching decreased from 7 to 2, and the number of rollback triggers decreased from 9 to 3, demonstrating the scheduling mechanism's higher ability to identify "true instability" and "false instability".

[0032] In terms of computational efficiency, after the attention acceleration mechanism took effect, the single-batch latency improved from 213ms to 164ms, and the GPU utilization rate decreased from 92% to 78%, indicating that the computational pressure was substantially reduced and the efficiency of computing resource utilization was improved.

[0033] In terms of training performance, this invention achieves stable convergence in the 18th round, while the traditional method requires 28 rounds. This fully demonstrates that the frequency domain steady-state scoring can reduce the interference of oscillations on the training path, allowing the model to enter the correct optimization track more quickly. The fluctuation range of the minimum loss point decreased from 0.19 to 0.07, further proving that the model converges more stably and with stronger consistency under the method of this invention.

[0034] Based on the above data, this invention not only accelerates the training of large models, but also significantly enhances training stability, reduces resource consumption, and improves the quality of the final model, demonstrating clear and quantifiable technological advancements.

[0035] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for training large AI models based on deep learning, characterized in that, Includes the following steps: Acquire training data including text data, image data, and structured data, along with a large-scale deep learning model to be trained. Construct an attention acceleration mechanism and a multi-stage steady-state training scheduling mechanism. Input the training data into the large-scale deep learning model and perform the first forward propagation to obtain initial training state information. Based on the initial training state information, stability monitoring is performed using a multi-stage steady-state training scheduling mechanism to obtain loss change characteristics, gradient change characteristics, and attention distribution characteristics, generating training stability indicators, and calculating steady-state score sequences through frequency domain analysis. The steady-state score sequence is input into the attention acceleration mechanism to generate the attention acceleration configuration corresponding to the current training stage. The attention structure of the large-scale deep learning model is updated based on the attention acceleration configuration to obtain the staged training model structure. Based on the phased training model structure, forward and backward propagation are performed to obtain the phased training results, which are then fed back to the multi-phase steady-state training scheduling mechanism. Based on the stage training results, the steady-state score sequence is recalculated, the stage switching determination is completed, the scheduling instruction for the next training cycle is generated, and when training instability is detected, a rollback operation is performed to update the attention acceleration configuration, resulting in the updated stage training configuration. The updated stage training configuration is input into the large-scale deep learning model and the training cycle is repeated until the termination condition is met, and finally the trained large-scale deep learning model is obtained. The generation of the staged training model structure specifically includes: Read the steady-state score corresponding to the current training cycle from the steady-state score sequence and compare it with the pre-set thresholds for the warm-up stage, the transition stage, and the steady-state stage in the multi-stage steady-state training scheduling mechanism to determine the training stage label to which the current training cycle belongs and obtain the current training stage status information. Based on the current training phase state information, the steady-state score is normalized, and the target attention sparsity corresponding to the current training cycle is calculated according to the normalized steady-state score and attention sparsity configuration. Based on the preset hierarchical sensitivity parameters, hierarchical scaling is performed, and the hierarchical sensitivity parameters of each attention layer are adjusted to obtain the hierarchical target attention sparsity for each attention layer. Based on the linear mapping relationship between the normalized steady-state score and the minimum and maximum label merging ratio thresholds corresponding to the label merging configuration in the current training phase, calculate the target label merging ratio for the current training period. Based on the threshold division relationship between the normalized steady-state score and the numerical accuracy configuration in the current training phase, the target numerical accuracy level corresponding to the current training cycle is determined. The hierarchical target attention sparsity, target label merging ratio and target numerical accuracy level are combined to form the attention acceleration configuration corresponding to the current training cycle, and the updated attention acceleration configuration is obtained. The updated attention acceleration configuration is applied to each attention layer of the large-scale deep learning model. The attention connection retention mask, redundant label merging rules and attention calculation numerical precision mode of each attention layer are uniformly updated. The updated large-scale deep learning model is output as a staged training model structure.

2. The method for training a large AI model based on deep learning according to claim 1, characterized in that, The generation of the initial training state information specifically includes: Acquire large-scale, multi-source training data for training, distinguish them according to data type, unify the format of the original training data, handle missing values ​​and remove outlier samples to obtain a structured training sample set. Deep learning preprocessing operations are performed on the structured training sample set, and the training sample set is divided according to preset rules to obtain the batch training data set; Construct the network structure of the large-scale deep learning model to be trained, initialize the parameter set of the large-scale deep learning model, associate the parameter set with the network structure, and obtain the initialized large-scale deep learning model. Based on the initialization of the large-scale deep learning model, the internal configuration items of the attention acceleration mechanism and the multi-stage steady-state training scheduling mechanism are predefined and set to the initial state to obtain the initial mechanism configuration information. The first batch of training data is selected from the batch training dataset, associated with the initialized large-scale deep learning model and the initial mechanism configuration information, and the first forward propagation is performed to obtain the output representation and initial loss value of the large-scale deep learning model on the first batch of training data. During the first forward propagation, the attention weight distribution corresponding to each attention layer of the large-scale deep learning model is calculated to obtain the attention distribution matrix set corresponding to the first batch of training data. The initial training state information is constructed based on the initial loss value and the attention distribution matrix set.

3. The method for training a large AI model based on deep learning according to claim 1, characterized in that, The generation of the steady-state score sequence specifically includes: Read the initial loss value and initial attention distribution matrix set from the initial training state information. In the multi-stage steady-state training scheduling mechanism, establish a loss sequence cache and attention distribution sequence cache corresponding to the training rounds, and write the initial loss value and initial attention distribution matrix set into them respectively to obtain the basic training state sequence. Establish a training stability index value sequence cache that stores the training stability index values ​​of each training cycle. After completing several subsequent training cycles, the loss value of the current training cycle and the loss value of the previous training cycle are extracted from the basic training state sequence, and the difference is calculated as the loss change of the current training cycle. Using an exponentially weighted moving average, the loss value of the current training cycle and the smoothed loss value of the previous training cycle are weighted and combined according to a preset smoothing coefficient to obtain the smoothed loss value of the current training cycle. This smoothed loss value is then combined with the smoothed loss value to form a loss change feature, which is written into the loss change feature cache. During the backpropagation process of the current training cycle, the gradients of the large-scale deep learning model on the parameters of each layer are obtained and combined into the overall gradient representation of the current training cycle. The L2 norm of the overall gradient representation of the current training cycle is calculated as the gradient norm of the current training cycle and compared with the gradient norm of the previous training cycle to obtain the gradient change magnitude. This is combined with the gradient change magnitude of the current training cycle to form the gradient change feature and written into the gradient change feature cache. Read the set of attention distribution matrices corresponding to the current training period from the attention distribution sequence cache, calculate the information entropy corresponding to the attention distribution probability at each position, and calculate the attention sparsity of each attention layer based on the non-zero proportion of the attention distribution probability at each position. Normalize and aggregate to obtain the attention distribution features of the current training period and add them to the attention distribution feature cache. The loss change features, gradient change features, and attention distribution features of the current training cycle are read from the loss change feature cache, gradient change feature cache, and attention distribution feature cache, respectively. After normalization, a linear weighted sum is performed based on the preset weighting coefficients to obtain the initial training stability index value of the current training cycle. The initial training stability index value of the current training cycle and the initial training stability index values ​​corresponding to several historical training cycles are written together into the training stability index value sequence cache. Frequency domain oscillation analysis is performed to extract the frequency domain oscillation features that reflect the degree of high-frequency oscillation. The energy ratio of high-frequency components is calculated and a correction coefficient is generated and applied to the initial training stability index value of the current training cycle to obtain the training stability index value of the current training cycle. The difference between the training stability index value of the current training cycle and the preset upper bound value is used as the steady-state score of the current training cycle, thus obtaining the steady-state score sequence.

4. The method for training a large AI model based on deep learning according to claim 1, characterized in that, The generation of the training results for the aforementioned stage specifically includes: Select the batch training data corresponding to the current training period from the batch training data set, input the staged training model structure and perform forward propagation to obtain the output representation of the current batch training data and the set of loss values ​​and attention distribution matrices corresponding to the current training period, and combine them to form the forward propagation result of the current training period; The supervision labels and training objectives corresponding to the current batch of training data are used together with the output representation to construct the loss function input for the current training cycle. Backpropagation is performed based on the staged training model structure. The gradients of the trainable parameters in each attention layer, feedforward transformation layer and normalization layer are calculated in sequence. The gradients are aggregated according to the network structure to form the overall gradient representation of the current training cycle and obtain the gradient norm of the current training cycle. Based on the overall gradient representation and gradient norm of the current training cycle, gradient updates are performed on each trainable parameter in the staged training model structure to obtain the set of parameters of the large-scale deep learning model after the current training cycle. The loss values, overall gradient representation, gradient norm and attention distribution matrix before and after the update are organized to form stage training statistics. The stage training statistics are written into the loss sequence cache and attention distribution sequence cache in the multi-stage steady-state training scheduling mechanism. The overall gradient representation and gradient norm corresponding to the current training cycle are written into the intermediate data cache used for training stability index calculation in the multi-stage steady-state training scheduling mechanism to obtain the stage training results used for subsequent stability monitoring.

5. The method for training a large AI model based on deep learning according to claim 1, characterized in that, The generation of the updated stage training configuration specifically includes: The training results of each stage are read from the multi-stage steady-state training scheduling mechanism. The loss change characteristics, gradient change characteristics and attention distribution characteristics, including the current training cycle and several historical training cycles, are updated and calculated to obtain the training stability index value and steady-state score corresponding to the current training cycle, forming an updated steady-state score sequence. The steady-state score corresponding to the current training cycle is compared with the pre-set thresholds for the warm-up stage, transition stage, and steady-state stage in the multi-stage steady-state training scheduling mechanism to determine the target training stage label for the next training cycle and obtain the stage switching judgment result. Based on the stage switching determination result, the attention sparsity configuration boundary, label merging configuration boundary and numerical accuracy configuration level set corresponding to the target training stage label are found in the multi-stage steady-state training scheduling mechanism. The steady-state score and the target training stage label are written into the scheduling control logic to generate a scheduling instruction for the next training cycle. Based on the training stability index value and stage training statistics of the current training cycle, training instability detection is performed in the multi-stage steady-state training scheduling mechanism. The training stability index value is compared with the preset instability threshold, and the loss change and gradient norm change magnitude of several adjacent training cycles are statistically analyzed. When the training stability index value exceeds the instability threshold, the current training cycle is marked as a training unstable state, and a training instability label is generated. When the training instability flag is true, a rollback operation is performed based on the current training stage state information and the target training stage label to update the attention acceleration configuration. The target training stage label is rolled back from the steady state stage to the transition stage, or from the transition stage to the warm-up stage. In the corresponding training stage after rollback, the attention acceleration configuration is readjusted according to the preset rollback ratio to obtain the updated stage training configuration. When the training instability is marked as false, the target training stage label and attention acceleration configuration corresponding to the stage switching judgment result are output as the updated stage training configuration.

6. The method for training a large AI model based on deep learning according to claim 1, characterized in that, The generation of the trained large-scale deep learning model specifically includes: Read the updated stage training configuration, associate it with the parameter set of the large-scale deep learning model in the current training cycle, obtain the stage training initialization information for the next training cycle, and write the target training stage label and the corresponding attention sparsity configuration, label merging configuration and numerical precision configuration into the stage training initialization information. Based on the initialization information of the stage training, the attention sparsity configuration, label merging configuration and numerical precision configuration in the attention acceleration mechanism are loaded synchronously and applied to each attention layer of the large-scale deep learning model. They are then associated with the next batch of training data to complete forward propagation and backward propagation and obtain the stage training results. The training results of each stage are written into the loss sequence cache, attention distribution sequence cache and intermediate data cache in the multi-stage steady-state training scheduling mechanism, the training stability index value and steady-state score sequence are updated, and stability evolution information for termination condition determination is generated. Based on stability evolution information, the training stability index value and validation loss value of the current and historical training cycles are read from the multi-stage steady-state training scheduling mechanism, their change range is calculated, and the termination condition is compared with the maximum training round, the minimum training stability change threshold and the minimum validation loss change threshold to obtain the termination condition judgment result. When the termination condition is not met, the updated stage training configuration is written into the multi-stage steady-state training scheduling mechanism, the training round count is updated, and the staged training model structure and stage training initialization information are retained to form a cross-cycle training closed loop. When the termination condition is met, the parameter set of the large-scale deep learning model in the current training cycle is frozen, and the staged training model structure is regarded as the large-scale deep learning model that has been trained.

Citation Information

Patent Citations

  • Large model accelerated training method based on staged learning

    CN118940802A

  • Deep learning model training method and deep learning model training system

    WO2025112801A1