Large model training method based on second-order matrix optimization

By adopting row-column vector decomposition, dynamic sparse sampling and fusion compensation mechanisms in deep learning, the problem of high second-order matrix storage and computing costs in large model training is solved, and the memory usage is reduced and the computing efficiency is improved.

CN120124701AInactive Publication Date: 2025-06-10BEIJING DIGITAL FUTURE TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510291568.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In deep learning, as the scale of model parameters increases, the storage and calculation costs of second-order matrices increase sharply. The existing technology reduces complexity through matrix decomposition or sparseness, but there are problems such as accumulation of estimation errors and static sparse sampling that cannot adapt to the low synchronization efficiency in training dynamics and distributed computing.

Method used

A large-model training method based on second-order matrix optimization is proposed. Through row-column vector decomposition, dynamic sparse sampling and fusion compensation mechanism, the video memory usage is reduced, the computing efficiency is improved and the training process is stable. Specific steps include statistical vector generation in row and column directions, distributed chunking calculation, low-rank approximation matrix reconstruction, dynamic sampling rate adjustment and asynchronous execution mechanism.

Benefits of technology

The video memory usage is greatly reduced, the computing efficiency is improved, and the training process is stable, which is suitable for large models with a scale of 100 billion parameters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure QLYQS_3
    Figure QLYQS_3
  • Figure QLYQS_5
    Figure QLYQS_5
  • Figure QLYQS_7
    Figure QLYQS_7
Patent Text Reader

Abstract

The invention discloses a large model training method based on second-order matrix optimization, and belongs to the field of deep learning model training technology optimization, and the large model training method based on second-order matrix optimization comprises the following steps: S1, decomposing a second-order matrix into row and column vectors, and carrying out sliding average and distributed block reduction storage; s2, generating a statistical row vector by combining row gradient aggregation with a historical attenuation factor; s3, carrying out block distributed statistics in the column direction and synchronously generating column vectors across equipment; s4, a low-rank matrix is constructed through outer products of row and column vectors, and estimation precision is improved through noise suppression; s5, performing dynamic sparse sampling, performing initial high-density focusing, and stabilizing the sampling rate of a key layer; s6, the sampling points execute time sequence attenuation updating, and asynchronous calculation is carried out to improve the resource utilization rate; s7, performing Gaussian kernel smoothing on a neighborhood value compensation coverage gap in an unsampled region; s8, fusing low-rank estimation and sparse data, and balancing global precision by self-adaptive weight; the method has the beneficial effects of reducing video memory occupation, and improving distributed computing efficiency and balance training precision and speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning model training, and more specifically, to a large model training method based on second-order matrix optimization. Background Art

[0002] In deep learning, adaptive optimization algorithms (such as Adam, AdaGrad) rely on second-order matrices (statistics of gradient squares) to adjust the learning rate. As the scale of model parameters increases (such as trillions of parameters), the storage and computational costs of second-order matrices rise sharply. Traditional methods need to store the complete m×n matrix, resulting in high video memory occupancy, and the computational complexity of full matrix update is O(mn), making it difficult to scale to large model scenarios. Existing technologies attempt to reduce complexity through matrix decomposition or sparsification, but there are the following problems: low-value approximation ignores local features, leading to the accumulation of estimation errors; static sparse sampling cannot adapt to the training dynamics, and the accuracy in key areas is insufficient; in distributed computing, the cross-device synchronization efficiency is low, and the consistency of statistical results is poor. Therefore, there is an urgent need for a second-order matrix processing method that takes into account global statistical accuracy, local computational efficiency, and video memory optimization. Summary of the Invention

[0003] The present invention proposes a large model training method based on second-order matrix optimization, aiming to solve the storage and computational bottlenecks of second-order matrices in large model training. Through row-column vector decomposition, dynamic sparse sampling, and fusion compensation mechanism, it realizes the reduction of video memory occupancy, the improvement of computational efficiency, and the stability of the training process.

[0004] Technical Solution: A large model training method based on second-order matrix optimization includes the following steps:

[0005] S1. Decompose the original second-order matrix with dimensions of m×n into statistical vectors in the row and column directions, where the row vectors are generated by sliding average of historical gradient square statistics along the rows, and the column vectors are generated by distributed block calculation and multi-device synchronization;

[0006] S2. Aggregate gradient data along the row direction of the parameter matrix, use a predefined decay factor to balance the contributions of current and historical gradients, and generate row vectors representing the statistical features in the row direction through parallel computing of distributed row partitioning and multi-device synchronization, providing a row dimension basis for matrix approximate reconstruction;

[0007] S3. Perform distributed statistics on the column direction of the parameter matrix, allocate matrix blocks to multiple computing devices to independently process local data, and integrate the global column statistical results through a cross-device synchronization protocol to generate vectors representing column features, which together with the row vectors support low-rank approximation;

[0008] S4. Generate a low-rank approximation matrix by performing an outer product operation on the row vector and the column vector, replacing the storage and calculation of the original second-order matrix. During this process, introduce a noise suppression mechanism to improve the estimation accuracy, and achieve vector-level storage to carry matrix-level statistical information;

[0009] S5. Traverse the parameter matrix based on the principle of spatial continuity. At the initial stage of training, select local continuous regions as the calculation focus with a higher density, dynamically reduce the sampling density as the training progresses, and maintain a stable sampling rate for the key network layers to locate the subset of parameters that need to be accurately calculated;

[0010] S6. Perform real-time second-moment updates with historical decay on the selected sampling positions, retain the temporal evolution characteristics, skip the calculation of non-sampling regions to reduce the computational amount, and improve the utilization rate of computing resources through an asynchronous execution mechanism;

[0011] S7. Compensate the data for the non-sampled positions through a neighborhood weighted aggregation algorithm, use a Gaussian kernel within a preset range to smoothly propagate the calculated accurate values, and generate a spatially continuous global estimation result to make up for the coverage gap caused by sparse sampling;

[0012] S8. Integrate the estimation results of low-rank approximation and sparse propagation, balance global statistics and local accuracy through a dynamic weight mechanism, and combine learning rate adaptive adjustment and momentum correction techniques to drive the parameter update process to converge stably.

[0013] Preferably, the S1 specifically includes the following steps:

[0014] The method for generating the gradient statistical vector of row-column decomposition, that is, the row-direction statistical vector and the column-direction statistical vector are calculated according to the following rules:

[0015] S1-1. Generate the direction statistical vector. For each row (dimension 1×n) of the parameter matrix, calculate the weighted statistic of the historical gradient square through exponential moving average (EMA). The formula is:

[0016]

[0017] where β 2 is the decay factor, is the square of the gradient of the i-th row and j-th column of the parameter matrix;

[0018] S1-2. Generate the column-direction statistical vector. For each column (dimension m×1) of the parameter matrix, calculate the column statistic of the current gradient square through global averaging. The formula is:

[0019]

[0020] Generate the column vector Avoid historical hysteresis;

[0021] The video memory occupancy is reduced from m×n to m + n, providing a basis for subsequent low-rank approximation, i.e., claim 4.

[0022] Preferably, the S2 specifically includes the following steps:

[0023] Through the row-direction statistical vector generation method in claim 2, it is further implemented by parallel computing with distributed row-block partitioning and multi-device synchronization:

[0024] S2-1. Row-block partitioning: The parameter matrix is partitioned into K sub-blocks by rows, and each sub-block is assigned to an independent GPU to calculate local row vectors:

[0025]

[0026] where k is the GPU number and K is the number of partitions; is the squared gradient value of the i-th row and j-th column in the k-th sub-block;

[0027] β 2 is the decay factor (default 0.999) defined in claim 2, which is used to balance the historical gradient statistic and the current gradient contribution;

[0028] By introducing the predefined decay factor β in claim 2 2 , the dynamic balance between the historical gradient statistic and the current gradient contribution is achieved in the distributed row-block calculation, ensuring the temporal consistency of row vector generation and providing a reliable input for subsequent low-rank approximation, i.e., claim 5.

[0029] S2-2. Global integration: Aggregate the local row vectors of all GPUs through the AllReduce protocol to generate a global row vector:

[0030]

[0031] Distributed computing accelerates the generation of row statistics, and the global row vector is consistent with the mathematical definition in claim 2.

[0032] Preferably, the S3 specifically includes the following steps:

[0033] Through the column-direction statistical vector generation method in claim 2, it is further implemented by distributed block calculation:

[0034] S3-1. Column-block partitioning: The parameter matrix is partitioned into B sub-blocks by columns, and each sub-block is assigned to an independent GPU to calculate local column statistics:

[0035]

[0036] where b is the sub - block number and B is the number of sub - blocks;

[0037] S3 - 2. Global integration, summarize the local statistical results of all sub - blocks to generate a global column vector:

[0038]

[0039] Global column vector is consistent with the mathematical definition in Claim 2.

[0040] Preferably, S4 specifically includes the following steps:

[0041] S4 - 1. Outer - product reconstruction, perform an outer - product operation on the row vector r in Claim 2 (t) and the column vector c in Claim 4 (t) to generate a rank - 1 approximate matrix:

[0042]

[0043] S4 - 2. Noise suppression and regularization correction, by introducing a noise suppression mechanism, add an adaptive regularization term to suppress the noise error and rank - deficiency problem in the low - rank approximation, and improve the estimation accuracy:

[0044]

[0045] where I is the m×n identity matrix.

[0046] The storage requirement is reduced from m×n to m + n. The regularization term enhances the numerical stability of the reconstructed matrix through the noise suppression mechanism and provides input for the fusion in Claim 8.

[0047] S4 - 3. Vector - level storage implementation, by only storing the row vector and the column vector (a total of m + n elements), replace the m×n storage requirement of the original second - order matrix, carry matrix - level statistical information with vector - level storage, and at the same time compensate for the low - rank approximation error through regularization correction to ensure the integrity of the statistical information.

[0048] Preferably, S5 specifically includes the following steps:

[0049] S5 - 1. Parameter subset localization, traverse the parameter matrix based on the principle of spatial continuity, and select a locally continuous region as the calculation focus with a high sampling density at the initial stage of training;

[0050] S5 - 2. Sampling rate decay, the sampling rate d(t) decays exponentially with the number of training epochs t:

[0051] d(t)=d 0 ·e -γt +d min

[0052] where d 0 = 0.8 is the initial sampling rate, γ = 0.05 is the decay rate, and d min = 0.2 is the lower limit;

[0053] S5-3. Constant sampling for key layers, maintaining d(t) ≥ 0.7 for the attention layer and MLP layer in the Transformer, while allowing it to drop to d for the Embedding layer min .

[0054] The computational complexity is reduced by 50%–70%, and at the same time, high-precision sampling is maintained in the key area, providing input for the local update of claim 7.

[0055] Preferably, S6 specifically includes the following steps:

[0056] S6-1. Sampling area update, performing exponential moving average (EMA) update with historical decay on the subset of dynamically located parameters in S5, i.e., the sampling area set Ω at the current time step t t , retaining the time series evolution characteristics:

[0057]

[0058] If (i,j) ∈ Ω t , where β 2 is the decay factor (default 0.999), consistent with the row vector EMA parameter in claims 2 and 3.

[0059] S6-2. Non-sampling area compensation, directly inheriting the second moment value of the previous moment for the non-sampled positions :

[0060]

[0061] By skipping the calculation of non-critical areas, the total computational complexity is reduced by 40%–60%, and at the same time, the time series evolution characteristics are retained, providing a basis for the compensation mechanism of claim 8.

[0062] S6-3. Asynchronous execution mechanism, decoupling the EMA update of the sampling area from the freezing operation of the non-sampling area, and implementing asynchronous execution through GPU pipeline parallelism and communication hiding technology, reducing the device waiting time and increasing the resource utilization rate by 20%–30%;

[0063] Preferably, S7 specifically includes the following steps:

[0064] S7-1. Neighborhood weighted aggregation algorithm compensation calculation, for the non-sampled position (i,j), generating Gaussian kernel weights according to the spatial distance of the sampled points within its neighborhood :

[0065]

[0066] Among them The bandwidth σ(t) of the Gaussian kernel linearly decays with the number of training rounds t:

[0067]

[0068] S7-2. Gaussian kernel smooth propagation and global estimation, compensating the second moment value of the unsampled position through neighborhood weighted averaging:

[0069]

[0070] Generate spatially continuous global statistical results The estimation error of the covered gap is reduced from 15% to 5%, and the computational overhead is controlled to 3% of the total computational amount;

[0071] S7-3. Continuity verification and error feedback, performing local consistency verification on the compensation result. If the difference in statistics between adjacent regions exceeds the threshold δ, the neighborhood range is dynamically adjusted Ensure that the spatial continuity meets the preset conditions.

[0072] Preferably, the S8 specifically includes the following steps:

[0073] S8-1. Dynamic fusion, performing weighted fusion on the low-rank approximation matrix in claim 5 and the sparse sampling result in claim 7, where the weight linearly decays with training:

[0074]

[0075] Among them Linearly decays with training;

[0076] S8-2. Learning rate adaptation, adjusting the learning rate according to the trace of the fusion matrix:

[0077]

[0078] S8-3. Momentum correction, introducing a momentum term to accelerate parameter update and suppress oscillations during parameter update. When calculating the parameter update amount, not only the current gradient information is considered, but also the previous momentum accumulation is combined. Let the momentum coefficient be μ (the value range is usually between 0 and 1, for example, it can be set to 0.9), the momentum at the previous moment is m (t-1) , and the gradient at the current moment is h (t) , then the momentum update formula at the current moment is:

[0079] m (t) = μ·m (t-1) +(1 - μ)·g(t)

[0080] When updating the parameters, adjust according to the momentum and learning rate, and the update formula is:

[0081] θ (t) = θ (t-1) - η(t)·m (t)

[0082] Among them, θ (t) represents the parameter value at the current moment, and θ (t-1) represents the parameter value at the previous moment. In this way, combined with the adaptive adjustment of the learning rate, jointly drive the parameter update process to converge stably, and fully implement the content described in S8 of Claim 1.

[0083] Compared with the prior art, the advantages of the present invention are:

[0084] (1) Significantly reduced video memory occupancy: The traditional method needs to store a complete m×n second-order matrix (video memory complexity O(mn)). The present invention reduces the storage requirement to O(m + n) through row and column vector decomposition, and the video memory occupancy is reduced by more than 90%, especially suitable for large models with hundreds of billions of parameters.

[0085] (2) Efficient cooperation in distributed computing: Existing low-rank approximation technologies are often inefficient due to limited computing power of a single device. The present invention adopts row / column block distributed computing and combines the AllReduce protocol to synchronize global results, which not only avoids memory overflow but also improves parallel efficiency.

[0086] (3) Dynamic adaptability: Existing sparse sampling is mostly at a fixed density (such as randomly sampling 30% of the positions) and cannot adapt to the requirements during the training phase. The present invention dynamically adjusts the sampling rate through an exponential decay function: d(t) = d 0 ·e -γt + d min High-density sampling (d 0 = 0.8) at the initial stage captures key features, and later reduces to d min = 0.2, and the total calculation amount is reduced by 50% - 70%.

[0087] (4) Scalability for industrial applications: Supports multi-GPU / TPU cluster deployment, suitable for training large models such as GPT, BERT, and ViT. The hardware cost is reduced by 40%, and the training cycle is shortened by 30%. Specific implementation mode

[0088] Example

[0089] A large model training method based on second-order matrix optimization, comprising the following steps:

[0090] S1. Decompose the original second-order matrix with dimensions of m×n into statistical vectors in the row and column directions, where the row vectors are generated by sliding and averaging the historical gradient squared statistics along the rows, and the column vectors are generated through distributed block calculation and multi-device synchronization;

[0091] S2. Aggregate the gradient data along the row direction of the parameter matrix, use a predefined decay factor to balance the current and historical gradient contributions, and generate row vectors representing the statistical characteristics in the row direction through parallel calculations of distributed row partitioning and multi-device synchronization, providing a row dimension basis for matrix approximation reconstruction;

[0092] S3. Perform distributed statistics on the column direction of the parameter matrix, allocate the matrix blocks to multiple computing devices to independently process the local data, and integrate the global column statistical results through a cross-device synchronization protocol to generate vectors representing the column characteristics, jointly supporting low-rank approximation with the row vectors;

[0093] S4. Generate a low-rank approximation matrix by performing an outer product operation on the row vectors and column vectors, replacing the storage and calculation of the original second-order matrix. Introduce a noise suppression mechanism during this process to improve the estimation accuracy, and achieve vector-level storage to carry matrix-level statistical information;

[0094] S5. Traverse the parameter matrix based on the principle of spatial continuity, select local continuous regions as the calculation focus with a higher density at the initial stage of training, dynamically reduce the sampling density as the training progresses, maintain a stable sampling rate for key network layers, and locate the subset of parameters that need to be accurately calculated;

[0095] S6. Perform real-time second-order moment updates with historical decay on the selected sampling positions, retain the temporal evolution characteristics, skip the calculation of non-sampling regions to reduce the computational amount, and improve the utilization rate of computing resources through an asynchronous execution mechanism;

[0096] S7. Compensate the data for the non-sampled positions through a neighborhood weighted aggregation algorithm, use a Gaussian kernel within a preset range to smoothly spread the calculated accurate values, and generate a spatially continuous global estimation result to fill the coverage gaps caused by sparse sampling;

[0097] S8. Integrate the estimation results of low-rank approximation and sparse propagation, balance the global statistics and local accuracy through a dynamic weight mechanism, and combine the learning rate adaptive adjustment and momentum correction techniques to drive the parameter update process to converge stably.

[0098] The first type of optimization embodiment, the low-rank approximation reconstruction method includes steps S1, S2, S3, and S4:

[0099] The specific steps of S1 are as follows:

[0100] The method for generating gradient statistical vectors for row-column decomposition, that is, the row direction statistical vectors and the column direction statistical vectors The calculation follows the following rules:

[0101] S1-1. Generate the direction statistical vector. For each row of the parameter matrix (dimension 1×n), calculate the weighted statistic of the historical gradient square through the Exponential Moving Average (EMA). The formula is:

[0102]

[0103] Parameter description:

[0104] represents the element of the row direction statistical vector corresponding to the i-th row of the parameter matrix at the t-th moment;

[0105] is the corresponding element at the (t-1)-th moment, used to retain historical information;

[0106] β 2 is the decay factor, and its value range is usually between 0 and 1. Here, the default value is 0.999; it controls the weight of the historical gradient statistic in the current calculation. The closer β 2 is to 1, the greater the influence of the historical gradient statistic; conversely, the greater the influence of the current gradient;

[0107] represents the square of the gradient of the i-th row and j-th column of the parameter matrix, reflecting the gradient change amplitude of the parameter at this position. By taking the weighted average of the squares of the gradients of all columns in a row (the weight is ), and combining with the historical statistic, the element of the statistical vector of the current row is obtained.

[0108] S1-2. Generate the column direction statistical vector. For each column of the parameter matrix (dimension m×1), calculate the column statistic of the current gradient square through the global average. The formula is:

[0109]

[0110] Generate the column vector Avoid historical hysteresis;

[0111] The video memory occupancy is reduced from m×n to m + n, providing a basis for the subsequent low-rank approximation, that is, claim 4;

[0112] represents the element of the column direction statistical vector corresponding to the j-th column of the parameter matrix at the t-th moment;

[0113] For all m rows in the j-th column, the squares of the gradients are globally averaged (divided by m) to obtain the element of the statistical vector of this column. This calculation method avoids historical hysteresis and only focuses on the column gradient information at the current moment.

[0114] Technical effect:

[0115] Reduce the video memory occupancy of the original second-order matrix from m×n to m + n, providing a basis for subsequent low-rank approximation operations, reducing storage overhead, and facilitating large model training under resource constraints.

[0116] Preferably, step S2 specifically includes the following steps:

[0117] Through the row direction statistical vector generation method in claim 2, further implemented through parallel computing of distributed row block division and multi-device synchronization:

[0118] S2-1. Row block division, divide the parameter matrix into K sub-blocks by rows, and allocate each sub-block to an independent GPU to calculate the local row vector:

[0119]

[0120] where k is the GPU number and K is the number of blocks; is the squared gradient value of the i-th row and j-th column in the k-th sub-block;

[0121] β 2 is the decay factor (default 0.999) defined in claim 2, used to balance the historical gradient statistic and the current gradient contribution;

[0122] Technical effect:

[0123] By introducing the predefined decay factor β in claim 2 2 , achieve the dynamic balance of historical gradient statistic and current gradient contribution in distributed row block calculation, ensure the temporal consistency of row vector generation, and provide a reliable input for subsequent low-rank approximation, i.e., claim 5.

[0124] S2-2. Global integration, aggregate the local row vectors of all GPUs through the AllReduce protocol to generate the global row vector:

[0125]

[0126] Distributed computing accelerates row statistic generation, and the global row vector is consistent with the mathematical definition in claim 2.

[0127] Step S3 specifically includes the following steps:

[0128] Through the column direction statistical vector generation method in claim 2, further implemented through distributed block calculation:

[0129] S3-1. Column Block Division: Divide the parameter matrix into B sub-blocks by column, and allocate each sub-block to an independent GPU to calculate the local column statistics:

[0130]

[0131] Parameter Description:

[0132] b is the sub-block number, and B is the number of blocks. The parameter matrix is divided into B sub-blocks by column, and each sub-block is calculated for local column statistics by an independent GPU. Each GPU processes n / B columns of data;

[0133] represents the j-th element of the local column statistic vector corresponding to the sub-block numbered b at the t-th moment;

[0134] is the squared gradient value of the i-th row and j-th column in the b-th sub-block, which is used to calculate the elements of the local column statistic vector.

[0135] S3-2. Global Integration: Aggregate the local statistical results of all sub-blocks to generate a global column vector:

[0136]

[0137] represents the element of the global column vector at the t-th moment. By aggregating the local column statistics of all B sub-blocks and then taking the average, a global column vector consistent with the mathematical definition in Claim 2 is obtained.

[0138] Technical Effect: The distributed calculation of column statistics avoids the problem of out-of-memory on a single device, enables the efficient calculation of large-scale matrix column statistics on multiple devices, and at the same time ensures the calculation accuracy of the global column vector, providing accurate column vector data for low-rank approximation.

[0139] The specific steps of S4 are as follows:

[0140] S4-1. Outer Product Reconstruction: Perform an outer product operation on the row vector r of Claim 2 (t) and the column vector c of Claim 4 (t) to generate an approximate matrix of rank 1:

[0141]

[0142] S4-2. Noise Suppression and Regularization Correction: By introducing a noise suppression mechanism and adding an adaptive regularization term to suppress the noise error and rank deficiency problem in low-rank approximation, the estimation accuracy is improved:

[0143]

[0144] Parameter description:

[0145] Update based on the original approximate matrix; λ is the coefficient of the adaptive regularization term, and I is the m×n identity matrix;

[0146] ‖r (t) ‖ 2 and ‖c (t) ‖ 2 are the 2-norms of the row vector r (t) and the column vector c (t) respectively, which are used to measure the length of the vector. The coefficient of the regularization term λ is determined by the relationship between the product of the 2-norms of the row vector and the column vector and the matrix dimensions m and n.

[0147] Technical effect:

[0148] The storage requirement is reduced from m×n to m + n, significantly reducing the storage overhead. The numerical stability of the reconstructed matrix is enhanced through the noise suppression mechanism. At the same time, the matrix-level statistical information is carried by the vector-level storage, ensuring the integrity of the statistical information and providing a reliable input for the fusion of claim 8.

[0149] S4-3. Implementation of vector-level storage, by storing only the row vector and the column vector (a total of m + n elements), replacing the m×n storage requirement of the original second-order matrix, carrying the matrix-level statistical information by the vector-level storage, and at the same time compensating for the low-rank approximation error through regularization correction to ensure the integrity of the statistical information.

[0150] The second type of optimization embodiment, the sparse sampling compensation method includes steps S5, S6, and S7:

[0151] The specific steps of S5 are as follows:

[0152] S5-1. Parameter subset positioning, traversing the parameter matrix based on the principle of spatial continuity, and selecting a locally continuous region as the calculation focus with a high sampling density at the initial stage of training;

[0153] S5-2. Sampling rate decay, the sampling rate d(t) decays exponentially with the number of training rounds t:

[0154] d(t) = d 0 ·e -γt +d min

[0155] Parameter description:

[0156] d(t) represents the sampling rate at the training round t, and d 0 is the initial sampling rate, which is taken as 0.8 here, representing the sampling ratio at the initial stage of training;

[0157] γ is the decay rate, with a value of 0.05, which controls the decay speed of the sampling rate with respect to the number of training rounds t, and e -γt makes the sampling rate decrease exponentially as the number of training rounds increases, d min is the lower limit of the sampling rate, with a value of 0.2, ensuring that the sampling rate does not decrease infinitely.

[0158] S5-3. Constant sampling for key layers. For the attention layer and MLP layer in Transformer, maintain d(t) ≥ 0.7, while for the Embedding layer, allow it to drop to d min 。

[0159] Technical effect: The computational cost is reduced by 50% - 70%. By dynamically adjusting the sampling rate, the computational cost is gradually reduced during training. At the same time, a stable sampling rate is maintained for key network layers, ensuring that key regions can maintain high-precision sampling, providing an accurate subset of parameters for subsequent local updates, and not affecting the training effect of key parts of the model while reducing computational resource consumption.

[0160] The said S6 specifically includes the following steps:

[0161] S6-1. Sampling region update. For the subset of parameters dynamically located in S5, i.e., the set of sampling regions Ω at the current time step t t , perform an exponential moving average (EMA) update with historical decay to retain the temporal evolution characteristics:

[0162]

[0163] Parameter description:

[0164] represents the second moment value at position (i, j) in the parameter matrix at the t-th moment, is the second moment value at this position in the previous moment;

[0165] β 2 is the decay factor, with a default value of 0.999, used to balance the weights of the historical second moment value and the current gradient square in the update. This update operation is only performed when the position (i, j) is within the set of sampling regions Ω at the current time step t t inside.

[0166] S6-2. Non-sampling region compensation. For the un-sampled positions directly inherit the second moment value of the previous moment:

[0167]

[0168] By skipping non-critical region calculations, the total computational load is reduced by 40%–60%, while preserving the temporal evolution characteristics, providing a basis for the compensation mechanism of claim 8.

[0169] S6-3. Asynchronous execution mechanism, which decouples the EMA update of the sampling region from the freezing operation of the non-sampling region, and realizes asynchronous execution through GPU pipeline parallelism and communication hiding technology, reducing the device waiting time and increasing the resource utilization rate by 20%–30%;

[0170] The said S7 specifically includes the following steps:

[0171] S7-1. Neighborhood weighted aggregation algorithm compensation calculation. For the unsampled position (i,j), according to the spatial distance of the sampled points within its neighborhood generate Gaussian kernel weights:

[0172]

[0173] w p,q is the Gaussian kernel weight corresponding to the sampled point (p,q) within the neighborhood N(i,j) of the unsampled position (i,j). ‖(p,q)-(i,j)‖ 2 represents the square of the Euclidean distance between the sampled point (p,q) and the unsampled point (i,j);

[0174] where the Gaussian kernel bandwidth σ(t) decays linearly with the number of training epochs t, and the formula is:

[0175]

[0176] T is the total number of training epochs. As the training progresses, the bandwidth gradually decreases, making the weights of the nearby neighbor points within the neighborhood relatively increase and the weights of the far neighbor points relatively decrease.

[0177] S7-2. Gaussian kernel smoothing propagation and global estimation, compensating the second moment value of the unsampled position through neighborhood weighted averaging:

[0178]

[0179] By performing weighted averaging on the second moment values of the sampled points (p,q) within the neighborhood N(i,j) p,q according to the Gaussian kernel weights w

[0180] the second moment estimated value of the unsampled position (i,j) is obtained Ensure that the spatial continuity meets the preset conditions.

[0181] Technical effect: The data is compensated by the neighborhood weighted aggregation algorithm to generate a spatially continuous global estimation result. The estimation error covering the gap is reduced from 15% to 5%, and the computational overhead is controlled at 3% of the total computational volume. By dynamically adjusting the neighborhood range to ensure that the spatial continuity meets the preset conditions, the problem of coverage gaps caused by sparse sampling is effectively compensated, and the accuracy of global estimation is improved.

[0182] S8 specifically includes the following steps:

[0183] S8-1. Dynamic fusion: The low-rank approximation matrix in claim 5 and the sparse sampling result in claim 7 are weighted and fused, where the weight decays linearly with training:

[0184] Dynamic fusion formula:

[0185] where

[0186] Parameter description:

[0187] is the final fusion result matrix at the t-th moment, is the low-rank approximation matrix in claim 5, is the sparse sampling result matrix in claim 7;

[0188] α(t) is the fusion weight, which decays linearly with the training round t. T is the total number of training rounds. At the beginning of training, α(t) is close to 1, and the weight of the low-rank approximation matrix is larger, emphasizing global statistical information; as training progresses, α(t) decreases, and the weight of the sparse sampling result matrix increases, paying more attention to local accuracy.

[0189] S8-2. Learning rate adaptation: Adjust the learning rate according to the trace of the fusion matrix:

[0190]

[0191] Parameter description:

[0192] η(t) is the learning rate at the t-th moment, η 0 is the initial learning rate, is the final fusion matrix of the trace, used to measure a certain global feature of the matrix;

[0193] ∈ is a very small positive number, usually set to 1e-8, etc., to prevent the denominator from being 0 and ensure the stability of the learning rate calculation.

[0194] S8-3. Momentum correction, introducing a momentum term to accelerate parameter update and suppress oscillations during the parameter update process. When calculating the parameter update amount, not only the current gradient information is considered, but also the previous momentum accumulation is combined. Let the momentum coefficient be μ (the value range is usually between 0 and 1, for example, it can be set to 0.9), and the momentum at the previous moment is m (t-1) , the gradient at the current moment is g (t) , then the momentum update formula at the current moment is:

[0195] m (t) = μ·m (t-1) +(1 - μ)·g (t)

[0196] When updating the parameters, adjust according to the momentum and the learning rate. The update formula is:

[0197] θ (t) = θ (t-1) - η(t)·m (t)

[0198] Among them, θ (t) represents the parameter value at the current moment, and θ (t-1) represents the parameter value at the previous moment. In this way, combined with the adaptive adjustment of the learning rate, jointly drive the parameter update process to converge stably, and fully implement the content described in S8 of claim 1.

[0199] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A large model training method based on second-order matrix optimization, characterized in that: The following steps are involved: S1. Decompose the original second-order matrix of dimension m×n into statistical vectors in row and column directions, where the row vectors are generated by sliding the historical gradient square statistics along the row, and the column vectors are generated by distributed block computing and multi-device synchronization; S2. Aggregate gradient data along the row direction of the parameter matrix, use a predefined attenuation factor to balance the current and historical gradient contributions, generate row vectors that represent the statistical characteristics of the row direction through distributed row block partitioning and multi-device synchronous parallel computing, and provide a row dimension basis for matrix approximate reconstruction; S3. Perform distributed statistics on the column direction of the parameter matrix, distribute the matrix blocks to multiple computing devices to independently process local data, integrate the global column statistics results through the cross-device synchronization protocol, generate vectors representing column features, and support low-rank approximation together with the row vector; S4. Generate a low-rank approximate matrix by performing outer product operations on row vectors and column vectors, replacing the storage and calculation of the original second-order matrix. In this process, a noise suppression mechanism is introduced to improve the estimation accuracy, so that matrix-level statistical information can be stored at the vector level. S5. Based on the principle of spatial continuity, the parameter matrix is ​​traversed. At the beginning of training, local continuous areas are selected as the calculation focus with higher density. The sampling density is dynamically reduced as the training progresses. The sampling rate of key network layers is maintained at a stable level to locate the parameter subsets that need to be accurately calculated. S6. Perform real-time second-order moment update with historical attenuation on the selected sampling positions, retain the time series evolution characteristics, skip calculations in non-sampling areas to reduce the amount of calculations, and improve computing resource utilization through asynchronous execution mechanism; S7. Compensate the data of the unsampled positions through the neighborhood weighted aggregation algorithm, use the Gaussian kernel in the preset range to smooth the propagation of the calculated precise value, generate a spatially continuous global estimation result, and make up for the coverage gap caused by sparse sampling; S8. By integrating the estimation results of low-rank approximation and sparse propagation, the global statistics and local accuracy are balanced through a dynamic weight mechanism. The learning rate adaptive adjustment and momentum correction technology are combined to drive the parameter update process to converge stably.

2. A large model training method based on second-order matrix optimization according to claim 1, characterized in that: The S1 specifically includes the following steps: The method of generating gradient statistics vectors by row-column decomposition is row-direction statistics vector and column-wise statistics vector The calculation follows the following rules: S1-1. Generate directional statistical vectors. For each row of the parameter matrix (dimension is 1×n), calculate the weighted statistics of the historical gradient square by exponential moving average (EMA). The formula is: Where β2 is the attenuation factor, is the square of the gradient of the i-th row and j-th column of the parameter matrix; S1-2. Generate column-wise statistical vectors. For each column of the parameter matrix (dimension is m×1), the column statistics of the current gradient square are calculated by global averaging. The formula is: Generate column vector Avoid historical lags; The memory occupancy is reduced from m×n to m+n, providing a basis for the subsequent low-rank approximation, i.e., claim 4.

3. A large model training method based on second-order matrix optimization according to claim 1, characterized in that: The S2 specifically includes the following steps: The method for generating row-wise statistical vectors in claim 2 is further implemented by distributed row block division and parallel computing synchronized with multiple devices: S2-1. Row-by-row partitioning: Divide the parameter matrix into K sub-blocks by row, and assign each sub-block to an independent GPU to calculate the local row vector: Where k is the GPU number and K is the number of blocks; is the square value of the gradient of the i-th row and j-th column in the k-th sub-block; β2 is the attenuation factor defined in claim 2 (default 0.999), which is used to balance the historical gradient statistics and the current gradient contribution; By introducing the predefined attenuation factor β2 in claim 2, a dynamic balance between historical gradient statistics and current gradient contributions is achieved in distributed row block calculations, ensuring the temporal consistency of row vector generation and providing reliable input for subsequent low-rank approximation, i.e. claim 5. S2-2. Global integration: Aggregate the local row vectors of all GPUs through the AllReduce protocol to generate a global row vector: Distributed computing accelerates row statistics generation, and global row vectors This is consistent with the mathematical definition of claim 2.

4. The large model training method based on second-order matrix optimization according to claim 1 is characterized in that: The S3 specifically includes the following steps: The column-wise statistical vector generation method in claim 2 is further implemented through distributed block calculation: S3-1. Column-by-block partitioning: Divide the parameter matrix into B sub-blocks by column, and assign each sub-block to an independent GPU to calculate local column statistics: Where b is the sub-block number and B is the number of sub-blocks; S3-2. Global integration, summarizing the local statistical results of all sub-blocks to generate a global column vector: Global column vector This is consistent with the mathematical definition of claim 2.

5. The large model training method based on second-order matrix optimization according to claim 1 is characterized in that: The S4 specifically comprises the following steps: S4-1. Outer product reconstruction, the row vector r of claim 2 (t) With the column vector c of claim 4 (t) Perform an outer product operation to generate an approximate matrix of rank 1: S4-2. Noise suppression and regularization correction: By introducing a noise suppression mechanism and adding an adaptive regularization term to suppress the noise error and rank deficiency problem in the low-rank approximation, the estimation accuracy is improved: Where I is the m×n identity matrix. The storage requirement is reduced from m×n to m+n, and the regularization term enhances the numerical stability of the reconstructed matrix through a noise suppression mechanism, providing input for the fusion of claim 8. S4-3. Vector-level storage implementation by storing only row vectors With column vector (a total of m+n elements), replacing the m×n storage requirement of the original second-order matrix, and carrying matrix-level statistical information with vector-level storage. At the same time, regularization correction is used to compensate for the low-rank approximation error to ensure the integrity of the statistical information.

6. A large model training method based on second-order matrix optimization according to claim 1, characterized in that: The S5 specifically includes the following steps: S5-1. Parameter subset positioning, traversing the parameter matrix based on the principle of spatial continuity, selecting local continuous areas as calculation focus with high sampling density at the beginning of training; S5-2. Sampling rate decay, the sampling rate d(t) decays exponentially with the training round t: d(t)=d0·e -γt +d min Where d0 = 0.8 is the initial sampling rate, γ = 0.05 is the decay rate, d min =0.2 is the lower limit; S5-3. Constant sampling of key layers, maintaining d(t) ≥ 0.7 for the attention layer and MLP layer in Transformer, and allowing it to drop to d for the Embedding layer min . The computational effort is reduced by 50%–70%, while key areas are sampled with high precision, providing input for the local update of claim 7.

7. The large model training method based on second-order matrix optimization according to claim 1 is characterized in that: The S6 specifically comprises the following steps: S6-1. Update the sampling area, the parameter subset dynamically positioned in S5, that is, the sampling area set Ω of the current time step t t , perform exponential moving average (EMA) updates with historical decay, preserving the time series evolution characteristics: If (i,j)∈Ω t , where β2 is the decay factor (default 0.999), consistent with the row vector EMA parameters in claims 2 and 3. S6-2. Non-sampled area compensation, for unsampled positions Directly inherit the second-order moment value of the previous moment: By skipping calculations in non-critical areas, the total amount of computation is reduced by 40%–60% while retaining the timing evolution characteristics, providing a basis for the compensation mechanism of claim 8. S6-3. Asynchronous execution mechanism, decouples the EMA update of the sampling area from the freeze operation of the non-sampling area, and realizes asynchronous execution through GPU pipeline parallelism and communication hiding technology, reducing device waiting time and improving resource utilization by 20%-30%.

8. The large model training method based on second-order matrix optimization according to claim 1 is characterized in that: The S7 specifically comprises the following steps: S7-1. Neighborhood weighted aggregation algorithm compensation calculation, for the unsampled position (i, j), according to its neighborhood The spatial distance of the sampled points within the Gaussian kernel is generated: in The Gaussian kernel bandwidth σ(t) decays linearly with the training round t: S7-2. Gaussian kernel smoothing and global estimation, compensating the second-order moment value of unsampled positions by neighborhood weighted averaging: Generate spatially continuous global statistics The estimation error of the coverage gap is reduced from 15% to 5%, and the computational overhead is controlled to 3% of the total computational workload. S7-3. Continuity verification and error feedback: local consistency check of compensation results. If the difference in statistics of adjacent regions exceeds the threshold δ, the neighborhood range is dynamically adjusted. Ensure that spatial continuity meets preset conditions.

9. The large model training method based on second-order matrix optimization according to claim 1 is characterized in that: The S8 specifically includes the following steps: S8-1. Dynamic fusion, by weighted fusion of the low-rank approximation matrix in claim 5 and the sparse sampling result in claim 7, where the weights decay linearly with training: in Decays linearly with training; S8-2. Learning rate adaptation, adjust the learning rate according to the trace of the fusion matrix: S8-3. Momentum correction: The momentum term is introduced to accelerate parameter update and suppress oscillation during parameter update. When calculating the parameter update amount, not only the current gradient information is considered, but also the previous momentum accumulation is combined. The momentum coefficient is set to μ (the value range is usually between 0 and 1, for example, it can be set to 0.9), and the momentum at the previous moment is m (t-1) , the gradient at the current moment is g (t) , then the momentum update formula at the current moment is: m (t) =μ·m (t-1) +(1-μ)·g (t) When the parameters are updated, they are adjusted according to the momentum and learning rate, and the update formula is: i (t) =θ (t-1) -η(t)·m (t) Among them, θ (t) represents the parameter value at the current moment, θ (t-1) In this way, combined with the adaptive adjustment of the learning rate, the parameter update process is driven to converge stably, and the content described in S8 of claim 1 is fully realized.

Citation Information

Cited By

  • Model training method and system based on alternative adaptation momentum optimization

    CN121615718A

  • Data processing method and device, computer equipment, storage medium and computer program product

    CN122113918A