Methods, apparatus, devices and storage media for updating preset models

By acquiring the parameter gradient information of computing nodes and predicting future gradient changes using a prediction model, and dynamically updating the preset model parameters, the problems of resource waste and inefficiency in existing technologies are solved, thereby improving the efficiency and stability of model training.

CN120892068BActive Publication Date: 2025-12-02INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511405256.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2025-12-02
Estimated Expiration
2045-09-29

AI Technical Summary

Technical Problem

Existing preset model update methods rely on fixed intervals or manual intervention, resulting in low resource utilization, excessive memory and bandwidth consumption, and low efficiency in updating all parameters.

Method used

By acquiring parameter gradient information from multiple computing nodes, a global gradient index is determined, and a prediction model is used to predict gradient changes in future training batches. Parameter update information is dynamically generated, and model parameters are updated only under specific conditions.

Benefits of technology

It achieves efficient use of resources, reduces the pressure on GPU memory and bandwidth, and improves the efficiency and system stability of large-scale model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120892068B_ABST
    Figure CN120892068B_ABST
Patent Text Reader

Abstract

This application provides a method, apparatus, device, and storage medium for updating a preset model, which can be applied in the field of computers. The method for updating the preset model includes: acquiring parameter gradient information of a preset model generated by multiple computing nodes in multiple training batches; determining a global gradient index for any training batch based on the multiple parameter gradient information of any training batch, thereby obtaining multiple global gradient indices for multiple training batches; predicting the multiple global gradient indices according to a preset prediction model, thereby obtaining predicted values ​​and confidence levels of the global gradient indices for future training batches; generating parameter update information based on the current parameter version number of the preset model to be updated when the global gradient index of the current training batch is lower than a preset index threshold, multiple predicted values ​​are lower than a preset prediction threshold, and multiple confidence levels are higher than a preset confidence threshold; and sending the parameter update information to multiple computing nodes to update the parameters of the preset model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computers, and more specifically to a method, apparatus, device, and storage medium for updating a preset model. Background Technology

[0002] As the scale of deep learning models continues to expand, training models with hundreds of billions of parameters has become the norm. However, related model update methods rely on fixed intervals or manual intervention, and update only all parameters, leading to wasted computing resources and excessive consumption of GPU memory and bandwidth. Summary of the Invention

[0003] This application provides a method, apparatus, device, and storage medium for updating a preset model, which can effectively solve the problems of low resource utilization and lack of adaptive timing of updates caused by existing preset model update methods that rely on fixed update intervals or manual intervention. At the same time, it can significantly reduce the memory and bandwidth pressure caused by full parameter updates and improve the efficiency of large-scale model training.

[0004] The first aspect of this application provides a method for updating a preset model, comprising: acquiring parameter gradient information of a preset model generated by multiple computing nodes in multiple training batches; determining a global gradient index for any training batch based on the multiple parameter gradient information of any training batch, thereby obtaining multiple global gradient indices for multiple training batches, wherein the global gradient indices are used to characterize the degree of change of parameter gradients of the preset model during training; predicting the multiple global gradient indices according to a preset prediction model to obtain predicted values ​​and corresponding confidence levels of the global gradient indices for future training batches; generating parameter update information based on the current parameter version number of the preset model to be updated when the global gradient index of the current training batch is lower than a preset index threshold, multiple predicted values ​​are lower than a preset prediction threshold, and multiple confidence levels are higher than a preset confidence threshold; and sending the parameter update information to multiple computing nodes to update the parameters of the preset model.

[0005] A second aspect of this application provides a device for updating a preset model, comprising: an acquisition module for acquiring parameter gradient information of a preset model generated by multiple computing nodes in multiple training batches; a determination index module for determining a global gradient index for any training batch based on the multiple parameter gradient information of any training batch, thereby obtaining multiple global gradient indices for multiple training batches, wherein the global gradient index is used to characterize the degree of change of parameter gradients of the preset model during training; a prediction module for predicting the multiple global gradient indices according to a preset prediction model, thereby obtaining predicted values ​​and corresponding confidence levels of the global gradient indices for future training batches; a determination update module for generating parameter update information based on the current parameter version number of the preset model to be updated when the global gradient index of the current training batch is lower than a preset index threshold, multiple predicted values ​​are lower than a preset prediction threshold, and multiple confidence levels are higher than a preset confidence threshold; and an update module for sending the parameter update information to multiple computing nodes to update the parameters of the preset model.

[0006] A third aspect of this application provides an electronic device comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method described above.

[0007] A fourth aspect of this application also provides a computer-readable storage medium having a computer program or instructions stored thereon, which, when executed by a processor, implement the steps of the above-described method.

[0008] The method for updating a preset model based on embodiments of this application obtains the gradient information of preset model parameters generated by multiple computing nodes in multiple training batches, determines global gradient indices based on the gradient information of parameters in each training batch, and obtains multiple global gradient indices. Then, it predicts the multiple global gradient indices based on a preset prediction model to obtain predicted values ​​and corresponding confidence levels for future training batches. If the global gradient indices of the current training batch are lower than a preset threshold, the predicted values ​​are lower than a prediction threshold, and the confidence level is higher than a preset confidence threshold, parameter update information is generated based on the current parameter version number and updated parameters are sent to the computing nodes. This achieves dynamic update decision-making, avoids resource waste, improves distributed training efficiency and resource utilization, significantly reduces the memory and bandwidth pressure caused by full parameter updates, and improves the efficiency and system stability of large-scale model training. Attached Figure Description

[0009] The above-mentioned contents, other objects, features and advantages of this application will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:

[0010] Figure 1The illustration schematically depicts an application scenario of an update method, apparatus, device, and storage medium for a preset model according to embodiments of this application.

[0011] Figure 2 A flowchart illustrating an update method for a preset model according to an embodiment of this application is shown schematically.

[0012] Figure 3 A flowchart illustrating the determination of global gradient information according to an embodiment of this application is shown schematically.

[0013] Figure 4 A flowchart illustrating the determination of a global gradient index according to an embodiment of this application is shown schematically.

[0014] Figure 5 A flowchart illustrating the updating of parameter values ​​according to an embodiment of this application is shown schematically;

[0015] Figure 6 This schematically illustrates the component architecture diagram of an update method for a preset model according to an embodiment of this application;

[0016] Figure 7 This illustration shows a schematic diagram of the parameters of a preset model during the training process according to an embodiment of this application;

[0017] Figure 8 This schematic diagram illustrates a structural block diagram of an update device for a preset model according to an embodiment of the present application;

[0018] Figure 9 A block diagram schematically illustrates an electronic device suitable for implementing an update method for a preset model according to an embodiment of this application. Detailed Implementation

[0019] The embodiments of this application will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of this application. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of this application for ease of explanation. However, it will be apparent that one or more embodiments may be implemented without these specific details. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of this application.

[0020] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The terms “comprising,” “including,” etc., as used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0021] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0022] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).

[0023] Embodiments of this application provide a method, apparatus, device, and storage medium for updating a preset model.

[0024] Figure 1 The illustration shows an application scenario diagram of the method, apparatus, device, and storage medium for updating a preset model according to embodiments of this application.

[0025] like Figure 1 As shown, the application scenario 100 according to this embodiment may include computing nodes 101, 102, 103 and a preset model update hub 104, wherein each computing node and the update hub are interconnected through a high-speed network to form a distributed communication link.

[0026] The preset model update hub 104 and the computing nodes establish bidirectional communication links (including in-band / out-of-band management networks) for information transmission and scheduling of the preset model update method.

[0027] Each computing node in computing nodes 101, 102, and 103 is equipped with relevant functions to collect parameter gradient information of a preset model generated in multiple training batches in real time, and transmit the gradient information to the update center 104 via a communication link. The update center 104 determines a global gradient index representing the training state of the model based on the parameter gradient information from each computing node; predicts the historical global gradient index based on a preset time-series prediction model to obtain the gradient prediction value and corresponding confidence level for future training batches; generates an update strategy based on preset index thresholds, prediction thresholds, and confidence level thresholds, and constructs parameter update information based on the current parameter version number.

[0028] The update hub 104 also distributes parameter update information to each computing node via a communication link. After receiving the update information, each computing node performs atomic update operations on the preset model parameters stored locally according to the update strategy, realizing collaborative update and state synchronization of parameters across multiple nodes in a distributed environment.

[0029] It should be noted that, Figure 1 The number of compute nodes shown is for illustrative purposes only and can be dynamically adjusted in actual deployment. The update hub 104 can be deployed using a highly available cluster architecture to ensure the continuity and reliability of system services.

[0030] The following will be based on Figure 1 The described scene, through Figures 2-7 The method for updating the preset model in the disclosed embodiments will be described in detail.

[0031] Figure 2 A flowchart illustrating an update method for a preset model according to an embodiment of this application is shown schematically.

[0032] like Figure 2 As shown, the update of the preset model in this embodiment includes operations S210 to S250.

[0033] In operation S210, parameter gradient information of a preset model generated by multiple computing nodes in multiple training batches is obtained.

[0034] In this embodiment, a computing node can be a physical or logical unit that undertakes part of the model's computational tasks. A training batch refers to the training process of a set of data processed by one forward and backward propagation during model training. The preset model is a machine learning model to be optimized. Parameter gradient information refers to the values ​​of the partial derivatives of the loss function of the preset model with respect to each parameter, used to describe the direction and magnitude of parameter updates.

[0035] In operation S220, for the gradient information of multiple parameters of any training batch, the global gradient index of any training batch is determined, and multiple global gradient indices of multiple training batches are obtained.

[0036] In this embodiment, the global gradient metric characterizes the degree of change in the parameter gradient of the preset model during training. Based on the global gradient metrics of multiple consecutive training batches, the fluctuation of these metrics can be determined, thereby judging the stability of the parameter gradient during model training. When the fluctuation of the global gradient metrics across multiple training batches is small, it indicates that the training of the preset model is in a relatively stable state; conversely, when the fluctuation of the global gradient metrics across multiple training batches is large, it indicates that the training of the preset model is in an unstable state.

[0037] For example, suppose there are four training batches: A, B, C, and D. For training batch A, based on its corresponding parameter gradient information b, c, and d, the global gradient index a corresponding to training batch A is obtained by following the method for determining the global gradient index for any training batch in operation S220 above. Similarly, using the same method, the global gradient index b corresponding to training batch B, the global gradient index c corresponding to training batch C, and the global gradient index d corresponding to training batch D are determined sequentially.

[0038] In operation S230, multiple global gradient indicators are predicted according to the preset prediction model to obtain the predicted values ​​and corresponding confidence levels of the global gradient indicators for multiple future training batches.

[0039] In the embodiments of this application, the preset prediction model is a model for prediction trained based on historical data. The predicted value of the global gradient index is an estimate of the global gradient index values ​​for multiple future training batches. The confidence level is used to measure the reliability of the predicted value. The confidence interval of the prediction result indicates the probability that the predicted value is close to the true value.

[0040] In operation S240, if the global gradient metric of the current training batch is lower than the preset metric threshold, multiple predicted values ​​are lower than the preset prediction threshold, and multiple confidence scores are higher than the preset confidence threshold, parameter update information is generated based on the current parameter version number of the preset model to be updated.

[0041] For example, when the global gradient metric of the current training batch is detected to be 0.04, and the preset metric threshold is 0.05, the predicted values ​​of the next three training batches are 0.02, 0.03, and 0.04, respectively, the preset prediction threshold is 0.06 and the corresponding confidence levels are 92%, 95%, and 93%, respectively, and the preset confidence threshold is 90%.

[0042] When the aforementioned predetermined conditions are met, parameter update information is generated based on the current parameter version number of the preset model to be updated. The current parameter version number is a specific code used to uniquely identify the parameter set state adopted by the model at the current moment during the training and update iteration process of the preset model. The current parameter version number records the specific version status of the preset model's parameters within this training cycle, providing a clear identification basis for model parameter tracking and update operations, ensuring the accuracy and traceability of model training and management.

[0043] During operation S250, parameter update information is sent to multiple computing nodes to update the parameters of the preset model.

[0044] In this embodiment of the application, the parameter update information is distributed to each computing node participating in distributed training through a communication network. After receiving the parameter update information, each computing node performs an atomic update operation on the parameters of the preset model stored locally, thereby completing the version synchronization and state switching of the model parameters.

[0045] It should be noted that the parameter update information is transmitted in an encrypted format, and each computing node must perform the update operation after integrity verification to ensure the consistency of parameter updates across multiple nodes in the distributed system.

[0046] The update method based on the preset model in this application obtains the preset model parameter gradient information generated by multiple computing nodes in multiple training batches, determines global gradient indicators based on the parameter gradient information of each training batch, and obtains multiple global gradient indicators. Then, it predicts the multiple global gradient indicators based on the preset prediction model to obtain the predicted values ​​and corresponding confidence levels of multiple future training batches. When the global gradient indicator of the current training batch is lower than a preset threshold, the predicted value is lower than the prediction threshold, and the confidence level is higher than a preset confidence threshold, parameter update information is generated based on the current parameter version number and the updated parameters are sent to the computing nodes. This achieves dynamic update decision-making to avoid resource waste, improves distributed training efficiency and resource utilization, and reduces the update risk during training instability periods.

[0047] The following combination Figure 3 The process of determining global gradient information is described, along with specific embodiments.

[0048] Figure 3 A flowchart illustrating the determination of global gradient information according to an embodiment of this application is shown schematically.

[0049] like Figure 3 As shown, the above operation S220 may include operations S301 to S304.

[0050] In operation S301, the gradient information of multiple parameters in any training batch is summed to obtain the summed gradient information of the parameters.

[0051] For example, the parameter gradient information for any training batch is as follows: =0.05、 =0.03、 =0.01, sum the gradient information of the above multiple parameters, i.e. 0.05+0.03+0.01=0.09, and obtain the summed parameter gradient information as 0.09. The gradient of each parameter in this training batch is determined by summing.

[0052] In operation S302, the square root of the summed parameter gradient information is taken to obtain the global gradient information of any training batch.

[0053] For example, when the summed gradient information is 0.09, performing a square root operation on the summed gradient information yields a global gradient information of 0.3 for this training batch. By performing a square root operation on the summed gradient information, the numerical range of the global gradient information becomes more reasonable, facilitating subsequent analysis of the overall gradient changes during model training.

[0054] In operation S303, multiple global gradient information for consecutive training batches are determined. Consecutive training batches include any training batch and multiple historical training batches preceding any training batch.

[0055] For example, the current training batch is the 4th batch. The consecutive training batches include the 4th batch and the previous 3rd, 2nd, and 1st batches. That is, the global gradient information of the 1st batch is 0.2, the global gradient information of the 2nd batch is 0.25, the global gradient information of the 3rd batch is 0.28, and the global gradient information of the 4th batch is 0.3, thus determining the multiple global gradient information of the consecutive training batches.

[0056] In operation S304, the global gradient metric for any training batch is determined based on multiple global gradient information.

[0057] The update method of the preset model based on the embodiments of this application integrates multiple parameter gradient information into global gradient information through operations such as summation and square root, and determines the global gradient index by combining continuous batch information, so as to more comprehensively and accurately reflect the changes in model parameter gradient.

[0058] The following combination Figure 4 The process of determining the global gradient index is described, along with specific implementation examples.

[0059] Figure 4 A flowchart illustrating the determination of a global gradient index according to an embodiment of this application is shown.

[0060] like Figure 4 As shown, the above operation S304 may include operations S401 to S403.

[0061] In operation S401, the standard deviation and root mean square value of multiple global gradient information are calculated respectively.

[0062] In operation S402, the volatility coefficient of the global gradient index is determined based on the standard deviation. The volatility coefficient is used to quantify the volatility of gradient information.

[0063] In operation S403, the global gradient index is determined based on the global gradient information, root mean square value, and fluctuation coefficient of any training batch.

[0064] For example, the standard deviation of multiple global gradient information is calculated as follows: The root mean square value of multiple global gradient information is The sum of the standard deviation and the constant is used as the constant. The fluctuation coefficient of the global gradient index is obtained by taking the logarithm with base 0. The global gradient index is determined according to the following formula. :

[0065]

[0066] The update method of the preset model based on the embodiments of this application determines the fluctuation coefficient by calculating the root mean square value and standard deviation of global gradient information, and determines the global gradient index by combining the global gradient information. By comprehensively considering factors such as root mean square value and standard deviation, the degree of gradient information fluctuation is quantified so that the determined global gradient index can more accurately reflect the model training status.

[0067] In this embodiment, multiple global gradient indicators from multiple training batches are combined in the order of training batches to obtain a temporal gradient vector; the temporal gradient vector is split to obtain multiple local temporal vectors; the multiple local temporal vectors are processed according to a preset temporal prediction model to obtain the predicted values ​​and corresponding confidence levels of the global gradient indicators for multiple future training batches.

[0068] For example, during the training of a pre-defined model, a global gradient metric representing the change in parameter gradients is calculated for each training batch. The global gradient metrics corresponding to multiple training batches are arranged and combined according to the order in which the training batches occur to construct a temporal gradient vector. This temporal gradient vector represents the dynamic changes in the global gradient metric as the pre-defined model progresses through training batches. The temporal gradient vector is then split into multiple local temporal vectors according to pre-defined rules. These local temporal vectors are processed according to a pre-defined temporal prediction model, and the resulting output is the predicted value of the global gradient metric for future training batches, along with the confidence level for each predicted value.

[0069] The update method of the preset model based on the embodiments of this application obtains the temporal gradient vector by sequentially combining global gradient indicators. The split temporal gradient vector is then used to predict the predicted values ​​and confidence levels of global gradient indicators for multiple future training batches. This allows for the discovery of data temporal patterns to accurately predict future global gradient indicators, laying the foundation for determining the parameter update time of the preset model.

[0070] In this embodiment, a state vector is obtained by extracting features from multiple local temporal vectors using a long short-term memory layer; the state vector is then processed using a fully connected layer to determine the predicted values ​​of the global gradient index for multiple future training batches and the multiple prediction distribution parameters corresponding to the multiple predicted values; and the confidence level is determined based on the variance of the multiple prediction distribution parameters.

[0071] For example, feature extraction is performed on multiple local time-series vectors based on the gating mechanism in the input long short-term memory layer. Specifically, information from multiple local time-series vectors at different time steps is filtered and integrated to determine the long-term dependencies between multiple local time-series vectors. Through these long-term dependencies, the state vectors of multiple local time-series vectors are determined.

[0072] After performing a nonlinear transformation on the state vector using a fully connected layer, the predicted values ​​of the global gradient index for multiple future training batches are determined by weight calculation and activation functions, resulting in multiple prediction distribution parameters corresponding to these predicted values. The confidence level is determined based on the variance of these multiple prediction distribution parameters; the smaller the variance, the more stable and reliable the predicted value, and the higher the corresponding confidence level.

[0073] The update method of the preset model based on the embodiments of this application makes full use of the temporal characteristics of the data through the collaborative processing of the long short-term memory layer and the fully connected layer, accurately obtains the predicted value and confidence level, provides solid and reliable data support for subsequent model parameter updates, and ensures the accuracy of the model update direction.

[0074] In this embodiment, the current parameter version number of the incremented preset model is used as the target parameter version number; parameter difference information is generated by comparing the parameter data corresponding to the current parameter version number with the parameter data corresponding to the target parameter version number; and the parameter difference information is used as parameter update information.

[0075] For example, during the iterative update of the preset model, when the preset update conditions are met, the current parameter version number of the preset model is incremented, and the incremented version number is used as the target parameter version number. A comparison algorithm is used to perform item-by-item comparison analysis between the parameter data corresponding to the current parameter version number and the parameter data corresponding to the target parameter version number. The differences between the parameter data corresponding to the current parameter version number and the parameter data corresponding to the target parameter version number are identified. These differences include changes in parameter values, adjustments to parameter structures, etc. The difference information is then summarized to generate parameter difference information, which is used as parameter update information to guide subsequent model parameter update operations.

[0076] The method for updating the preset model based on the embodiments of this application determines the parameter differences by accurately comparing the version numbers, which can clearly and comprehensively grasp the changes in model parameters, ensure that the parameter update process is accurate, avoid the degradation of model performance caused by parameter update errors, and ensure the stability and effectiveness of model training.

[0077] In operation S240, after generating parameter update information based on the current parameter version number of the preset model to be updated, the numerical range of multiple parameter values ​​is checked; if any parameter value exceeds the preset numerical range, the parameter value is reset to the preset initial value.

[0078] As an example, this method performs range validation on multiple parameter values. During the validation process, each parameter value is compared to a preset range based on a predefined threshold. When any parameter value is found to exceed the preset range, a parameter value reset mechanism is triggered. Specifically, the value exceeding the preset range is reset to a preset initial value to prevent abnormal parameters from spreading and affecting subsequent model training, thus causing bias or errors in model training.

[0079] The update method of the preset model based on the embodiments of this application can effectively ensure that the parameter values ​​are always within a reasonable range through strict numerical range verification and timely reset operation, avoid instability factors caused by abnormal parameters in the model training process, improve the reliability and robustness of model training, and ensure that the model can be optimized and updated in the right direction.

[0080] The following combination Figure 5 The process of updating parameter values ​​is described, along with specific embodiments.

[0081] Figure 5 A flowchart illustrating the updating of parameter values ​​according to an embodiment of this application is shown schematically.

[0082] like Figure 5 As shown, the above operation S240 may include operations S501 to S503.

[0083] In operation S501, the target parameter names in the preset model are traversed according to the parameter names in the mapping relationship.

[0084] In operation S502, if the target parameter name exists in the mapping relationship, the parameter value corresponding to the target parameter name is obtained from the parameter update information, and the parameter value is overwritten with the original parameter value.

[0085] In operation S503, if the target parameter name does not exist in the mapping relationship, the parameter value corresponding to the target parameter name is initialized.

[0086] In this embodiment, the parameter list of the preset model is traversed, and the target parameter name of each target parameter is read sequentially according to the model structure hierarchy. The target parameter name is matched and verified item by item with the mapping relationship contained in the parameter update information. The mapping relationship is stored in the form of key-value pairs to correspond to the parameter name and the parameter value.

[0087] When a match is successful, according to the key-value correspondence rules in the mapping relationship, the parameter value that completely corresponds to the parameter name is extracted from the parameter update information, and the parameter value is updated to the parameter storage area of ​​the preset model through direct memory write operation, overwriting its original value.

[0088] When a match fails, a preset initial value is assigned to the target parameter according to the parameter initialization strategy, and the initial value is written to the storage location corresponding to the target parameter in the parameter storage area through a memory write operation.

[0089] It should be noted that the traversal process uses a depth-first search algorithm to sequentially access the parameters of each level of the model, ensuring that all parameters are completely traversed. The parameter storage area is the model parameter storage space pre-allocated in the video memory. The parameter initialization strategy includes, but is not limited to, zero-value initialization, random initialization, or variance scaling initialization based on the model structure.

[0090] The update method for the preset model based on the embodiments of this application achieves targeted updates and intelligent initialization of parameters by traversing the model parameter names and matching them with mapping relationships. For existing parameters, a numerical overwrite method is used for precise updates. For newly added parameters, an initialization operation is automatically performed. This effectively avoids the resource overhead of updating all parameters, significantly improves the efficiency and accuracy of large-scale model updates, and ensures the integrity and training stability of the updated model.

[0091] In this embodiment of the application, the operation S502 described above, which involves obtaining the parameter value corresponding to the target parameter name from the parameter update information and overwriting the original value of the parameter with the parameter value, may include: if the target parameter name is not an updated target parameter name, stopping the overwriting of the original value of the parameter with the parameter value; and if the target parameter name is an updated target parameter name, obtaining the parameter value corresponding to the target parameter name from the parameter update information and overwriting the original value of the parameter with the parameter value.

[0092] For example, during the process of traversing the parameter list of the preset model, the following judgment and operation are performed on each target parameter name: when it is identified that the current target parameter name belongs to a non-updating target parameter name, the numerical update process for the current target parameter is terminated, all numerical transmission and writing operations are skipped, and the original value of the current target in the video memory storage area remains unchanged.

[0093] When it is identified that the current target parameter name belongs to the updated target parameter name, the corresponding parameter value is retrieved by using the target parameter name as the key value according to the mapping relationship between parameter name and value recorded in the parameter update information. The corresponding parameter value is then extracted and a direct memory access operation is initiated through the video memory controller to write the parameter value to the physical address unit corresponding to the parameter in the model parameter storage area, overwriting the original parameter value in the storage unit.

[0094] The update method of the preset model based on the embodiments of this application realizes differentiated update processing of model parameters by distinguishing between the names of updated and non-updated target parameters: for non-updated parameters, the original value remains unchanged, and for parameters that need to be updated, the corresponding value is extracted from the update information and overwritten with the original value, so as to realize precise update at the parameter level, avoid the computation and communication overhead caused by full model parameter refresh, reduce unnecessary video memory write operations, reduce hardware resource consumption and improve update efficiency.

[0095] In this embodiment, after updating the parameters of the preset model, multiple computing nodes are controlled to pause the training task after completing the current training batch; after pausing the training task, the current parameter information of the preset model is obtained and used as backup parameter data; the global gradient information of the first training batch of the updated preset model is determined; the temperature information of multiple computing nodes and the memory usage of the graphics processors in multiple computing nodes are obtained; if the global gradient information is greater than a preset gradient threshold, the memory usage exceeds a second preset usage threshold, or the temperature information exceeds a second preset temperature threshold, the backup parameter data is sent to multiple computing nodes to update the parameters of the preset model.

[0096] For example, multiple computing nodes can be coordinated to synchronously pause the training task after completing the current training batch, ensuring that the model state is momentarily frozen. Based on this consistent state, all current parameter information of the preset model can be completely extracted and backed up to generate a recoverable parameter snapshot, which is then used as backup parameter data.

[0097] After the global gradient information of the first training batch of the updated preset model is completed, the global gradient information of the first batch after restarting training is monitored in real time, and temperature sensor data and memory usage of each computing node are continuously collected.

[0098] If the global gradient magnitude exceeds the safety threshold, the memory usage rate rises abnormally, or the hardware temperature exceeds the operating limit, it is determined that there is a stability risk in this update. The pre-stored backup parameter data is distributed to all computing nodes and the existing parameters are overwritten, so that the preset model is restored to the stable state before the update.

[0099] It should be noted that the update method of the preset model based on the embodiments of this application can not only effectively avoid problems such as gradient anomalies, memory overflow or device overheating caused by parameter updates, but also minimize training interruption time and improve the availability and reliability of the distributed training system in the continuous evolution process.

[0100] In this embodiment, the backup parameter information includes multiple parameter names, multiple parameter values, and a mapping relationship between parameter names and parameter values. The optimizer state data is loaded asynchronously. Each parameter of the preset model is traversed. If there is parameter data in the backup parameter information that matches the current parameter name, the corresponding parameter value is copied to the graphics processor where the current parameter is located and its original value is overwritten. If there is no parameter data in the backup parameter information that matches the current parameter name, the parameter is initialized according to the preset initialization strategy. Each tensor in the optimizer state data stored in the backup parameter information is moved to the current graphics processor device to complete the loading and restoration of the optimizer state.

[0101] For example, the preset initialization strategy may include zero-value initialization, random initialization, or variance scaling initialization based on the model structure; the asynchronous loading process allows computation and data transmission to overlap to reduce training interruption time; optimizer state data includes, but is not limited to, training state information such as momentum and second-order moment estimation, and optimizer state data is of great significance for maintaining the convergence characteristics of the model.

[0102] The update method of the preset model based on the embodiments of this application performs parameter matching and restoration based on the mapping relationship in the backup parameter information, performs numerical overwriting on the matching parameters, initializes the unmatched parameters according to the strategy, and asynchronously completes the loading and restoration of the optimizer state, thereby achieving accurate restoration at the parameter level, avoiding the waste of computing resources caused by resetting all model parameters, and improving the efficiency of the restoration process and reducing training interruption time through asynchronous loading.

[0103] In this embodiment, before sending parameter update information to multiple computing nodes, the temperature information of multiple computing nodes and the video memory usage rate of the graphics processors in multiple computing nodes are obtained; if the temperature information is higher than a first preset temperature threshold or the video memory usage rate is higher than a first preset usage rate threshold, the sending of parameter update information to multiple computing nodes is stopped.

[0104] For example, before sending parameter update information to multiple computing nodes, the real-time readings of the internal temperature sensors and the current memory occupancy of each computing node are acquired in parallel; when the temperature data of any computing node is continuously higher than the first preset temperature threshold, or its memory occupancy exceeds the first preset occupancy threshold, the network distribution process of parameter update information is terminated and a hardware status alarm is sent.

[0105] The update method based on the preset model in the embodiments of this application effectively avoids a series of problems that may be caused by forcibly updating when the system resources are bottlenecked or the heat dissipation capacity is insufficient, by real-time perception and prediction of the thermal state and memory resources of computing nodes. For example, it prevents the graphics processor from throttling or permanent hardware damage caused by overheating, and eliminates numerical calculation errors caused by resource contention.

[0106] In this embodiment, the time interval between data reported by multiple computing nodes is monitored; if the time interval between data reported by any computing node exceeds a preset time threshold, the sending of parameter update information to multiple computing nodes is stopped.

[0107] For example, when the time interval between the arrival of heartbeat data packets of any computing node is detected to exceed a preset time threshold, any computing node that exceeds the preset time threshold is marked as a target computing node, the target computing node is marked as an abnormal state and an abnormal event log with a timestamp is generated, and the sending of parameter update information to multiple computing nodes is stopped.

[0108] It should be noted that after stopping the sending of parameter update information to multiple computing nodes, the abnormal status of the target computing node can be broadcast to the computing node cluster through a distributed consensus protocol to ensure that all control units are aware of the node abnormality of the target computing node and release the update bandwidth and computing resources allocated to the target computing node synchronously.

[0109] The update method of the preset model based on the embodiments of this application continuously monitors the time interval of data reported by computing nodes, and immediately triggers the circuit breaker mechanism when a heartbeat timeout is detected, that is, when the time interval exceeds the preset threshold, to stop distributing parameter update information to all nodes. By predictively interrupting the update, it avoids model state splitting and training process crashes caused by network partitions or node downtime, reduces invalid communication and computing resource consumption, and improves cluster resource utilization efficiency and system fault tolerance.

[0110] The following combination Figure 6 The component architecture of the method for updating the preset model is introduced, along with specific embodiments.

[0111] Figure 6 The schematic diagram illustrates the component architecture of a method for updating a preset model according to an embodiment of this application.

[0112] In the embodiments of this application, such as Figure 6 As shown, the component architecture mainly includes the following modules: status monitoring module, decision engine module, backup reference data module, distributed storage module, update execution module, storage service module, training cluster module, and circuit breaker rollback module. These modules work together to achieve an efficient and adaptive pre-defined model update process.

[0113] The training cluster module includes multiple computing nodes, which are used to execute the training tasks of the preset model and report the generated parameter gradient information after each training batch is completed.

[0114] The status monitoring module is used to collect and monitor the operational status data of multiple computing nodes in real time, including but not limited to the temperature information of the computing nodes, the memory usage of the graphics processor, and the time interval between data reports from each node. The status monitoring module transmits the monitoring data to the decision engine module and the circuit breaker rollback module in real time, providing a basis for updating decisions and system protection.

[0115] The distributed storage module is used to persistently store the massive amounts of data generated during training, including parameter gradient information reported by each computing node, model parameters of different versions, and backup reference data of global gradient index sequences.

[0116] Figure 7 The diagram illustrates the parameters of a preset model during training according to an embodiment of this application.

[0117] In the embodiments of this application, such as Figure 7 As shown, 701 represents the computation node, 702 represents the parameter gradient information of the second training batch of a single computation node in the training process of the preset model, 703 represents the fifth training batch, 704 represents the global gradient information of the second training batch, GCI1 represents the global gradient exponent of the third training batch, and similarly, GCI2 represents the global gradient exponent of the fourth training batch and GCI3 represents the global gradient exponent of the fifth training batch.

[0118] The decision engine module receives parameter gradient information from the training cluster, aggregates it by training batch, and calculates the global gradient metric. It then calls a built-in preset prediction model, such as a time-series prediction model, to predict the gradient metric for future batches. Combining the current metric, predicted values, confidence levels, and system status monitoring information, it determines whether the preset model update conditions are met. If the conditions are met, parameter update information is generated. The decision engine module also verifies parameter values ​​and performs a system status check before sending updates.

[0119] The backup reference data module is used to perform snapshot data backup of the current model's parameter version and related data before the decision engine triggers an update, and provides a backup data interface to the circuit breaker rollback module to ensure rapid recovery in the event of an update failure.

[0120] The update execution module is responsible for receiving parameter update information from the decision engine and distributing it to each computing node in the training cluster. Based on the mapping relationship in the parameter update information, the update execution module accurately traverses and updates the parameters of the preset model in the computing nodes, handling parameter overwriting and initialization operations.

[0121] The storage service module provides underlying data access services for the entire architecture, encapsulates read and write operations on distributed storage, and provides efficient and consistent data access interfaces for other modules.

[0122] The circuit breaker rollback module monitors the global gradient information and real-time status of computing nodes in the first training batch, such as temperature and memory usage. When an anomaly is detected, such as gradient explosion or hardware overload, a rollback operation is triggered. It automatically retrieves backup parameter data from the backup reference data module and calls the update execution module to revert the model parameters to the state before the update, ensuring system stability.

[0123] Through the close cooperation of the above modules, the component architecture provided in this application embodiment can realize an adaptive, safe, and efficient preset model update method, effectively improving the automation and reliability of the training process of large-scale machine learning systems.

[0124] Based on the above-mentioned method for updating the preset model, this application also provides a device for updating the preset model. The following will be combined with... Figure 8 The device is described in detail.

[0125] Figure 8 The diagram illustrates a structural block diagram of a preset model update device according to an embodiment of this application.

[0126] like Figure 8 As shown, the preset model update device 800 of this embodiment includes an acquisition module 810, an indicator determination module 820, a prediction module 830, an update determination module 840, and an update module 850.

[0127] The acquisition module 810 is used to acquire the parameter gradient information of a preset model generated by multiple computing nodes in multiple training batches. In one embodiment, the acquisition module 810 can be used to perform the operation S210 described above, which will not be repeated here.

[0128] The indicator determination module 820 is used to determine the global gradient indicator for any training batch based on the gradient information of multiple parameters in any training batch, thereby obtaining multiple global gradient indicators for multiple training batches. The global gradient indicator is used to characterize the degree of change of the parameter gradient of the preset model during training. In one embodiment, the indicator determination module 820 can be used to perform the operation S220 described above, which will not be repeated here.

[0129] The prediction module 830 is used to predict the multiple global gradient indicators according to a preset prediction model, so as to obtain the predicted values ​​and corresponding confidence levels of the global gradient indicators for multiple future training batches. In one embodiment, the prediction module 830 can be used to perform the operation S230 described above, which will not be repeated here.

[0130] The update determination module 840 is used to generate parameter update information based on the current parameter version number of the preset model to be updated when the global gradient index of the current training batch is lower than a preset index threshold, the multiple predicted values ​​are lower than a preset prediction threshold, and the multiple confidence scores are higher than a preset confidence threshold. In one embodiment, the update determination module 840 can be used to perform the operation S240 described above, which will not be repeated here.

[0131] The update module 850 is used to send the parameter update information to the plurality of computing nodes to update the parameters of the preset model. In one embodiment, the update module 850 can be used to perform the operation S250 described above, which will not be repeated here.

[0132] According to embodiments of this application, any multiple modules among the acquisition module 810, indicator determination module 820, prediction module 830, update determination module 840, and update module 850 can be combined into one module, or any one of these modules can be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules can be combined with at least part of the functionality of other modules and implemented in one module. According to embodiments of this application, at least one of the acquisition module 810, indicator determination module 820, prediction module 830, update determination module 840, and update module 850 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, at least one of the acquisition module 810, the indicator determination module 820, the prediction module 830, the update determination module 840, and the update module 850 may be implemented at least partially as a computer program module, which can perform corresponding functions when the computer program module is run.

[0133] Figure 9 A block diagram schematically illustrates an electronic device suitable for implementing an update method for a preset model according to an embodiment of this application.

[0134] like Figure 9As shown, an electronic device 900 according to an embodiment of this application includes a processor 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage portion 908 into a random access memory (RAM) 903. The processor 901 may include, for example, a general-purpose microprocessor (e.g., a central processing unit), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 901 may also include onboard memory for caching purposes. The processor 901 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of this application.

[0135] RAM 903 stores various programs and data required for the operation of electronic device 900. Processor 901, ROM 902, and RAM 903 are interconnected via bus 904. Processor 901 executes various operations of the method flow according to embodiments of this application by executing programs in ROM 902 and / or RAM 903. It should be noted that programs may also be stored in one or more memories other than ROM 902 and RAM 903. Processor 901 may also execute various operations of the method flow according to embodiments of this application by executing programs stored in one or more memories.

[0136] According to embodiments of this application, the electronic device 900 may further include an input / output (I / O) interface 905, which is also connected to a bus 904. The electronic device 900 may also include one or more of the following components connected to the input / output (I / O) interface 905: an input section 906 including a keyboard, mouse, etc.; an output section 907 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a LAN card, modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the input / output (I / O) interface 905 as needed. A removable medium 911, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 910 as needed so that computer programs read from it can be installed into the storage section 908 as needed.

[0137] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of this application.

[0138] According to embodiments of this application, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this application, the computer-readable storage medium may include ROM 902 and / or RAM 903 and / or one or more memories other than ROM 902 and RAM 903 described above.

[0139] Embodiments of this application also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code is used to enable the computer system to implement the update method of the preset model provided in the embodiments of this application.

[0140] When the computer program is executed by the processor 901, it performs the functions defined in the system / apparatus of this application embodiment. According to the embodiments of this application, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0141] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and downloaded and installed via the communication section 909, and / or installed from a removable medium 911. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.

[0142] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 909, and / or installed from the removable medium 911. When the computer program is executed by the processor 901, it performs the functions defined in the system of this application embodiment. According to the embodiments of this application, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0143] According to embodiments of this application, program code for executing the computer programs provided in the embodiments of this application can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C", or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0144] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0145] Those skilled in the art will understand that the features described in the various embodiments of this application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, the features described in the various embodiments of this application can be combined and / or combined in various ways without departing from the spirit and teachings of this application. All such combinations and / or combinations fall within the scope of this application.

[0146] The embodiments of this application have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of this application. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Without departing from the scope of this application, those skilled in the art can make various substitutions and modifications, all of which should fall within the scope of this application.

Claims

1. A method for updating a preset model, characterized in that, The method includes: Obtain the parameter gradient information of a preset model generated by multiple computing nodes in multiple training batches; For multiple parameter gradient information of any training batch, determine the global gradient index of the training batch, and obtain multiple global gradient indices for multiple training batches. The global gradient index is used to characterize the degree of change of parameter gradient of the preset model during training. The global gradient metrics from multiple training batches are combined in the order of the training batches to obtain the temporal gradient vector. The temporal gradient vector is split into multiple local temporal vectors; Based on the long short-term memory layer, feature extraction is performed on the multiple local temporal vectors to obtain the state vector; The state vector is processed by the fully connected layer to determine the predicted values ​​of the global gradient index for multiple future training batches and multiple predicted distribution parameters corresponding to the multiple predicted values. Based on the variance of the multiple predicted distribution parameters, multiple confidence levels are determined; If the global gradient index of the current training batch is lower than the preset index threshold, the multiple predicted values ​​are lower than the preset prediction threshold, and the multiple confidence scores are higher than the preset confidence threshold, parameter update information is generated according to the current parameter version number of the preset model to be updated. The parameter update information is sent to the multiple computing nodes to update the parameters of the preset model.

2. The method according to claim 1, characterized in that, The determination of the global gradient index for any training batch, based on the gradient information of multiple parameters for any training batch, includes: Summing the gradient information of multiple parameters for any training batch yields the summed gradient information of the parameters. The square root of the summed parameter gradient information is used to obtain the global gradient information for any training batch. Determine multiple global gradient information for consecutive training batches, wherein the consecutive training batches include any training batch and multiple historical training batches preceding any training batch; Calculate the standard deviation and root mean square value of the multiple global gradient information respectively; Based on the standard deviation, the fluctuation coefficient of the global gradient index is determined, and the fluctuation coefficient is used to characterize the degree of fluctuation of the gradient information; The global gradient index is determined based on the global gradient information, root mean square value, and fluctuation coefficient of any training batch. The global gradient index is positively correlated with the global gradient information and root mean square value, and negatively correlated with the standard deviation.

3. The method according to claim 1, characterized in that, Based on the current parameter version number of the preset model, parameter update information is generated, including: The current parameter version number of the preset model after incrementing is used as the target parameter version number; Parameter difference information is generated by comparing the parameter data corresponding to the current parameter version number with the parameter data corresponding to the target parameter version number. The parameter difference information is used as the parameter update information.

4. The method according to claim 1, characterized in that, The parameter update information includes multiple parameter values; after generating the parameter update information, the method further includes: The numerical range of the multiple parameters is verified; If any of the multiple parameter values ​​exceeds the preset value range, the value of that parameter will be reset to the preset initial value.

5. The method according to claim 1, characterized in that, The parameter update information includes multiple parameter names, multiple parameter values, and multiple mapping relationships composed of parameter names and parameter values; Sending the parameter update information to the plurality of computing nodes to update the parameters of the preset model includes: Based on the parameter names in the mapping relationship, the target parameter names in the preset model are traversed; If the target parameter name exists in the mapping relationship, obtain the parameter value corresponding to the target parameter name from the parameter update information, and overwrite the original value of the parameter with the parameter value; If the target parameter name does not exist in the mapping relationship, the parameter value corresponding to the target parameter name is initialized.

6. The method according to claim 5, characterized in that, The target parameter name includes the updated target parameter name and the non-updated target parameter name; Obtaining the parameter value corresponding to the target parameter name from the parameter update information, and overwriting the original value of the parameter with the parameter value includes: If the target parameter name is not an updated target parameter name, stop overwriting the original value of the parameter with the parameter value; If the target parameter name is an updated target parameter name, the parameter value corresponding to the target parameter name is obtained from the parameter update information, and the parameter value is used to overwrite the original value of the parameter.

7. The method according to claim 1, characterized in that, After updating the parameters of the preset model, the method further includes: Control the multiple computing nodes to pause the training task after completing the current training batch; After pausing the training task, obtain the current parameter information of the preset model and use the current parameter information as backup parameter data; Determine the global gradient information for the first training batch of the updated preset model; Obtain the temperature information of the multiple computing nodes and the video memory usage of the graphics processors in the multiple computing nodes; Under the condition that the global gradient information is greater than a preset gradient threshold, the memory usage exceeds a second preset usage threshold, or the temperature information exceeds a second preset temperature threshold, the backup parameter data is sent to the multiple computing nodes to update the parameters of the preset model.

8. The method according to claim 1, characterized in that, Before sending the parameter update information to the plurality of computing nodes, the method further includes: Monitor the status indicators of the plurality of computing nodes, the status indicators including at least one of the following: temperature information of the plurality of computing nodes, video memory utilization rate of the graphics processor in the plurality of computing nodes, and time interval of data reporting by the plurality of computing nodes; If the temperature information is higher than a first preset temperature threshold, the time interval between data reported by any computing node exceeds a preset time threshold, or the memory usage rate is higher than a first preset usage rate threshold, the transmission of the parameter update information to the plurality of computing nodes shall be stopped.

9. An electronic device, comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Deep learning model training method and device, equipment and storage medium

    CN118627600A

  • Distributed training method and device, equipment and storage medium

    CN120654849A