Detection Method and Device for Distributed Training Model, Storage Medium and Electronic Device
By evaluating the parameter differences between the reference model and the current model, the target error value and parameter benchmark error value are generated, which solves the problem of poor evaluation of distributed training results, and realizes consistency evaluation of model parameters and training strategy optimization.
Patent Information
- Application Number
- CN202510191309.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-02-20
AI Technical Summary
In the prior art, the evaluation effect of distributed training results is poor, resulting in low consistency of training results of each node and the inability to adjust the training process in time.
By obtaining the degree of parameter differences between the reference model and the current model, the target error value and parameter benchmark error value are generated, the parameter consistency of the model in a distributed training environment is evaluated, and the hyperparameters are adjusted according to the evaluation results to optimize the training process.
The consistency evaluation of model parameters in a distributed training environment is realized, the accuracy and consistency of training results are improved, and the training strategy is optimized.
Smart Images

Figure CN119669715B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of computers, and more particularly, to a method and apparatus for detecting a distributed training model, a storage medium, and an electronic device. Background Art
[0002] With the rapid development of artificial intelligence technology, the technology of distributed training of a model through multiple nodes has been widely applied.
[0003] The current evaluation of the distributed training result of a model is carried out through the accuracy of the trained model on a validation set. However, in a distributed training scenario, there are certain differences between the outputs of different nodes, and the output result of the model on the validation set cannot quantify the above differences. Therefore, when the output differences of different nodes are large, the training process cannot be adjusted in time, resulting in low consistency of the training results of each node, and the aggregated training result deviates from the expected value. That is to say, there is a technical problem in the prior art that the evaluation effect of the distributed training result is not good.
[0004] For the above problems, no effective solution has been proposed yet. Summary of the Invention
[0005] Embodiments of the present application provide a method and apparatus for detecting a distributed training model, a storage medium, and an electronic device, so as to at least solve the technical problem that the evaluation effect of the distributed training result in the related art is not good.
[0006] According to one aspect of the embodiments of the present application, a method for detecting a distributed training model is provided, including:
[0007] Obtaining at least one reference model parameter when a reference model reaches a training convergence state, and at least one current parameter obtained after a current model performs a distributed node training operation, where the structural similarity between the model structure of the reference model and the model structure of the current model is greater than a target threshold, and the distributed node training operation is a training operation based on multiple distributed nodes with a shared memory pool;
[0008] Generating at least one target error value between at least one current parameter and the corresponding reference model parameter, where the target error value is used to indicate the degree of difference between the current parameter and the corresponding reference model parameter;
[0009] Generating a training detection result of the current model according to at least one target error value and a parameter reference error value corresponding to the reference model, where the parameter reference error value is used to indicate the degree of difference allowed between the reference model parameter and the current parameter.
[0010] Optionally, a training detection result of the current model is generated according to at least one target error value and a parameter reference error value corresponding to a reference model, including: obtaining a product between a scaling reference value for adjusting model parameters and the reference model parameters; determining the sum of the product and the error reference value as the parameter reference error value; when the target error value generated based on the current parameters is greater than the parameter reference error value corresponding to the current parameters, generating a training detection result based on the ratio of the number of target model parameters to the number of current parameters, or generating a training detection result based on the number of target model parameters.
[0011] Optionally, after generating a training detection result based on the ratio of the number of target model parameters to the number of current parameters, the method further includes: when the ratio of the number of target model parameters to the number of current parameters is greater than or equal to a first target adjustment threshold, adjusting the current parameters of the current model; or, after generating a training detection result based on the number of target model parameters, the method further includes: when the number of target model parameters is greater than or equal to a second target adjustment threshold, adjusting the current parameters of the current model.
[0012] Optionally, obtaining at least one reference model parameter when the reference model reaches a training convergence state, and at least one current parameter obtained after the current model performs a distributed node training operation, including: obtaining a model structure corresponding to the current model and a model structure corresponding to a candidate model, where the candidate model has the same executable task type as the current model; determining a structural similarity between the current model and the candidate model according to the model structure corresponding to the current model and the model structure corresponding to the candidate model; when the structural similarity is greater than a target threshold, determining the candidate model as the reference model; when the reference model reaches training convergence, obtaining at least one reference model parameter of the reference model; after the current model performs a distributed node training operation, obtaining at least one current parameter of the current model.
[0013] Optionally, generating at least one target error value between at least one current parameter and the corresponding reference model parameter, including: obtaining the respective differences between at least one current parameter and the corresponding reference model parameter; determining the absolute value of at least one difference as at least one target error value.
[0014] Optionally, when the training detection result indicates adjusting the current parameters, adjusting at least one hyperparameter matching the current model, where the hyperparameter is used to indicate the training method of the current model.
[0015] Optionally, in the process of generating the training detection result of the current model according to the target error value and the parameter reference error value corresponding to the reference model, it further includes at least one of the following: obtaining the data transmission delay of the current model performing the training operation based on the training sample set; obtaining the first training duration of the current model performing the training operation based on the training sample set, and the second training duration of the reference model performing the synchronous training operation based on the training sample set; obtaining the consistency synchronization description parameter of the current model performing the training operation based on the training sample set; obtaining the first sample prediction accuracy of the current model based on the training sample set, and obtaining the second sample prediction accuracy of the current model based on the validation sample set.
[0016] Optionally, the above detection method for a distributed training model further includes at least one of the following: adjusting at least one hyperparameter matching the current model when the data transmission delay meets the first sub-goal adjustment condition; adjusting at least one hyperparameter matching the current model when the first training duration and the second training duration meet the second sub-goal adjustment condition; adjusting at least one hyperparameter matching the current model when the consistency synchronization description parameter meets the third sub-goal adjustment condition; adjusting at least one hyperparameter matching the current model when the first sample prediction accuracy and the second sample prediction accuracy meet the fourth sub-goal adjustment condition.
[0017] Optionally, adjusting at least one hyperparameter matching the current model includes at least one of the following: adjusting the structure parameter for indicating the model structure of the current model; adjusting the node description parameter for indicating the number of nodes of the distributed node; adjusting the memory description parameter for indicating the memory size of the shared memory pool; adjusting the process description parameter for indicating the distribution of training processes running in multiple distributed nodes.
[0018] Optionally, adjusting at least one hyperparameter matching the current model further includes: obtaining the hyperparameter sequences corresponding to each of the multiple current models in the training state, where the hyperparameter sequence includes at least one hyperparameter, and the hyperparameter sequence is used to indicate the training methods corresponding to each of the multiple current models in the training state; determining at least one target hyperparameter sequence from the multiple hyperparameter sequences, where the at least one target hyperparameter sequence is determined based on the sorting result of the multiple hyperparameter sequences according to the target fitness, and the target fitness is the proportion of the target parameter in at least one current parameter of each of the multiple current models in the training state; adjusting at least one hyperparameter matching the current model according to the at least one target hyperparameter sequence.
[0019] Optionally, after generating the target error value between the reference model parameters and the current parameters, it includes: obtaining a plurality of target error values; determining the target error values with numerical values greater than the first target threshold as the first reference target error values, and determining the target error values with numerical values less than the second target threshold as the second reference target error values; adjusting the error reference value and / or the scaling reference value according to the quantities of the first reference target error values and the second reference target error values.
[0020] According to another aspect of the embodiments of the present application, there is provided a detection device for a distributed training model, including:
[0021] An obtaining unit, configured to obtain at least one reference model parameter when the reference model reaches the training convergence state, and at least one current parameter obtained after the current model performs the distributed node training operation, where the structural similarity between the model structure of the reference model and the model structure of the current model is greater than the target threshold, and the distributed node training operation is a training operation based on a plurality of distributed nodes with a shared memory pool;
[0022] A generating unit, configured to generate at least one target error value between at least one current parameter and the corresponding reference model parameter, where the target error value is used to indicate the degree of difference between the current parameter and the corresponding reference model parameter;
[0023] A detecting unit, configured to generate a training detection result of the current model according to at least one target error value and a parameter benchmark error value corresponding to the reference model, where the parameter benchmark error value is used to indicate the degree of difference allowed between the reference model parameter and the current parameter.
[0024] According to still another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium, in which a computer program is stored, where the computer program is configured to execute the above-mentioned evaluation method of the distributed training result when running.
[0025] According to still another aspect of the embodiments of the present application, there is provided a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the evaluation method of the distributed training result as above.
[0026] According to still another aspect of the embodiments of the present application, there is also provided an electronic device, including a memory and a processor, where a computer program is stored in the memory, and the processor is configured to execute the above-mentioned evaluation method of the distributed training result through the computer program.
[0027] Through the above embodiments of the present application, at least one reference model parameter when the reference model reaches the training convergence state is first obtained, as well as at least one current parameter obtained after the current model performs the distributed node training operation, wherein the structural similarity between the model structure of the reference model and the model structure of the current model is greater than the target threshold, and the distributed node training operation is a training operation based on multiple distributed nodes with a shared memory pool; further, at least one target error value between at least one current parameter and the corresponding reference model parameter is generated, where the target error value is used to indicate the degree of difference between the current parameter and the corresponding reference model parameter; and a training detection result of the current model is generated according to at least one target error value and the parameter benchmark error value corresponding to the reference model. Thus, the consistency evaluation of each model parameter generated by the model in the distributed training environment is realized, thereby solving the technical problem in the prior art that the evaluation effect of the distributed training result is not good. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:
[0029] Figure 1 is a hardware structure block diagram of a server device for a method of detecting a distributed training model according to an embodiment of the present application;
[0030] Figure 2 is a flowchart of a method of detecting a distributed training model according to an embodiment of the present application;
[0031] Figure 3 is a flowchart of another method of detecting a distributed training model according to an embodiment of the present application;
[0032] Figure 4 is a flowchart of yet another method of detecting a distributed training model according to an embodiment of the present application;
[0033] Figure 5 is a flowchart of yet another method of detecting a distributed training model according to an embodiment of the present application;
[0034] Figure 6 is a flowchart of yet another method of detecting a distributed training model according to an embodiment of the present application;
[0035] Figure 7 is a schematic structural diagram of a device for detecting a distributed training model according to an embodiment of the present application;
[0036] Figure 8 is a schematic structural diagram of an electronic device for detecting a distributed training model according to an embodiment of the present application. Detailed implementation manners
[0037] In the following, embodiments of the present application will be described in detail with reference to the accompanying drawings and in conjunction with the embodiments.
[0038] It should be noted that the terms "first", "second", etc. in the specification, claims and the above-mentioned drawings of the present application are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence.
[0039] The method embodiments provided in the embodiments of the present application can be executed on a server device or a similar computing device. Taking the operation on a server device as an example, Figure 1 is a hardware structural block diagram of a server device for an evaluation method of distributed training results in an embodiment of the present application. As Figure 1 shown, the server device may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above-mentioned server device may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned server device. For example, the server device may further include more or fewer components than those shown in Figure 1 the figure, or have a different configuration from that shown in Figure 1 the figure.
[0040] The memory 104 can be used to store computer programs. For example, software programs and modules of application software, such as the computer program corresponding to the data processing method of the memory in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, the above-mentioned method is implemented. The memory 104 may include a high-speed random access memory, and may further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely provided with respect to the processor 102, and these remote memories can be connected to the server device through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0041] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of a server device. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 may be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0042] As an alternative implementation, as Figure 2 shown, the above-mentioned method for evaluating distributed training results includes:
[0043] S202, obtaining at least one reference model parameter when the reference model reaches the training convergence state, and at least one current parameter obtained after the current model performs the distributed node training operation, where the structural similarity between the model structure of the reference model and the model structure of the current model is greater than the target threshold, and the distributed node training operation is a training operation based on multiple distributed nodes with a shared memory pool;
[0044] S204, generating at least one target error value between at least one current parameter and the corresponding reference model parameter, where the target error value is used to indicate the degree of difference between the current parameter and the corresponding reference model parameter;
[0045] S206, generating a training detection result of the current model according to at least one target error value and a parameter reference error value corresponding to the reference model, where the parameter reference error value is used to indicate the degree of difference allowed between the reference model parameter and the current parameter.
[0046] It should be noted that in the above step S202, the above-mentioned structural similarity is a quantification of the structural similarity degree between the current model and the reference model. For example, when the model structures of the current model and the reference model are exactly the same, the above-mentioned structural similarity is 1. Another example is that when there are differences in the model structures of the current model and the reference model, the above-mentioned structural similarity is determined by the respective structural parameters of the current model and the reference model, such as the number of layers, the number of neurons in each layer, etc. Optionally, assuming that the number of network layers of the current model and the reference model with the same network layer is x layers, and the number of network layers with different network layers is y layers, then the above-mentioned structural similarity is x / (x + y).
[0047] It should be noted that the above current model is an artificial neural network model trained in a distributed environment with a shared memory pool, and the reference model is a converged artificial neural network model. The reference model can have a training environment different from that of the current model. For example, the above reference model can be trained through multiple distributed nodes without a shared memory pool, or can be trained based on a single node.
[0048] It should be noted that the above shared memory pool includes the memory in each distributed node, the video memory in the GPU of each distributed node, etc. Each distributed node can directly access the data in the shared memory pool.
[0049] The reference model can also have the same training environment as the current model, that is, the reference model can be trained through multiple distributed nodes with a shared memory pool. For example, the reference model is a model with high accuracy trained in distributed nodes with a shared memory pool. To improve the training and inference efficiency, a structured pruning operation (removing some convolutional kernels) is performed on the reference model, and the structural similarity between the pruned reference model and the reference model before pruning is greater than the target threshold. The pruned reference model is continued to be trained in a distributed training scenario with a shared memory pool to obtain the current model.
[0050] To evaluate the training effect of the current model, at least one reference model parameter when the reference model reaches the training convergence state, and at least one current parameter obtained after the current model performs the distributed node training operation are obtained. Further, the reference model is a model with high accuracy, and the current model is a model whose training effect is to be evaluated. The training effect of the current model in a distributed scenario with a shared memory pool is evaluated by comparing the degree of difference in model parameters between the current model and the reference model. That is, the role of the above reference model is to provide a benchmark. By comparing the parameters with the current model, the parameter deviation in the training process of the current model can be quantified, and then the current training effect can be optimized by modifying hyperparameters and other methods.
[0051] Optionally, generating at least one target error value between at least one current parameter and the corresponding reference model parameter includes: obtaining the respective differences between at least one current parameter and the corresponding reference model parameter; determining the absolute value of at least one difference as at least one target error value. That is, 。
[0052] Optionally, when there are some current parameters in the current model that do not have corresponding reference model parameters due to the inconsistent model structures of the reference model and the current model, the value of 0 can be regarded as the parameter value of the corresponding reference model parameter. That is, for the current parameter without a corresponding reference model parameter, the corresponding target error value is the absolute value of itself, that is, 。
[0053] Further, perform the above step S206 to generate a training detection result of the current model according to the target error value and the parameter reference error value corresponding to the reference model. It should be noted that the above parameter reference error value is used to indicate the allowable difference degree between the reference model parameters and the current parameters. The above parameter reference error value can be a pre-specified value or can be dynamically determined according to the reference model parameters.
[0054] Optionally, the process of dynamically determining the parameter reference error value according to the reference model parameters is as follows: obtain the product between the scaling reference value for adjusting the model parameters and the reference model parameters; determine the sum of the product and the error reference value as the parameter reference error value.
[0055] It can be understood that the calculation formula of the above parameter reference error value is as follows:
[0056]
[0057] It should be noted that the above scaling reference value is a pre-specified value, such as 10 -3 、10 -4 etc., and the above error reference value is a pre-specified offset amount, and the value of the error reference value can be set to 10 -5 etc., and no specific limitation is made here.
[0058] Optionally, generating the training detection result of the current model according to at least one target error value and the parameter reference error value corresponding to the reference model further includes: in the case where the target error value generated based on the current parameters is greater than the parameter reference error value corresponding to the current parameters, determining the current parameters as the target model parameters; generating a training detection result based on the ratio of the number of target model parameters to the number of current parameters, or generating a training detection result based on the number of target model parameters.
[0059] Through the above embodiments of the present application, at least one reference model parameter when the reference model reaches the training convergence state is first obtained, as well as at least one current parameter obtained after the current model performs the distributed node training operation. Among them, the structural similarity between the model structure of the reference model and the model structure of the current model is greater than the target threshold, and the distributed node training operation is a training operation based on multiple distributed nodes with a shared memory pool; further, a target error value between at least one current parameter and the corresponding reference model parameter is generated, where the target error value is used to indicate the degree of difference between the current parameter and the corresponding reference model parameter; and a training detection result of the current model is generated according to the target error value and the parameter reference error value corresponding to the reference model. Thus, the consistency evaluation of each model parameter generated by the model in the distributed training environment is realized, thereby solving the technical problem in the prior art that the evaluation effect of the distributed training result is not good.
[0060] Optionally, after generating the training detection result based on the ratio of the number of target model parameters to the number of current parameters, it further includes: when the ratio of the number of target model parameters to the number of current parameters is greater than or equal to the first target adjustment threshold, adjusting the current parameters of the current model.
[0061] Optionally, after generating the training detection result based on the number of target model parameters, it further includes: when the number of target model parameters is greater than or equal to the second target adjustment threshold, adjusting the current parameters of the current model.
[0062] It can be understood that when the target error value between the current parameter in the current model and the corresponding reference model parameter in the reference model is less than or equal to the parameter reference error value, it indicates that the current parameter has a high consistency with the corresponding reference model parameter. The more current parameters with high consistency, the better the training effect of the current model. Similarly, when the target error value between the current parameter in the current model and the corresponding reference model parameter in the reference model is greater than the parameter reference error value, it indicates that the current parameter has a low consistency with the corresponding reference model parameter. The more current parameters (target model parameters) with low consistency, the worse the training effect of the current model. Further, when the ratio of the number of current parameters (target model parameters) with low consistency to all current parameters in the current model is greater than or equal to the preset first target adjustment threshold, or when the number of target model parameters is greater than or equal to the preset second target threshold, it is necessary to adjust the current parameters in the current model to make the current model have a good inference accuracy.
[0063] Optionally, as Figure 3As shown, obtain at least one reference model parameter when the reference model reaches the training convergence state, and at least one current parameter obtained after the current model performs the distributed node training operation, including:
[0064] S302, obtain the model structure corresponding to the current model and the model structure corresponding to the candidate model, where the candidate model has the same type of executable task as the current model;
[0065] S304, determine the structural similarity between the current model and the candidate model according to the model structure corresponding to the current model and the model structure corresponding to the candidate model;
[0066] S306, when the structural similarity is greater than the target threshold, determine the candidate model as the reference model;
[0067] S308, when the reference model reaches training convergence, obtain at least one reference model parameter of the reference model;
[0068] S310, after the current model performs the distributed node training operation, obtain at least one current parameter of the current model.
[0069] It should be noted that the above-mentioned candidate model having the same type of executable task as the current model means that the current model and the candidate model jointly have one or more of the same types of executable tasks. For example, if the current model can be used to perform the image-text matching task, then the candidate model can also be used to perform the image-text matching task; another example is that if the current model can perform the image-text matching task and the dialogue task, then the candidate model can perform at least one of the image-text matching task or the dialogue task.
[0070] Furthermore, obtain the model structures corresponding to the current model and the candidate model respectively, and determine the structural similarity between the current model and the candidate model according to the model structure corresponding to the current model and the model structure corresponding to the candidate model. Specifically, determine the structural parameters corresponding to the current model and the candidate model respectively according to the model structures corresponding to the current model and the candidate model, such as the number of layers, the number of neurons in each layer, etc. Optionally, assume that the number of network layers of the current model and the candidate model is the same as x layers, and the number of different network layers is y layers, then the structural similarity between the current model and the candidate model is .
[0071] Furthermore, when the structural similarity between the current model and the candidate model is greater than the preset target threshold, determine the candidate model as the reference model. Thus, when the reference model reaches training convergence, obtain at least one reference model parameter of the reference model; after the current model performs the distributed node training operation, obtain at least one current parameter of the current model.
[0072] As an alternative implementation, when the training detection result indicates an adjustment to the current parameters, the method further includes:
[0073] S1. Adjust at least one hyperparameter that matches the current model, where the hyperparameter is used to indicate the training method of the current model.
[0074] It can be understood that the above hyperparameters include the learning rate, batch size, optimizer type, etc., and are not specifically limited here.
[0075] It should be noted that the above training detection result is determined based on the target error value and the parameter reference error value. Therefore, when the ratio of the number of current parameters for which the corresponding target error value is greater than the parameter reference error value to the number of all current parameters is greater than the preset threshold, it indicates that the training effect of the current model is not good and the training strategy needs to be optimized. That is, it is determined that the training detection result indicates an adjustment to the current parameters. Furthermore, the current model is retrained or continued to be trained. Before retraining or continuing to train the current model, the above step S1 can be executed to modify the hyperparameters of the current model to adjust the training process of the current model, so that the current model after retraining or continued training has a good training effect.
[0076] Through the above implementation, when the training detection result indicates an adjustment to the current parameters, at least one hyperparameter that matches the current model is adjusted. Furthermore, when the training effect of the current model is not good in a distributed scenario based on a shared memory pool, the training strategy can be modified in a timely manner.
[0077] As an alternative implementation, in the process of generating the training detection result of the current model according to at least one target error value and the parameter reference error value corresponding to the reference model, the method further includes at least one of the following:
[0078] S1. Obtain the data transmission delay of the current model performing a training operation based on the training sample set;
[0079] S2. Obtain the first training duration of the current model performing a training operation based on the training sample set, and the second training duration of the reference model performing a synchronous training operation based on the training sample set;
[0080] S3. Obtain the consistency synchronization description parameter of the current model performing a training operation based on the training sample set;
[0081] S4. Obtain the first sample prediction accuracy of the current model based on the training sample set, and obtain the second sample prediction accuracy of the current model based on the validation sample set.
[0082] It should be noted that the above method of evaluating the parameter consistency of the current model by comparing the target error value with the parameter reference error value is one aspect of evaluating the distributed training effect of the current model based on a shared memory pool. It is also possible to more comprehensively evaluate the distributed training effect of the current model based on a shared memory pool through the data transmission delay, the first training duration, the second training duration, the consistency synchronization description parameter, the first sample prediction accuracy, and the second sample prediction accuracy obtained in the above steps S1 to S4.
[0083] Optionally, after obtaining the data transmission delay, the first training duration, the second training duration, the consistency synchronization description parameter, the first sample prediction accuracy, and the second sample prediction accuracy, the following steps S1-1 to S4-1 can be executed:
[0084] S1-1, when the data transmission delay meets the first sub-goal adjustment condition, adjust at least one hyperparameter matching the current model;
[0085] S2-1, when the first training duration and the second training duration meet the second sub-goal adjustment condition, adjust at least one hyperparameter matching the current model;
[0086] S3-1, when the consistency synchronization description parameter meets the third sub-goal adjustment condition, adjust at least one hyperparameter matching the current model;
[0087] S4-1, when the first sample prediction accuracy and the second sample prediction accuracy meet the fourth sub-goal adjustment condition, adjust at least one hyperparameter matching the current model.
[0088] For the above step S1-1, record the data transmission delay of the current model performing the training operation based on the training sample set, that is, the time required for data to be exchanged between multiple nodes through the shared memory pool. If the data transmission delay meets the first sub-goal adjustment condition, that is, the above data transmission delay exceeds the preset threshold or has a significant deviation from the expected delay, it indicates that the current training configuration (such as batch size, network communication performance, etc.) affects the data transmission efficiency.
[0089] For example, when Node A has completed parameter updates for some of the parameters in the current model, Node A can directly use the updated parameters for continued training. Other nodes, however, can only use the updated parameters for continued training when they access the shared memory pool and Node A has completed synchronizing the updated parameters to the shared memory pool. Otherwise, they will continue training with the parameters before Node A's update. Therefore, the lower the above-mentioned data transmission delay, the more timely each node updates the parameters output by other nodes, and the better the training effect of the current model. Conversely, if the data transmission delay is higher, each node updates the parameters output by other nodes more slowly, and the training effect of the current model is worse.
[0090] When the data transmission delay is greater than a preset delay threshold, the data transmission delay at this time meets the first sub-goal adjustment condition. At this time, at least one hyperparameter matching the current model is adjusted to optimize the data transmission efficiency. For example, reduce the batch size to reduce the data transmission delay because processing smaller data blocks requires less communication time. Another example is to adjust the network communication parameters (such as increasing the communication power and channel bandwidth) to improve the speed and stability of data transmission.
[0091] For the above step S2-1, record the training durations of the current model and the reference model when performing training operations based on the same training sample set, which are the first training duration and the second training duration respectively. If the first training duration and the second training duration meet the second sub-goal adjustment condition, that is, the difference between the two exceeds the set threshold, it means that the training efficiency of the current model is lower than expected. At this time, at least one hyperparameter is adjusted to increase the training speed of the current model and improve the model performance. For example, if the training duration of the current model is significantly longer than that of the reference model, adjust the learning rate, optimizer type, or batch size to increase the convergence speed of the current model, thereby reducing the training duration of the current model.
[0092] Regarding the above step S3-1, it should be noted that the above-mentioned consistency synchronization description parameter is used to quantify the consistency and stability of data synchronization between different nodes during model training. Specifically, the above-mentioned consistency synchronization description parameter is used to describe the delay of each distributed node synchronizing data to the above-mentioned shared memory pool. When the number of distributed nodes with a delay exceeding the preset time threshold exceeds the preset node number threshold, it is confirmed that the consistency synchronization description parameter meets the third sub-goal adjustment condition. At this time, at least one hyperparameter is adjusted to optimize the data synchronization process. For example, increase the data synchronization frequency or change the data transmission protocol to improve the efficiency of data synchronization to the shared memory pool.
[0093] It should be noted that the data transmission delay in step S1-1 can be the transmission delay of a certain node or the average data transmission delay of multiple nodes. When the data transmission delay is greater than a preset delay threshold (the first delay threshold), hyperparameter adjustment is triggered. The consistency synchronization description parameter in step S3-1 is used to record the delay of each node synchronizing data to the shared memory pool. For step S3-1, a lower-value delay threshold (the second delay threshold) can be preset. When the number of distributed nodes with a transmission delay greater than the second delay threshold exceeds the preset node number threshold, hyperparameter adjustment is triggered.
[0094] For the above step S4-1, record the first sample prediction accuracy of the current model based on the training sample set and the second sample prediction accuracy based on the validation sample set. If the above first sample prediction accuracy and second sample prediction accuracy meet the fourth sub-goal adjustment condition, that is, the difference between the prediction accuracy of the current model on the validation sample set and the prediction accuracy on the training sample set exceeds the set tolerance range, it indicates that the current model has overfitting during training. At this time, the actual training effect of the current model is not as expected. At this time, adjust at least one hyperparameter to improve the prediction accuracy of the current model. For example, by adjusting the learning rate, regularization parameter, or optimizer type, the current model can more stably learn the features of the data, thereby avoiding overfitting and improving the prediction accuracy of the current model.
[0095] Through the above implementation manners, adjustment conditions for multiple hyperparameters are provided for the current model. When the adjustment conditions for the hyperparameters are met, the hyperparameters of the current model are adjusted to improve the training effect of the current model.
[0096] As an optional implementation manner, adjusting at least one hyperparameter matching the current model includes at least one of the following:
[0097] Method 1: Adjust the structure parameter indicating the model structure of the current model;
[0098] Method 2: Adjust the node description parameter indicating the number of nodes of the distributed nodes;
[0099] Method 3: Adjust the memory description parameter indicating the memory size of the shared memory pool;
[0100] Method 4: Adjust the process description parameter indicating the distribution of training processes running in multiple distributed nodes.
[0101] It can be understood that the above methods 1 to 4 are specific hyperparameter adjustment methods. The functions of different hyperparameter adjustments are specifically described below.
[0102] For the above-mentioned first method, the above-mentioned structural parameters are the parameters defining the current model architecture, including but not limited to the number of network layers, the number of neurons in each layer, the type of activation function, the loss function, etc. During the process of distributed training, the adjustment of the above-mentioned structural parameters can be to fine-tune the complexity of the model to adapt to the computing resources and communication conditions in the multi-node environment. For example, increasing or decreasing the number of network layers to adjust the training speed and prediction accuracy of the model.
[0103] For another example, by replacing different activation functions or loss functions to adjust the training effect of the current model, so as to achieve the purpose of preventing gradient explosion and accelerating convergence. It can be understood that by timely adjusting the structural parameters, it can be ensured that the current model is neither too complex to cause difficulties in training during distributed training, and at the same time avoid that the current model is too simple in structure to fully capture the complexity of the data, so that the current model achieves the best balance between model performance and training efficiency.
[0104] For the above-mentioned second method, it can be understood that the number of nodes indicated by the above-mentioned node description parameters affects the parallelism and data processing ability of distributed training. Increasing the number of nodes can enhance the parallel processing ability of the above-mentioned current model training and accelerate data processing and feature synchronization. Therefore, the number of nodes can be dynamically adjusted according to the resource consumption, communication delay and data processing requirements of the current training.
[0105] For example, at the initial stage of the current model training, the number of nodes can be increased to accelerate the model's learning of data features; while in the later stage of training, if it is found that the communication delay becomes a bottleneck, the number of nodes can be reduced to reduce the communication overhead, thereby improving the stability and efficiency of model training.
[0106] For the above-mentioned third method, it can be understood that the memory size indicated by the above-mentioned memory description parameters determines the capacity of the shared memory pool, affecting the efficiency of data synchronization between distributed nodes and the utilization rate of computing resources of multiple nodes. Expanding the memory size can accommodate more data and intermediate calculation results, reducing the number of data transmissions between nodes, thereby reducing the communication delay. Therefore, according to the memory requirements of the current training task and the hardware conditions, when the memory requirements of the current training task are not high, the memory size of the shared memory pool can be reduced to avoid resource waste, while when the memory requirements of the current training task are large, the memory size of the shared memory pool can be increased to meet the real-time requirements of the training process.
[0107] For the above-mentioned fourth method, it should be noted that the process description parameters involve how to allocate and execute training processes on each node in multi-node distributed training. By adjusting this parameter, the load balancing of the training task among multiple nodes can be optimized to ensure that the computing resources of each node are fully utilized.
[0108] For example, according to the hardware performance of the nodes and the computing requirements of the current training task, the execution weights of the training processes are dynamically allocated, so that compute-intensive tasks run on high-performance nodes, while communication-intensive tasks are executed on nodes with high-bandwidth network connections. Another example is that when the resource utilization rate of a certain node is lower than a preset threshold, at least one new process is allocated to this node to increase the resource utilization rate of this node. Thus, by reasonably adjusting the process description parameters, more efficient distributed computing is achieved, and the overall speed of model training is improved.
[0109] As an alternative implementation, refer to Figure 4 , and adjust at least one hyperparameter that matches the current model, including:
[0110] S402, obtain the hyperparameter sequences corresponding to each of the multiple current models in the training state, where the hyperparameter sequence includes at least one hyperparameter, and the hyperparameter sequence is used to indicate the training methods corresponding to each of the multiple current models in the training state;
[0111] S404, determine at least one target hyperparameter sequence from the multiple hyperparameter sequences, where the at least one target hyperparameter sequence is determined based on the sorting result of the multiple hyperparameter sequences according to the target fitness, and the target fitness is the proportion of the target parameter in at least one current parameter of each of the multiple current models in the training state;
[0112] S406, adjust at least one hyperparameter that matches the current model according to the at least one target hyperparameter sequence.
[0113] It should be noted that the above steps S402 to S406 can achieve the automatic adjustment of specific hyperparameters, which will be specifically described below.
[0114] First, configure multiple different hyperparameter sequences for the model to be trained, and use different hyperparameter sequences to train the model to be trained respectively, so as to generate multiple different current models, and obtain the hyperparameter sequences and the corresponding current parameters corresponding to each of the different current models during the training process. It can be understood that each current model instance has a specific set of hyperparameter configurations during the training process, including but not limited to learning rate, optimizer type, regularization parameter, batch size, etc., which are not specifically limited here. The timing of the above acquisition can be after each current model completes the backward update.
[0115] After obtaining multiple hyperparameter sequences and their corresponding current parameters, these sequences are sorted according to the objective fitness, and finally at least one target hyperparameter sequence is determined. It should be noted that the above sorting process can be implemented by a genetic algorithm. Specifically, each hyperparameter sequence is regarded as an individual in the genetic algorithm, and different hyperparameters in each individual are regarded as a gene. By performing operations such as exchanging and mutating different genes in the individual, iterations are carried out to generate multiple new individuals, that is, multiple new hyperparameter sequences, and the current model is continuously trained according to the newly generated hyperparameter sequences. At the same time, after each backward update, the objective fitness corresponding to each hyperparameter sequence is calculated. Specifically, the current model corresponding to each hyperparameter sequence is determined, and the ratio of the target model parameters (the current parameters whose corresponding target error values are greater than the corresponding parameter benchmark error values) in the corresponding current model to all the current parameters is determined as the objective fitness. When the number of iterations reaches the preset number, the hyperparameter sequences corresponding to one or more individuals with the smallest objective fitness are used as one or more target hyperparameter sequences.
[0116] Furthermore, at least one hyperparameter matching the current model is adjusted according to at least one target hyperparameter sequence.
[0117] Through the above implementation manner, in distributed training, by determining at least one target hyperparameter sequence from multiple hyperparameter sequences, the model can find the optimal hyperparameter combination, thereby improving the training effect of the model.
[0118] As an alternative implementation manner, refer to Figure 5 , after determining the respective target error values between at least one reference model parameter and at least one current parameter, it includes:
[0119] S502, obtain multiple target error values;
[0120] S504, determine the target error values greater than the first target threshold as the first reference target error values, and determine the target error values less than the second target threshold as the second reference target error values;
[0121] S506, adjust the error reference value and / or the scaling reference value according to the quantities of the first reference target error values and the second reference target error values.
[0122] It should be noted that, as described above, If the target error value is large, it means that the difference between the current parameters and the reference model parameters is large. If the target error value is small, it means that the difference between the current parameters and the reference model parameters is small. Further, when there are many target error values with large values, it means that the training result of the current model is poor. In this case, the scaling reference value and / or the error reference value are reduced to strictly control the relative difference of the parameter update, thereby ensuring the consistency of the current model with the reference model.
[0123] In addition, when there are more target error values with smaller numerical values, it means that the training result of the current model is better. At this time, the error reference value and / or the scaling reference value can be appropriately increased to relax the consistency requirements between the previous model and the reference model, thereby improving the training efficiency of the model.
[0124] Through the above implementation, dynamic adjustment between the scaling reference value and the error reference value is achieved according to the numerical values of the acquired multiple target error values, so as to achieve a balance between the training effect and the training efficiency of the current model.
[0125] As an optional implementation manner, after determining the parameter reference error value, the method further includes:
[0126] S1, when the ratio of the number of target model parameters to the number of current parameters is less than or equal to a first target adjustment threshold, scaling down at least one of the error reference value and the scaling reference value;
[0127] S2, updating the parameter reference error value when at least one of the error reference value and the scaling reference value has completed the downscaling adjustment;
[0128] S3, determining whether the target error value is greater than the updated parameter reference error value.
[0129] It should be noted that, when the ratio of the number of target model parameters to the number of current parameters is less than or equal to the first target adjustment threshold, it means that the current model maintains good consistency as a whole, and there is no significant parameter deviation that requires special processing. Therefore, the error reference value and / or the scaling reference value can be reduced, and then according to the above formula: , achieving a reduction in the parameter benchmark error value, thereby further improving the evaluation criteria for model consistency.
[0130] Then, it is determined whether the target error value is greater than the adjusted parameter benchmark error value. If the ratio of the number of current parameters (new target model parameters) whose corresponding target error values are greater than the adjusted parameter benchmark error value to the number of all current parameters is once again less than or equal to the first target adjustment threshold, it means that the error reference value and / or scaling reference value can still be reduced to continue to improve the evaluation criteria for model consistency.
[0131] As an alternative implementation, obtaining the error reference value and the scaling reference value includes:
[0132] S1. Obtain the set batches of the training sample set, where the training sample set is used to train the current model;
[0133] S2. Obtain the error reference value and the scaling reference value that match the set batches, where the value of at least one of the error reference value and the scaling reference value is negatively correlated with the size of the set batches.
[0134] It should be noted that the values of the above error reference value and scaling reference value are negatively correlated with the size of the training sample set batches, that is, the larger the batch of the training sample set, the smaller the error reference value and scaling reference value.
[0135] Furthermore, for larger set batches, since more training samples are used for gradient calculation, the stability of model parameter updates will be improved. Therefore, a smaller error reference value can be set to reduce the error of model parameter updates and further improve the training accuracy of the model. Conversely, for smaller set batches, the model parameter updates are unstable and are easily affected by noise and local optima. At this time, the error reference value can be appropriately increased to allow a certain degree of error and improve the stability of model training.
[0136] Figure 6 is a flowchart of another detection method for a distributed training model according to an embodiment of the present application. The following is combined with Figure 6 to illustrate the detection process of a distributed training model.
[0137] S602. Train the benchmark model.
[0138] It should be noted that to evaluate the distributed training results of the current model, a model with the same model structure as the current model or a structure similarity greater than the threshold needs to be used as the benchmark model (reference model), and the benchmark model should have high accuracy. For example, train the benchmark model through a distributed environment based on the PCIE protocol, or through a single-node (non-distributed) environment based on the CXL protocol, and record the parameters of the obtained reference model as
[0139] The above two training methods are relatively mature, so better training effects can be obtained. Therefore, they can be used as the benchmark for evaluating the distributed training results of the current model.
[0140] S604. Perform CXL distributed training.
[0141] Train the current model using a distributed environment with a shared memory pool based on the CXL protocol. Obtain the current parameters of the current model 。
[0142] It should be noted that during this process, due to the low latency and high bandwidth characteristics of the CXL protocol, the synchronization of the current parameters and gradients is accelerated, while the shared memory mechanism ensures data consistency across nodes.
[0143] S606, calculate the deviation of the model parameters.
[0144] Specifically, compare the current parameters with the reference model parameters for their numerical deviation. The calculation formula for the numerical deviation is as follows:
[0145]
[0146] It should be noted that the above absolute error reflects the basic numerical difference between the two parameters. For example, it can be set to 1e-5, that is, the allowable absolute deviation after the decimal point is within 0.00001. The above relative error represents the error of the relative magnitude of the two parameter values. For example, it can be set to 1e-3 or 1e-4, corresponding to the relative error in the third or fourth digit after the decimal point, which is used to measure the ratio of the deviation to the reference value.
[0147] For each current parameter, if the above conditions are met, it is considered that the current parameter is consistent with the corresponding reference model parameter. After all the current parameters of the current model pass the verification of the above conditions, it can be determined that the consistency between the CXL distributed training and the benchmark training results is relatively high. It should be noted that the above absolute error and relative error can be selected according to the numerical stability of the experiment and the actual error requirements.
[0148] It should be noted that the above absolute error and relative error are used to ensure that the model can still produce sufficiently consistent outputs under the floating-point calculation error and a specific noise level in the distributed training environment. If large numerical fluctuations are found during the training process, the relative error and absolute error need to be adjusted to ensure the accuracy of the model.
[0149] S608, record the performance indicators of consistency.
[0150] It should be noted that in actual training, the performance indicators of various consistencies also need to be recorded to comprehensively evaluate the consistency of the distributed training results. Specifically, it includes recording the deviation ratio, transmission delay and training speed, memory consistency delay, and precision loss.
[0151] It should be noted that the above deviation ratio is used to count the proportion of parameters that meet the above relative error and absolute error requirements. It can be understood that the higher the deviation ratio, the better the numerical consistency between CXL distributed training and the baseline method. The above transmission delay and training speed specifically refer to the transmission delay of each training cycle and the total training time. Due to the low-latency feature provided by the CXL protocol, data synchronization across nodes is more rapid. By recording the transmission delay of each training cycle and the total training time, the speed advantage of CXL in distributed training can be evaluated.
[0152] It should be noted that the above memory consistency delay is used to describe the delay of each distributed node synchronizing data to the above shared memory pool. By recording the delay of memory consistency synchronization, the contribution of the consistency mechanism of CXL to performance can be further quantified. The above accuracy loss refers to the difference in the accuracy of the trained model on the validation set and the training set in a multi-node scenario.
[0153] In this embodiment, a strict numerical consistency evaluation mechanism is provided through relative error and absolute error, which can help detect subtle deviations in training, thereby ensuring the accuracy of model training. In addition, the low-latency feature of CXL enables high-consistency model parameters to be maintained in large-scale model training, with faster synchronization speed and a more stable training process, thus providing a scientific and reliable evaluation basis for distributed training.
[0154] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0155] According to another aspect of the embodiments of the present application, there is also provided a detection device for a distributed training model for implementing the above detection method of the distributed training model. As Figure 7 shown, the device includes:
[0156] An acquisition unit 702, which acquires at least one reference model parameter when the reference model reaches the training convergence state, and at least one current parameter obtained after the current model performs distributed node training operations, where the structural similarity between the model structure of the reference model and the model structure of the current model is greater than a target threshold, and the distributed node training operation is a training operation based on multiple distributed nodes with a shared memory pool;
[0157] A generating unit 704, configured to generate at least one target error value between at least one current parameter and a corresponding reference model parameter, where the target error value is used to indicate the degree of difference between the current parameter and the corresponding reference model parameter;
[0158] A detecting unit 706, configured to generate a training detection result of the current model according to at least one target error value and a parameter reference error value corresponding to the reference model, where the parameter reference error value is used to indicate the allowable degree of difference between the reference model parameter and the current parameter.
[0159] Optionally, the above-mentioned detecting unit 706 includes: a product obtaining module, configured to obtain the product between a scaling reference value for adjusting the model parameter and the reference model parameter; a parameter reference error value determining module, configured to determine the sum of the product and the error reference value as the parameter reference error value; a detection result determining unit, configured to, when the target error value generated based on the current parameter is greater than the parameter reference error value corresponding to the current parameter, determine the current parameter as the target model parameter; generate a training detection result based on the ratio of the number of target model parameters to the number of current parameters, or generate a training detection result based on the number of target model parameters.
[0160] Optionally, the above-mentioned apparatus further includes a current parameter adjusting unit, configured to adjust the current parameter of the current model when the ratio of the number of target model parameters to the number of current parameters is greater than or equal to a first target adjustment threshold; or adjust the current parameter of the current model when the number of target model parameters is greater than or equal to a second target adjustment threshold.
[0161] Optionally, the above-mentioned obtaining unit 702 includes: a model structure obtaining module, configured to obtain the model structure corresponding to the current model and the model structure corresponding to the candidate model, where the candidate model has the same type of executable task as the current model; a structure similarity determining module, configured to determine the structure similarity between the current model and the candidate model according to the model structure corresponding to the current model and the model structure corresponding to the candidate model; a reference model determining module, configured to, when the structure similarity is greater than the target threshold, determine the candidate model as the reference model; a reference model parameter obtaining module, configured to obtain at least one reference model parameter of the reference model when the reference model reaches training convergence; a current parameter determining module, configured to obtain at least one current parameter of the current model after the current model performs a distributed node training operation.
[0162] Optionally, the above-mentioned generating unit 704 includes: a difference obtaining module, configured to obtain the respective differences between at least one current parameter and the corresponding reference model parameter; a target error value determining module, configured to determine the absolute value of at least one difference as at least one target error value.
[0163] Optionally, the above device further includes a first reference value adjustment unit, configured to perform a reduction adjustment on at least one of an error reference value and a scaling reference value; update a parameter reference error value when at least one of the error reference value and the scaling reference value is adjusted; and determine whether a target error value is greater than the updated parameter reference error value.
[0164] Optionally, the above device further includes a first hyperparameter adjustment unit, configured to adjust at least one hyperparameter matching the current model when a training detection result indicates an adjustment to the current parameter, where the hyperparameter is used to indicate a training method of the current model.
[0165] Optionally, the above device further includes a reference acquisition unit, configured to acquire a data transmission delay of the current model performing a training operation based on a training sample set during a process of generating a training detection result of the current model according to at least one target error value and a parameter reference error value corresponding to a reference model; acquire a first training duration of the current model performing a training operation based on the training sample set and a second training duration of the reference model performing a synchronous training operation based on the training sample set; acquire a consistency synchronization description parameter of the current model performing a training operation based on the training sample set; acquire a first sample prediction accuracy of the current model based on the training sample set, and acquire a second sample prediction accuracy of the current model based on a validation sample set.
[0166] Optionally, the above device further includes a second hyperparameter adjustment unit, configured to adjust at least one hyperparameter matching the current model when the data transmission delay meets a first sub-goal adjustment condition; adjust at least one hyperparameter matching the current model when the first training duration and the second training duration meet a second sub-goal adjustment condition; adjust at least one hyperparameter matching the current model when the consistency synchronization description parameter meets a third sub-goal adjustment condition; adjust at least one hyperparameter matching the current model when the first sample prediction accuracy and the second sample prediction accuracy meet a fourth sub-goal adjustment condition.
[0167] Optionally, the above first hyperparameter adjustment unit or second hyperparameter adjustment unit is further configured to adjust a structure parameter for indicating a model structure of the current model; adjust a node description parameter for indicating a number of nodes of a distributed node; adjust a memory description parameter for indicating a memory size of a shared memory pool; adjust a process description parameter for indicating a distribution of training processes running in a plurality of distributed nodes.
[0168] Optionally, the above-mentioned second hyperparameter adjustment unit is further configured to obtain hyperparameter sequences corresponding to multiple current models in a training state, where each hyperparameter sequence includes at least one hyperparameter, and the hyperparameter sequence is used to indicate the training methods corresponding to the multiple current models in a training state; determine at least one target hyperparameter sequence from the multiple hyperparameter sequences, where the at least one target hyperparameter sequence is determined based on the sorting result of the multiple hyperparameter sequences according to the target fitness, and the target fitness is the proportion of the target parameter among at least one current parameter of the multiple current models in a training state; adjust at least one hyperparameter matching the current model according to the at least one target hyperparameter sequence.
[0169] Optionally, the above-mentioned device further includes a batch size reference unit, configured to obtain the set batch of the training sample set, where the training sample set is used to train the current model; obtain an error reference value and a scaling reference value that match the set batch of the training sample set, where the value of at least one of the error reference value and the scaling reference value is negatively correlated with the size of the set batch of the training sample set.
[0170] Optionally, the above-mentioned device further includes a second reference value adjustment unit, configured to obtain multiple target error values; determine the target error values greater than the first target threshold as the first reference target error values, and determine the target error values less than the second target threshold as the second reference target error values; adjust the error reference value and / or the scaling reference value according to the quantities of the first reference target error values and the second reference target error values.
[0171] Optionally, the above-mentioned second reference value adjustment unit is further configured to decrease the error reference value and / or the scaling reference value when the quantity of the first reference target error values is greater than the first quantity threshold; increase the error reference value and / or the scaling reference value when the quantity of the second reference target error values is greater than the second quantity threshold.
[0172] According to another aspect of the embodiments of the present application, there is also provided an electronic device for implementing the above-mentioned detection method for a distributed training model. The electronic device may be Figure 1 the terminal device or server shown. In this embodiment, the electronic device is taken as a mobile phone or a computer as an example. As Figure 8 shown, the electronic device includes a memory 802 and a processor 804. A computer program is stored in the memory 802, and the processor 804 is configured to execute the steps in any one of the above method embodiments through the computer program.
[0173] Optionally, in this embodiment, the above-mentioned electronic device may be at least one network device among multiple network devices in a computer network.
[0174] Optionally, in this embodiment, the above processor may be configured to perform the following steps by a computer program:
[0175] S1. Obtain at least one reference model parameter when the reference model reaches the training convergence state, and at least one current parameter obtained after the current model performs the distributed node training operation, where the structural similarity between the model structure of the reference model and the model structure of the current model is greater than the target threshold, and the distributed node training operation is a training operation based on multiple distributed nodes with a shared memory pool;
[0176] S2. Generate at least one target error value between at least one current parameter and the corresponding reference model parameter, where the target error value is used to indicate the degree of difference between the current parameter and the corresponding reference model parameter;
[0177] S3. Generate a training detection result of the current model according to at least one target error value and the parameter benchmark error value corresponding to the reference model, where the parameter benchmark error value is used to indicate the allowable degree of difference between the reference model parameter and the current parameter.
[0178] Optionally, those of ordinary skill in the art can understand that Figure 8 The structure shown is only schematic, and the electronic device can also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, and terminal devices such as Mobile Internet Devices (MID), PAD, etc. Figure 8 It does not limit the structure of the above electronic device. For example, the electronic device may further include more or fewer components (such as a network interface, etc.) than those shown Figure 8 in, or have a different configuration from that shown Figure 8 in.
[0179] Among them, the memory 802 can be used to store software programs and modules, such as the program instructions / modules corresponding to the detection method and device of the distributed training model in the embodiments of the present application. The processor 804 executes various functional applications and data processing by running the software programs and modules stored in the memory 802, that is, implements the above method for evaluating the distributed training result. The memory 802 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, a flash memory, or other non-volatile solid-state memories. In some instances, the memory 802 may further include a memory remotely set relative to the processor 804, and these remote memories can be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, an enterprise internal network, a local area network, a mobile communication network, and combinations thereof.
[0180] As an example, such asFigure 8 As shown, the above-mentioned memory 802 may include, but is not limited to, the acquisition unit 702, the generation unit 704, and the detection unit 706 in the above-mentioned device. In addition, it may also include, but is not limited to, other module units in the above-mentioned device, which will not be elaborated in this example.
[0181] Optionally, the above-mentioned transmission device 806 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wired network and a wireless network. In one example, the transmission device 806 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices and routers through a network cable, so as to communicate with the Internet or a local area network. In one example, the transmission device 806 is a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0182] In addition, the above-mentioned electronic device further includes: a connection bus 808, which is used to connect each module component in the above-mentioned electronic device.
[0183] In other embodiments, the above-mentioned terminal device or server may be a node in a distributed system. Among them, the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting multiple nodes through network communication. Among them, the nodes can form a point-to-point network, and any form of computing device, such as a server, a terminal, etc., can become a node in the blockchain system by joining the point-to-point network.
[0184] According to one aspect of the present application, a computer-readable storage medium is provided. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the above various optional implementation manners;
[0185] Optionally, in this embodiment, the above-mentioned computer-readable storage medium may be set to store a computer program for executing the following steps:
[0186] S1, obtain at least one reference model parameter when the reference model reaches the training convergence state, and at least one current parameter obtained after the current model performs the distributed node training operation, where the structural similarity between the model structure of the reference model and the model structure of the current model is greater than the target threshold, and the distributed node training operation is a training operation based on multiple distributed nodes with a shared memory pool;
[0187] S2. Generate at least one target error value between at least one current parameter and the corresponding reference model parameter, where the target error value is used to indicate the degree of difference between the current parameter and the corresponding reference model parameter;
[0188] S3. Generate a training detection result of the current model according to at least one target error value and a parameter reference error value corresponding to the reference model, where the parameter reference error value is used to indicate the allowable degree of difference between the reference model parameter and the current parameter.
[0189] Optionally, in the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.
[0190] Optionally, in this embodiment, those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware of the terminal device through a program, and the program can be stored in a computer-readable storage medium. The storage medium can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disc, etc.
[0191] If the integrated unit in the above embodiments is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in the above computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in the storage medium and includes several instructions for causing one or more computer devices (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present application.
[0192] In the above embodiments of the present application, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0193] In addition, in each embodiment of the present application, each functional unit may be integrated into a processing unit, may exist separately as individual physical units, or two or more units may be integrated into one unit. The above-mentioned integrated units may be implemented in the form of hardware or in the form of software functional units.
[0194] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A method for detecting a distributed training model, characterized in that: It includes: Obtain at least one reference model parameter when the reference model reaches the training convergence state, and at least one current parameter obtained after the current model performs the distributed node training operation, wherein the structural similarity between the model structure of the reference model and the model structure of the current model is greater than the target threshold, and the distributed node training operation is a training operation based on multiple distributed nodes with a shared memory pool. The reference model and the current model are used to perform the image-text matching task, and the reference model parameters and the current parameters correspond one by one; Generate at least one target error value between at least one of the current parameters and the corresponding reference model parameters, wherein the target error value is used to indicate the degree of difference between the current parameter and the corresponding reference model parameter; Generate a training detection result of the current model according to at least one of the target error values and the parameter benchmark error value corresponding to the reference model, so as to determine the accuracy of the current model in matching images and texts, wherein the parameter benchmark error value is used to indicate the allowable degree of difference between the reference model parameter and the current parameter; Wherein, the generating the training detection result of the current model according to at least one of the target error values and the parameter benchmark error value corresponding to the reference model includes: When the target error value generated based on the current parameter is greater than the parameter benchmark error value corresponding to the current parameter, determine the current parameter as the target model parameter; Generate the training detection result based on the ratio of the number of the target model parameters to the number of the current parameters, or generate the training detection result based on the number of the target model parameters.
2. The method according to claim 1, characterized in that: Determining the parameter benchmark error value corresponding to the reference model includes: Obtain the product between the scaling reference value for adjusting the model parameter and the reference model parameter; Determine the sum of the product and the error reference value as the parameter benchmark error value.
3. The method according to claim 2, characterized in that: After determining the parameter benchmark error value corresponding to the reference model, the method further includes: When the ratio of the number of the target model parameters to the number of the current parameters is less than or equal to the first target adjustment threshold, reduce at least one of the error reference value and the scaling reference value; When the reduction adjustment is completed, update the parameter benchmark error value; Determine the target model parameter based on the updated parameter benchmark error value.
4. The method according to claim 1, characterized in that: After generating the training detection result based on the ratio of the number of the target model parameters to the number of the current parameters, the method further includes: when the ratio of the number of the target model parameters to the number of the current parameters is greater than or equal to the first target adjustment threshold, adjust the current parameters of the current model; or After generating the training detection result based on the quantity of the target model parameters, the method further includes: when the quantity of the target model parameters is greater than or equal to a second target adjustment threshold, adjusting the current parameters of the current model.
5. The method according to claim 1, wherein The obtaining of at least one reference model parameter when the reference model reaches the training convergence state and at least one current parameter obtained after the current model performs a distributed node training operation includes: obtaining the model structure corresponding to the current model and the model structure corresponding to a candidate model, where the candidate model has the same executable task type as the current model; determining the structural similarity between the current model and the candidate model according to the model structure corresponding to the current model and the model structure corresponding to the candidate model; when the structural similarity is greater than the target threshold, determining the candidate model as the reference model; when the reference model reaches training convergence, obtaining at least one of the reference model parameters of the reference model; after the current model performs a distributed node training operation, obtaining at least one of the current parameters of the current model.
6. The method according to claim 1, wherein the generating of at least one target error value between at least one of the current parameters and the corresponding reference model parameters includes: obtaining the respective differences between at least one of the current parameters and the corresponding reference model parameters; determining the absolute value of at least one of the differences as at least one of the target error values.
7. The method according to claim 4, wherein when the training detection result indicates an adjustment to the current parameters, the method further includes: adjusting at least one hyperparameter matching the current model, where the hyperparameter is used to indicate the training manner of the current model.
8. The method according to claim 1, wherein in the process of generating the training detection result of the current model according to at least one of the target error values and the parameter benchmark error value corresponding to the reference model, the method further includes at least one of the following: obtaining the data transmission delay of the current model performing a training operation based on a training sample set; obtaining the first training duration of the current model performing the training operation based on the training sample set and the second training duration of the reference model performing a synchronous training operation based on the training sample set; obtaining the consistency synchronization description parameter of the current model performing a training operation based on the training sample set; obtaining the first sample prediction accuracy of the current model based on the training sample set and obtaining the second sample prediction accuracy of the current model based on a validation sample set.
9. The method according to claim 8, wherein The method further includes at least one of the following: when the data transmission delay meets a first sub-target adjustment condition, adjusting at least one hyperparameter matching the current model; When the first training duration and the second training duration meet the second sub-goal adjustment condition, at least one hyperparameter matched with the current model is adjusted; When the consistency synchronization description parameter meets the third sub-goal adjustment condition, at least one hyperparameter matched with the current model is adjusted; When the first sample prediction accuracy and the second sample prediction accuracy meet the fourth sub-goal adjustment condition, at least one hyperparameter matched with the current model is adjusted.
10. The method according to any one of claims 7 or 9, characterized in that The adjustment of at least one hyperparameter matched with the current model includes at least one of the following: Adjusting the structure parameter in the at least one hyperparameter for indicating the model structure of the current model; Adjusting the node description parameter in the at least one hyperparameter for indicating the number of nodes of the distributed nodes; Adjusting the memory description parameter in the at least one hyperparameter for indicating the memory size of the shared memory pool; Adjusting the process description parameter in the at least one hyperparameter for indicating the distribution of training processes running in multiple distributed nodes.
11. The method according to claim 10, characterized in that The adjustment of at least one hyperparameter matched with the current model further includes: Obtaining hyperparameter sequences corresponding to the current models in a plurality of training states, wherein the hyperparameter sequences include at least one of the hyperparameters, and the hyperparameter sequences are used to indicate the training methods corresponding to the current models in a plurality of training states; Determining at least one target hyperparameter sequence from the plurality of hyperparameter sequences, wherein the at least one target hyperparameter sequence is determined based on the sorting result of the plurality of hyperparameter sequences according to the target fitness, and the target fitness is the ratio of the number of target model parameters to the number of current parameters; Adjusting at least one hyperparameter matched with the current model according to the at least one target hyperparameter sequence.
12. The method according to claim 2, wherein After generating at least one target error value between the at least one reference model parameter and the current parameter, it includes: Obtaining a plurality of the target error values; Determining the target error values greater than the first target threshold as the first reference target error values, and determining the target error values less than the second target threshold as the second reference target error values; Adjusting the error reference value and / or the scaling reference value according to the number of the first reference target error values and the second reference target error values.
13. A detection device for a distributed training model, characterized in that Comprising: An acquisition unit, configured to acquire at least one reference model parameter when a reference model reaches a training convergence state, and at least one current parameter obtained after the current model performs a distributed node training operation, where a structural similarity between a model structure of the reference model and a model structure of the current model is greater than a target threshold, the distributed node training operation is a training operation based on a plurality of distributed nodes having a shared memory pool, the reference model and the current model are used to perform a graphic-text matching task, and the reference model parameters and the current parameters are in one-to-one correspondence; A generation unit, configured to generate at least one target error value between at least one of the current parameters and the corresponding reference model parameters, where the target error value is used to indicate a difference degree between the current parameter and the corresponding reference model parameter; A detection unit, configured to generate a training detection result of the current model according to at least one of the target error values and a parameter reference error value corresponding to the reference model, so as to determine an accuracy of matching between the current model and an image and text, where the parameter reference error value is used to indicate an allowable difference degree between the reference model parameter and the current parameter; The apparatus is further configured to, when a target error value generated based on the current parameter is greater than the parameter reference error value corresponding to the current parameter, determine the current parameter as a target model parameter; generate the training detection result based on a ratio of a number of the target model parameters to a number of the current parameters, or generate the training detection result based on a number of the target model parameters.
14. A computer-readable storage medium, characterized in that a computer program is stored in the computer-readable storage medium, where the computer program, when executed by a processor, implements the steps of the method according to any one of claims 1 to 12.
15. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that when the processor executes the computer program, the steps of the method according to any one of claims 1 to 12 are implemented.
16. A computer program product, comprising a computer program, characterized in that when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.
Citation Information
Patent Citations
Using distributed learning to develop a machine learning model
WO2024125787A1