A business processing method, test system, medium and product
By determining the target topology and dynamically adjusting model parameters, the problem of underutilization of computing accelerator resources is solved, and the efficiency and evaluation effect of model training are improved.
Patent Information
- Application Number
- CN202411621636.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-11-14
AI Technical Summary
In the prior art, the resource characteristics of the computing accelerator are not fully utilized during the model training process, resulting in a reduction in the operation efficiency of the training process, and the human experience selecting model parameters, the adjustment process takes a lot of time and poor evaluation results.
By obtaining the weight parameters of the target model, communication strategy and preset topology, the performance parameters are processed to determine the target topology, and the target model is trained based on the parallel training strategy and the target sample number, and the model parameters are dynamically adjusted to improve training efficiency and evaluation effect.
It realizes the full use of computing accelerator resources, improves the operation efficiency and evaluation effect of model training, and reduces the time and complexity of artificial parameter adjustment.
Smart Images

Figure CN119149245B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a business processing method, a test system, a medium and a product. Background Art
[0002] As the complexity and scale of model operations continue to increase, the usage, bandwidth, and computing power of hardware resources used by multi-performance computing accelerators (such as graphics processing units (GPUs) and field-programmable gate arrays (FPGAs)) during model training cannot change dynamically, and the resource characteristics of computing accelerators cannot be fully utilized, resulting in reduced efficiency in the training process. Model training parameters corresponding to user business needs are often based on the selection of initial parameters based on human experience parameters for training, and the training parameters of the current iteration are adjusted according to the results of each training to obtain the final parameter values of the model. The running process of the training model and the adjustment process of the training parameter selection make the entire process time-consuming, and the evaluation and testing effect of the training parameters is poor.
[0003] Therefore, how to utilize the resource characteristics of computing accelerators to improve the operating efficiency of the training process and evaluate the test results is an urgent problem that technical personnel in this field need to solve. Summary of the invention
[0004] The purpose of the present invention is to provide a business processing method, a test system, a medium and a product to solve the problem that in the model training process corresponding to conventional user business needs, the resource characteristics of the computing accelerator are not fully utilized, resulting in reduced operating efficiency of the training process, and the model parameters are selected based on human experience and the adjustment process is based on the adjustment after the entire operation process is completed, resulting in a long time consumption and poor evaluation effect.
[0005] In order to solve the above technical problems, the present invention provides a service processing method, comprising:
[0006] Obtain the target model corresponding to the current business needs, the weight parameters of each set communication strategy, and the preset topology structure composed of each computing accelerator;
[0007] Processing the performance parameters of each computing accelerator corresponding to each preset topology structure, the weight parameter and each collective communication strategy to determine a target topology structure corresponding to the benchmark test;
[0008] Based on the target topology, the target model is trained according to the parallel training strategies corresponding to each of the computing accelerators and the target number of samples of the target model to obtain the model parameters of the target model during the operation of the target model to complete the processing of the current business needs.
[0009] On the one hand, based on the target topology structure, the target model is trained according to the parallel training strategy corresponding to each of the computing accelerators and the target number of samples of the target model to obtain the model parameters of the target model during the operation of the target model, including:
[0010] Obtaining an initial parallel training strategy corresponding to each of the computing accelerators under the target topology structure;
[0011] Determining, within each initial parallel training strategy, a parallel training strategy corresponding to the target model of the current business requirement;
[0012] Based on the computing accelerator of the parallel training strategy, the training data corresponding to the target number of samples is input into the target model for training processing to determine the model parameters of the target model during the operation of the target model.
[0013] On the other hand, determining a target topology structure according to each preset topology structure, the weight parameter and the performance parameter of each computing accelerator corresponding to each collective communication strategy includes:
[0014] Processing each collective communication strategy in each of the preset topological structures to obtain the corresponding performance parameters;
[0015] Determine an evaluation score of each of the preset topological structures according to the performance parameter and the weight parameter of each of the collective communication strategies;
[0016] The target topology is determined according to the evaluation scores of the preset topologies.
[0017] On the other hand, the performance parameters include at least bandwidth and delay, and determining the evaluation score of each of the preset topological structures according to the performance parameters and the weight parameters of each of the collective communication strategies includes:
[0018] Determine the throughput corresponding to each of the collective communication strategies in the current preset topology structure according to the bandwidth corresponding to each of the collective communication strategies and the delay corresponding to each of the collective communication strategies;
[0019] The weight parameters and the throughput corresponding to each of the collective communication strategies are used to determine an evaluation score of the current preset topology structure.
[0020] On the other hand, before determining the target topology structure according to each preset topology structure, the weight parameter and the performance parameter of each computing accelerator corresponding to each collective communication strategy, the method further includes:
[0021] Determining the interconnection status of each of the computing accelerators in each of the preset topological structures;
[0022] When the interconnection state is normal, configuring link configuration information corresponding to the preset topology structure;
[0023] Testing the bandwidth between the computing accelerators and between the computing accelerators and the processor in each of the preset topological structures to obtain a first test result;
[0024] If the first test result satisfies the first preset condition, then proceeding to the step of determining the target topology structure according to each preset topology structure, the weight parameter and the performance parameter of each computing accelerator corresponding to each collective communication strategy;
[0025] If the first test result does not meet the first preset condition, return to the step of configuring the link configuration information corresponding to the preset topology structure to which the link configuration information belongs, and adjust the link configuration information until the first test result meets the first preset condition.
[0026] On the other hand, the first preset condition is a first sub-preset condition or a second sub-preset condition, the first sub-preset condition is a preset condition for bandwidth and delay balance of a topological structure within a connection domain; the second sub-preset condition is a preset condition for bandwidth and delay balance of a topological structure outside a connection domain, and a process for determining the first preset condition includes:
[0027] When each of the computing accelerators is in a connection domain topology structure, obtaining a first bandwidth standard deviation and a first delay standard deviation between the computing accelerators in the connection domain topology structure;
[0028] Obtaining a grouping method corresponding to computing accelerators in a topological structure within a connection domain;
[0029] If the number of groups corresponding to the grouping mode is 0, determine whether the first bandwidth standard deviation is less than or equal to the first bandwidth threshold and whether the first delay standard deviation is less than or equal to the first delay threshold; if the first bandwidth standard deviation is less than or equal to the first bandwidth threshold and the first delay standard deviation is less than or equal to the first delay threshold, determine that the first sub-preset condition is met; if the first bandwidth standard deviation is less than or equal to the first bandwidth threshold or the first delay standard deviation is less than or equal to the first delay threshold, determine that the first sub-preset condition is not met;
[0030] If the number of groups corresponding to the grouping mode is greater than 0, it is determined whether the first bandwidth standard deviation is less than or equal to the second bandwidth threshold and whether the first delay standard deviation is less than or equal to the second delay threshold; if the first bandwidth standard deviation is less than or equal to the second bandwidth threshold and the first delay standard deviation is less than or equal to the second delay threshold, it is determined that the first sub-preset condition is met; if the first bandwidth standard deviation is less than or equal to the second bandwidth threshold or the first delay standard deviation is less than or equal to the second delay threshold, it is determined that the first sub-preset condition is not met;
[0031] When each of the computing accelerators is in a topological structure outside the connection domain, obtaining a second bandwidth standard deviation and a second delay standard deviation between computing accelerators in the topological structure outside the connection domain and across the connection domain;
[0032] Determine whether the second bandwidth standard deviation is less than or equal to a third bandwidth threshold and whether the second delay standard deviation is less than or equal to a third delay threshold; if the second bandwidth standard deviation is less than or equal to the third bandwidth threshold and the second delay standard deviation is less than or equal to the third delay threshold, determine that the second sub-preset condition is met; if the second bandwidth standard deviation is less than or equal to the third bandwidth threshold or the second delay standard deviation is less than or equal to the third delay threshold, determine that the second sub-preset condition is not met.
[0033] On the other hand, determining the parallel training strategy corresponding to the target model of the current business requirement within each initial parallel training strategy includes:
[0034] Obtaining a preset segmentation ratio of the target model on each of the computing accelerators;
[0035] Preprocessing the target model according to different preset segmentation ratios to determine performance evaluation parameters corresponding to each of the initial parallel training strategies;
[0036] The parallel training strategy is determined according to the performance evaluation parameters.
[0037] On the other hand, the performance evaluation parameters include at least training iteration time, model performance and resource utilization; the preprocessing of the target model according to different preset segmentation ratios to determine the corresponding performance evaluation parameters includes:
[0038] Preprocessing the target model according to different preset segmentation ratios to obtain the training iteration time, model performance and resource utilization rate corresponding to the parallel computing accelerator and the single-row computing accelerator in each of the initial parallel training strategies;
[0039] Determine the model parallel efficiency corresponding to each of the initial parallel training strategies based on the training iteration time and model performance corresponding to the parallel computing accelerator and the single-row computing accelerator in each of the initial parallel training strategies;
[0040] The performance evaluation parameter is determined according to the model parallel efficiency and the resource utilization rate corresponding to each of the initial parallel training strategies.
[0041] On the other hand, the training iteration time is the time for forward propagation, backward propagation and parameter update of the target model; the process of determining the training iteration time includes:
[0042] Get the time of each iteration and the total number of iterations;
[0043] Determine the training iteration time according to the time of each iteration and the total number of iterations;
[0044] Correspondingly, the model performance is the accuracy and loss index performance of the target model in the validation set or the training set; the process of determining the model performance includes:
[0045] Obtaining model parameters and a validation set of the target model;
[0046] The model parameters and the validation set are evaluated according to a performance evaluation function to obtain the model performance.
[0047] On the other hand, the process of determining the parallel efficiency of the model includes:
[0048] Obtain a first training iteration time corresponding to the parallel computing accelerator and a second training iteration time corresponding to the single-row computing accelerator;
[0049] Obtaining a first model performance corresponding to a parallel computing accelerator and a second model performance corresponding to a single-row computing accelerator;
[0050] Determining a parallel iteration time according to a first training iteration time and the first model performance corresponding to the parallel computing accelerator;
[0051] Determine a single-row iteration time according to a second training iteration time and the second model performance corresponding to the single-row computing accelerator;
[0052] The model parallel efficiency is determined according to the parallel iteration time and the single-row iteration time.
[0053] On the other hand, the computing accelerator based on the parallel training strategy inputs the training data corresponding to the target sample quantity into the target model for training processing to determine the model parameters of the target model during the operation of the target model, including:
[0054] Get the current preset sample quantity;
[0055] Input the training data corresponding to the current preset number of samples into the target model;
[0056] During the operation of the target model, determining whether the operating memory exceeds the preset memory;
[0057] If the preset memory is exceeded, the next preset number of samples is obtained, and the process returns to the step of inputting the training data corresponding to the current preset number of samples into the target model, until the running memory corresponding to the training data corresponding to the current preset number of samples during the running of the target model does not exceed the preset memory; wherein the next preset number of samples is an integer multiple of the current preset number of samples;
[0058] If it does not exceed the preset memory, the current preset number of samples is used as the target number of samples, and the loss function of each segment is obtained by training according to the target number of samples; the convergence speed and loss fluctuation of the target model are determined according to the loss value of the loss function of each segment; if the convergence speed and loss fluctuation meet the second preset condition, the initial model parameters of the target model are adjusted in sequence to obtain the final model parameters.
[0059] On the other hand, there are multiple types of initial model parameters; and sequentially adjusting the initial model parameters of the target model to obtain the final model parameters includes:
[0060] After the current initial model parameters are adjusted, the current initial model parameters are applied to the target model, and applied to the target model to obtain the training test results of the target model;
[0061] If the training test result meets the third preset condition, the next initial model parameter is obtained as the current initial model parameter for adjustment, and the step of applying the current initial model parameter to the target model is returned to until the various initial model parameters of the target model are adjusted to obtain the final model parameters.
[0062] In order to solve the above technical problems, the present invention further provides a testing system, comprising:
[0063] Memory for storing computer programs;
[0064] A processor is used to implement the steps of the business processing method as described above when executing the computer program.
[0065] In order to solve the above technical problem, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the business processing method as described above are implemented.
[0066] In order to solve the above technical problem, the present invention also provides a computer program product, including a computer program / instruction, which implements the steps of the business processing method when executed by a processor.
[0067] The beneficial effects of the present invention are that, on the one hand, the performance parameters of the computing accelerators are processed by obtaining the preset topology structure composed of computing accelerators, weight parameters and each set communication strategy, and the target topology structure corresponding to the benchmark test is obtained. Before the actual business demand processing is performed, the corresponding preset topology structures are evaluated, and the target topology structure is obtained by fully considering the different performance parameters of each computing accelerator, so that the resources of each computing accelerator are fully utilized. On the other hand, based on the determined target topology structure, in the actual target model of the current business demand, the dynamic change characteristics of each computing accelerator in the model training test and the model parameter adjustment characteristics during the model operation are taken into account, and the target model is trained and processed according to the parallel training strategy corresponding to each computing accelerator and the target sample number of the target model to determine the model parameters. The target sample number parameter is used to realize the corresponding training speed, memory usage, training stability and other factors in the model parameter operation process, adjust the model parameters, and avoid the selection of initial parameters based on human experience. The present invention completes the adjustment of model parameters during operation, and there is no need for conventional models to adjust and return to iterate after the entire training process is completed, thereby reducing time consumption and improving the accuracy of model training parameters. At the same time, through the determination of the parallel training strategy of each computing accelerator, it is applied to the specific target model corresponding to the business needs, and the configuration parameters of the model that is truly suitable for different business needs are adapted, thereby improving the operation efficiency of the training process and the evaluation effect of the model training parameters corresponding to user needs.
[0068] Secondly, for the application test of the target model of business requirements, the parallel training model corresponding to the target model is first selected from multiple initial parallel training strategies, so that the parallel training strategy with balanced training time and model effect can be obtained by combining the target model under the target topology of the basic test. Under this parallel training strategy, the final model parameters corresponding to the target model operation process are determined by the maximum target sample number (MBS), which improves the model training speed, reduces the complexity and bias of manual parameter adjustment, and improves the authority. At the same time, it also improves the efficiency and effect of model training, so that users can find the optimal model parameter combination in a short time. Based on the evaluation scores determined by the performance parameters and weight parameters corresponding to different collective communication strategies, the target topology is determined according to the evaluation scores, and the comprehensive efficiency of the communication operation of the topology is improved by weighted scoring. Based on the evaluation scores determined by the performance parameters and weight parameters corresponding to different collective communication strategies, the throughput is determined by the ratio of corresponding bandwidth and delay to improve the accuracy of data judgment and improve data communication efficiency. During the benchmark test, the bandwidth test corresponding to the configuration information needs to meet the first preset condition to achieve the balance and expected performance of the preset topology before the collective communication test can be performed to improve the accuracy of the collective communication test. The determination process of topological balance within the connection domain and outside the connection domain within the topological structure is carried out to achieve the balance of the topological structure and the performance reaching the standard, so as to improve the accuracy and authority of the subsequent communication test. The performance evaluation parameters are determined by different preset split ratios, and then the final parallel training strategy is determined based on the performance evaluation parameters. The initial parallel training strategy under the performance evaluation parameters determined by different preset split ratios is traversed to avoid conventional human experience selection and improve the authority of data selection. At the same time, the final parallel training strategy is determined by the performance evaluation parameters to improve the accuracy of the actual performance of the training process of the topological structure. The performance evaluation parameters are determined based on three performance parameters, so that the accuracy is improved in the process of determining the final parallel training strategy to obtain a balance between training time and model effect. The determination process of training iteration time and model performance is to facilitate the subsequent calculation of model parallel efficiency and improve the accuracy of data verification. The determination process of model parallel efficiency is to find the parameter combination corresponding to the appropriate parallel training strategy. At the same time, the weighted sum is combined with the resource utilization rate to obtain the comprehensive scoring performance, so as to determine the parallel training strategy suitable for the current target model and reduce the training cost. The entire process of sequentially adjusting the initial model parameters of the target model to obtain the final model parameters. The currently adjusted initial model parameters are adjusted after the previous initial model parameters are adjusted, thereby improving the accuracy of model parameter adjustment and training efficiency.
[0069] In addition, the present invention also provides a testing system, a medium and a product, which have the same beneficial effects as the above-mentioned business processing method. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] In order to more clearly illustrate the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0071] Figure 1 A flowchart of a business processing method provided by an embodiment of the present invention;
[0072] Figure 2 A schematic diagram of a test node under a business requirement provided by an embodiment of the present invention;
[0073] Figure 3 A schematic diagram of an automated testing system architecture provided by an embodiment of the present invention;
[0074] Figure 4 A flow chart for determining a parallel training strategy provided by an embodiment of the present invention;
[0075] Figure 5 A benchmark test flow chart provided for an embodiment of the present invention;
[0076] Figure 6 A flow chart of adjusting model parameters of a target model provided by an embodiment of the present invention;
[0077] Figure 7 A structural diagram of a service processing device provided by an embodiment of the present invention;
[0078] Figure 8 A structural diagram of a test system provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0079] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0080] The core of the present invention is to provide a business processing method, a test system, a medium and a product to solve the problem that in the model training process corresponding to conventional user business needs, the resource characteristics of the computing accelerator are not fully utilized, resulting in reduced operating efficiency of the training process, and the model parameters are selected based on human experience and the adjustment process is based on the adjustment after the entire operation process is completed, resulting in a long time consumption and poor evaluation effect.
[0081] In order to enable those skilled in the art to better understand the scheme of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0082] In the field of high performance computing (HPC) and deep learning, computing accelerators are used in image processing, scientific computing, and training of large-scale machine learning models. As the complexity and scale of deep learning models continue to increase, the computing power of a single GPU is often unable to meet the demand, so multi-GPU parallel computing architectures have emerged. By assigning tasks to multiple GPUs for parallel processing, multi-GPU systems can significantly improve computing efficiency. However, performance testing and tuning of multi-GPU systems face a series of challenges. Conventional model operations fail to comprehensively evaluate the GPU's bandwidth, computing power, and continuously optimize training performance. Since training is time-consuming and costly, if the various parameters in the process are not correctly evaluated, it will affect the model's convergence speed and final accuracy, resulting in increased training costs.
[0083] Conventional training parameter evaluation relies on manual debugging, and the debugging process is a step-by-step test. It is necessary to select the initial parameter settings based on human experience, observe the test results after running the complete training process, and then make gradual adjustments based on the performance of the trained model, such as the loss value and iteration speed, which leads to more time-consuming and low efficiency, and is also affected by human factors. At the same time, it is also impossible to fully consider the utilization of hardware resources, such as the dynamic changes of GPU bandwidth and computing power, resulting in a waste of resources. In short, it limits the optimization speed of the model and the improvement of the overall performance. The business processing method provided by the present invention can solve the above technical problems.
[0084] Figure 1 A flowchart of a business processing method provided by an embodiment of the present invention, such as Figure 1 As shown, the method includes:
[0085] S11: Obtaining a target model corresponding to the current business demand, weight parameters of each set communication strategy, and a preset topology structure composed of each computing accelerator;
[0086] S12: Processing the performance parameters of each computing accelerator according to each preset topology structure, weight parameter and each collective communication strategy to determine a target topology structure corresponding to the benchmark test;
[0087] S13: Based on the target topology, the target model is trained according to the parallel training strategies corresponding to each computing accelerator and the target number of samples of the target model to obtain the model parameters of the target model during the operation of the target model to complete the processing of the current business needs.
[0088] The target model corresponding to the current business demand in step S11, it can be understood that the business demand is the model corresponding to the subsequent customized test case. The collective communication strategy is a series of algorithms used to implement data exchange and synchronization between multiple computing nodes (such as central processing unit (CPU) core, GPU, etc.), usually reflected in distributed computing and deep learning training. The number and type of collective communication strategies in this embodiment are not limited, and can be collective communication algorithms corresponding to application scenarios such as broadcast, gather, scatter, all gather, reduce, and all reduce.
[0089] Broadcast is a one-to-many communication mode, where one node sends data to all other nodes. In deep learning, it is often used to initialize parameters to ensure that all computing nodes start with the same parameter values. Gather is a many-to-one communication mode, where multiple nodes send data to one node. In distributed training, it can be used to collect updates from multiple nodes to a master node for processing. Disperse is a one-to-many communication mode, where one node distributes data to multiple nodes. In model parallelism, it can be used to send different parts of the model to different computing nodes. Gather All is a many-to-many communication mode, where each node sends data to all other nodes and receives data from all other nodes. In model parallelism, it can be used to synchronize the parameters of all nodes after forward computation. Reduce is a many-to-one communication mode, where multiple nodes send data to one node, which performs some operation on the data (such as summing, finding the maximum value, etc.). In distributed training, it can be used to aggregate gradient updates. Reduce All is a many-to-many communication mode, where each node sends data to all other nodes and receives data from all other nodes, and then performs a reduction operation on all the data. This is one of the most common collective communication operations in deep learning and is used to synchronize the gradients of all computing nodes.
[0090] The weight parameters of each collective communication strategy are different for different computing nodes corresponding to different business needs, and can be set according to actual conditions. The preset topology structure composed of various computing accelerators, the computing accelerator is a hardware device used to accelerate computing tasks, the type of computing accelerator is not limited, it can be GPU, FPGA, etc., or other accelerators, which are not limited here. The preset topology structure is the connection method and organizational form between the internal components of the accelerator, which can be a star topology, ring topology, tree topology, etc., which are not limited here.
[0091] In step S12, the performance parameters of each computing accelerator are processed according to each preset topological structure, weight parameter and each collective communication strategy, and the target topological structure corresponding to the benchmark test is determined. The performance parameters of the computing accelerator here mainly correspond to the connection bandwidth, delay and other parameters between different computing accelerators under the topological structure. Due to the different manufacturers of computing accelerators, the corresponding communication bandwidth and delay correspond to the realization of load balancing, etc., and it is necessary to adjust the corresponding target topological structure suitable for different business scenarios. The purpose of the benchmark test is to evaluate the communication capability, memory bandwidth and floating-point computing capability between different computing accelerators, and its performance indicators directly affect the overall efficiency of the parallel training system between different computing accelerators. It can automatically generate customized test cases according to different hardware configurations and application requirements, test different preset topological structures composed of each computing accelerator in an automated manner, count and analyze the collective communication bandwidth and delay, find the corresponding target topological structure, output test reports and tuning suggestions, and provide data support for subsequent training test performance tuning. The benchmark test in this embodiment is not only the basis for computing accelerator performance evaluation, but also a key component of the feedback mechanism in the training process, so as to facilitate the subsequent model optimization training strategy. It should be noted that the current benchmark test is a target topology structure determined corresponding to different topologies under the actual application scenarios of the target model, which fully considers the resource utilization of different computing accelerators and enables it to automatically adapt to the optimal training configuration according to the different computing accelerator manufacturers, hardware resources and training tasks.
[0092] In the benchmark test, the corresponding computing accelerator software stack will be included in the automatically generated virtualized test environment according to the different computing accelerator manufacturers, so as to support different computing accelerator tests on a set of systems.
[0093] Based on different preset topologies, the performance parameters and weight parameters of each computing accelerator corresponding to each collective communication strategy are summed up to obtain the evaluation scores under different preset topologies. For different evaluation scores, the preset topology with the largest score is determined as the target topology.
[0094] In step S13, under the target topology, the corresponding interconnection conditions and basic configuration performance between the corresponding computing accelerators under the target topology are known. This embodiment trains and tests the target model based on the selected parallel training strategy and the selected target sample quantity in each computing accelerator corresponding to the target topology, and obtains the model parameters of the target model under the current business needs. The parallel training strategy here is to divide a target model based on the corresponding preset ratio to multiple computing accelerators corresponding to the corresponding target topology. It is mainly aimed at the comprehensive parallel training strategy corresponding to the data parallelism, model parallelism or pipeline parallelism corresponding to the parallel training process of multiple computing accelerators. There are multiple parallel training strategies, and the appropriate parallel training strategy is selected from multiple preset parallel training strategies to achieve a balance between the training time and model effect of the target model. Corresponding to the selection of parallel training strategies, all preset parallel training strategies are performance evaluated in this embodiment to obtain the optimal parallel ratio.
[0095] The target sample number is the number of samples used to update the model parameters. It can be a micro-batch size (MBS). The number of samples used to update the model parameters in each iteration or training step can increase the training speed and utilize the parallel processing capabilities of the computing accelerator. It will increase the usage of memory and affect the stability of model training. At the same time, while increasing the number of samples, it will also pay attention to whether there is a memory overflow. While ensuring that the memory does not overflow, the MBS is maximized to adjust the model parameters of the subsequent target model.
[0096] It should be noted that the target sample number needs to be selected after each preset sample number to determine whether there is a memory overflow. In the absence of memory overflow, the stability of the model can be evaluated and determined by the convergence of the model. The specific convergence can be viewed through the loss function value, training set and validation set loss, overfitting or underfitting, etc., and there is no limitation here. As long as the preset conditions for convergence are met, the corresponding target sample number is also known. Under the target sample number, the initial model parameters of the target model are adjusted to determine the final model parameters. The model parameters here are adjusted during the operation of the target model. There is no need to adjust them after the current iteration ends. They are adjusted while running. First, one model parameter is adjusted, and then the next model parameter is adjusted after the adjustment is completed.
[0097] Based on the automated benchmark test, the automated training test generates customized test cases and collects and analyzes key performance parameters such as single iteration time, GPU floating point operations per second (TFLOPS), loss value change rate, communication bandwidth occupancy, and expansion ratio improvement during the training process. The performance evaluation and tuning method is used to optimize the training parameters in multiple training iterations to achieve closed-loop optimization. In multiple automated training and testing cycles, this feedback mechanism can gradually improve the training efficiency of the model, find the best training parameters, and output tuning suggestions.
[0098] Figure 2 A schematic diagram of a test node under a business requirement provided by an embodiment of the present invention, such as Figure 2 As shown in the figure, a decoupled design is adopted, which includes five nodes: user node, automated test control node, large model node, test case library and test node. The user node is the interactive entrance between the user and the system, responsible for issuing instructions, evaluating and adjusting test scripts, viewing test results and confirming operations. The automated test control node is the core node of the entire system, responsible for running the automated test process. It receives instructions from the user node and calls the pre-trained model node and test case library for test operations as needed. The pre-trained model node is a locally running node, which usually deploys a pre-trained model with a large parameter scale and is fine-tuned based on the knowledge related to computing accelerator testing to better handle the inference request sent by the automated test control node. In addition, the pre-trained model node can also interact directly with the user node to provide more intelligent support. The test case library collects a variety of test scripts for different hardware architectures, computing accelerator (such as GPU) models and different software versions, including officially released scripts, community contributed scripts and optimized scripts. These scripts are stored in the database for the automated test control node to call and use. The test node is a GPU node used to perform actual testing. Its main responsibility is to run the test scripts assigned by the automated test control node. It usually does not have other additional functions and only communicates with the automated test control node.
[0099] Figure 3 A schematic diagram of an automated testing system architecture provided by an embodiment of the present invention is shown in FIG. Figure 3 As shown in the figure, the startup script first collects system information such as host topology, computing accelerator topology, and network topology through the system interface, and receives relevant information such as model architecture, training data, training framework, and preset parameters input by the user. After processing, various types of information are combined into a standardized text file in JSON format to ensure data standardization and compatibility. Subsequently, the system automatically generates a virtualized test environment and benchmark test cases based on this data.
[0100] In the generated virtualized test environment, the corresponding GPU driver, GPU programming library, collective communication library, and a series of general test scripts provided by the manufacturer will be selected according to the GPU manufacturer to achieve rapid deployment and removal in the test environment. The benchmark test cases are designed to evaluate key indicators such as communication bandwidth, latency, and computing performance under the existing GPU topology. The automated test system is equipped with a dedicated test script generator, which can generate multiple benchmark test scripts with different resource ratios based on the host topology, GPU topology, and network topology of the current system. In addition, the test script generator also has the ability to call the locally running pre-trained model for script optimization and complex script writing, further improving the accuracy and coverage of the test. The generated benchmark test cases and virtualized test environment will be confirmed through report feedback, and a pre-trained model interaction method will be provided to re-optimize the test script. The benchmark test cases and virtualized test environment generated by the automated test system will be fed back to the user through reports for final confirmation to ensure their effectiveness and pertinence. Subsequently, the automated benchmark test begins to iterate, and the iteration process follows the benchmark test process. When the automated benchmark test is completed, the test system will output a test report and topology optimization suggestions. The test report records in detail the adjustments made by the test system to the topology and the allocation of GPU resources based on each test result, as well as the test results after each iteration. Users will decide which topology to use for training based on the topology optimization recommendations.
[0101] Automated training tests will be conducted based on the hardware and software topologies determined in automated benchmark tests. First, the system will match the most appropriate test cases in the local test case library based on the model architecture, training framework, and preset parameter settings provided by the user. Then, the test script generator will optimize and adjust these test cases based on the topological information obtained in the benchmark test. During the optimization process, key parameters such as the number of GPUs used in parallel training, the parallel computing ratio, and the batch size (MBS) are usually adjusted according to the specific hardware configuration. This is to maximize the utilization of hardware resources during training tests, thereby improving the convergence speed and effect of model training. Unlike benchmark tests, training tests are often more complex. In addition to considering the utilization of GPUs during training, the performance of the converged model must also be fully evaluated. In addition, for specific computing nodes or cluster network environments, automated training tests will adjust data distribution strategies, memory allocation methods, and communication protocols, and match corresponding drivers to ensure that the model can achieve the best training performance in the target environment. In addition, considering that user-defined training corpora are often not in a standardized text format, the automated training test system supports secondary processing of training corpora using local pre-trained models. The system can standardize the format of the corpus, or expand and enrich the corpus based on existing content to ensure that it meets the training requirements.
[0102] After multiple training and testing iterations, the automated training and testing system will conduct a comprehensive test on the training parameter combination of the current topology and model architecture, evaluate the optimal parameter combination, and conduct an in-depth analysis of the statistical results of each training iteration. Finally, these evaluation and analysis results will be presented to the user in a visual form.
[0103] The embodiment of the present invention provides a business processing method, which obtains a target model corresponding to the current business demand, weight parameters of each collective communication strategy, and a preset topology structure composed of each computing accelerator; processes the performance parameters of each computing accelerator corresponding to each preset topology structure, weight parameters, and each collective communication strategy to determine the target topology structure corresponding to the benchmark test; based on the target topology structure, the target model is trained and processed according to the parallel training strategy corresponding to each computing accelerator and the target sample number of the target model to obtain the model parameters of the target model during the operation of the target model to complete the processing of the current business demand. On the one hand, the preset topology structure composed of the computing accelerator, weight parameters, and each collective communication strategy are obtained to process the performance parameters of the computing accelerator to obtain the target topology structure corresponding to the benchmark test, and before the actual business demand processing is performed, each preset topology structure is evaluated, and the target topology structure is obtained by fully considering the different performance parameters of each computing accelerator, so that the resources of each computing accelerator are fully utilized. On the other hand, based on the determined target topology, in the actual target model of the current business needs, taking into account the dynamic change characteristics of each computing accelerator in the model training test and the model parameter adjustment characteristics during the model operation process, the target model is trained and processed to determine the model parameters according to the parallel training strategy corresponding to each computing accelerator and the target sample number of the target model. With the help of the parameter of the target sample number, the training speed, memory usage, training stability and other factors corresponding to the model parameter operation process are realized, and the model parameters are adjusted to avoid the selection of initial parameters based on human experience. The present invention completes the adjustment of model parameters during operation, and does not require the conventional model to adjust and return to iteration after the entire training process is completed, which reduces time consumption and improves the accuracy of model training parameters. At the same time, through the determination of the parallel training strategy of each computing accelerator, it is applied to the specific target model corresponding to the business needs, and the configuration parameters of the model that is truly suitable for different business needs are adapted, thereby improving the operation efficiency of the training process and the evaluation effect of the model training parameters corresponding to user needs.
[0104] In some embodiments, based on the target topology, the target model is trained according to the parallel training strategy corresponding to each computing accelerator and the target number of samples of the target model to obtain the model parameters of the target model during the operation of the target model, including:
[0105] Obtaining the initial parallel training strategy corresponding to each computing accelerator under the target topology structure;
[0106] Determine the parallel training strategy corresponding to the target model of the current business requirement within each initial parallel training strategy;
[0107] A computing accelerator based on a parallel training strategy inputs training data corresponding to a target number of samples into a target model for training processing to determine model parameters of the target model during the operation of the target model.
[0108] Specifically, for different types of models and different application scenarios, users may pay more attention to the timeliness of training or the actual effect of training models, and adjust relevant parameters according to the training test based on this information. Different initial parallel training strategies corresponding to each computing accelerator under the target topology structure, for example, a target model under different computing accelerators determines multiple initial parallel training strategies according to different preset ratios, and determines the final parallel training strategy corresponding to the target model of the current business needs under multiple initial parallel training strategies. Based on the computing accelerator of the parallel training strategy, select the training data corresponding to the target sample quantity and input it into the target model for model training testing, and adjust the model parameters during the model operation to meet the current business needs.
[0109] The initial parallel training strategy here can be data parallelism, model parallelism or pipeline parallelism, etc., which is not limited here. Compared with the selection of conventional parallel training strategies, the frequently used parallel training strategies selected by human experience are adopted. In this embodiment, all initial parallel training strategies need to be traversed to find the final parallel training strategy to improve authority.
[0110] Figure 4 A flow chart for determining a parallel training strategy provided by an embodiment of the present invention, such as Figure 4 As shown, the steps include:
[0111] S21: Setting the target topology structure and target model information;
[0112] S22: Determine whether the initial parallel training strategy has been exhausted; if so, proceed to step S26; if not, proceed to step S23;
[0113] S23: Initial parallel training strategy adjustment for training;
[0114] S24: perform training test;
[0115] S25: Parallel training performance evaluation, and return to step S22;
[0116] S26: Based on the computing accelerator of the parallel training strategy, the training data corresponding to the target number of samples is input into the target model for training processing to determine the model parameters of the target model during the operation of the target model.
[0117] The application test of the target model of the business demand provided in this embodiment first selects the parallel training model corresponding to the target model from multiple initial parallel training strategies, so that a parallel training strategy that balances the training time and model effect is obtained in combination with the target model under the target topology structure of the basic test. Under this parallel training strategy, the final model parameters corresponding to the target model operation process are determined by maximizing the target sample number (MBS), thereby improving the model training speed, reducing the complexity and bias of manual parameter adjustment, and improving the authority. At the same time, the efficiency and effect of model training are also improved, so that users can find the optimal model parameter combination in a short time.
[0118] In some embodiments, determining a target topology structure according to each preset topology structure, a weight parameter, and a performance parameter of each computing accelerator corresponding to each collective communication strategy includes:
[0119] Processing each collective communication strategy in each preset topology structure to obtain corresponding performance parameters;
[0120] Determine the evaluation score of each preset topology structure according to the performance parameters and weight parameters of each collective communication strategy;
[0121] The target topology is determined according to the evaluation scores of the preset topologies.
[0122] Specifically, the collective communication strategy affects the efficiency of data transmission between multiple computing nodes, and usually requires efficient use of available bandwidth to ensure the speed of data transmission. The corresponding delay parameter, the transmission time from the source to the destination, needs to minimize the delay to improve communication efficiency.
[0123] According to each collective communication strategy, the corresponding performance parameters are processed in each preset topology structure, and the corresponding evaluation scores under each preset topology structure are obtained by weighting based on different performance parameters and weight parameters set for corresponding services. The evaluation scores under different preset topologies are arranged in order, and the preset topology structure corresponding to the maximum evaluation score is used as the target topology structure.
[0124] For example, there are two preset topologies (A and B), with a total of four collective communication strategies. Under the preset topology A, the first collective communication strategy is used to obtain the corresponding performance parameter A1, the second collective communication strategy is used to obtain the corresponding performance parameter A2, the third collective communication strategy is used to obtain the corresponding performance parameter A3, and the fourth collective communication strategy is used to obtain the corresponding performance parameter A4. The evaluation score corresponding to the preset topology A = A1*a1+A2*a2+A3*a3+A4*a4, where a1, a2, a3, and a4 correspond to weight parameters. By comparing the evaluation scores corresponding to the two preset topologies, the preset topology corresponding to the largest evaluation score is selected as the target topology.
[0125] This embodiment provides an evaluation score determined based on performance parameters and weight parameters corresponding to different set communication strategies, determines a target topology structure according to the evaluation score, and improves the overall efficiency of communication operations of the topology structure by weighted scoring.
[0126] In some embodiments, the performance parameters include at least bandwidth and delay, and the evaluation score of each preset topology is determined according to the performance parameters and weight parameters of each collective communication strategy, including:
[0127] Determine the throughput corresponding to each collective communication strategy in the current preset topology structure according to the bandwidth corresponding to each collective communication strategy and the delay corresponding to each collective communication strategy;
[0128] The weight parameters and throughput corresponding to each collective communication strategy are used to determine the evaluation score of the current preset topology.
[0129] It is understandable that the performance parameter is used to characterize the communication capability of the preset topology, and may be a bandwidth and delay parameter, or may include other parameters, which are not limited here. In this embodiment, the bandwidth and delay parameters under each collective communication strategy are divided to obtain the throughput corresponding to the current preset topology, and the weight parameter and the throughput are multiplied to determine the evaluation score of the current preset topology.
[0130] For example, All-Reduce is the most commonly used collective communication algorithm in many deep learning and scientific computing tasks. It is usually used for operations such as global gradient update and model averaging, so it has the highest weight. All-Gather has the second highest weight because it is often used to collect partial results on each GPU in data parallel training. According to different weight allocations, the collective communication performance is scored S, and the formula is as follows:
[0131] ;
[0132] in, Weight parameters corresponding to different collective communication strategies; , The bandwidth parameters of the preset topology corresponding to different collective communication strategies, , Corresponding to the delay parameters of the preset topology under different collective communication strategies, the throughput is the ratio of the bandwidth parameter to the delay parameter. The larger the throughput, the higher the evaluation bandwidth of the collective communication strategy and the lower the delay, indicating that the overall efficiency of the communication operation is higher.
[0133] This embodiment provides an evaluation score determined based on performance parameters and weight parameters corresponding to different collective communication strategies, and determines the throughput by corresponding to different bandwidth and delay ratios, so as to improve the accuracy of data judgment and the efficiency of data communication.
[0134] In some embodiments, before determining the target topology according to each preset topology, weight parameter and performance parameter of each computing accelerator corresponding to each collective communication strategy, the method further includes:
[0135] Determining the interconnection status of each computing accelerator in each preset topology structure;
[0136] When the interconnection status is normal, configure the link configuration information corresponding to the preset topology structure;
[0137] Testing the bandwidth between the computing accelerators in each preset topology structure and between the computing accelerator and the processor to obtain a first test result;
[0138] If the first test result satisfies the first preset condition, then proceeding to the step of determining the target topology structure according to each preset topology structure, weight parameter and performance parameter of each computing accelerator corresponding to each collective communication strategy;
[0139] If the first test result does not meet the first preset condition, the process returns to the step of configuring the link configuration information corresponding to the preset topology structure, and adjusts the link configuration information until the first test result meets the first preset condition.
[0140] Specifically, it is necessary to check all computing accelerators and interconnection status, and configure the link configuration information corresponding to the preset topology structure according to the type of computing accelerator and the communication bus form corresponding to the interconnection status. The interconnection status here can be Peripheral Component Interconnect Express (PCIE) and other manufacturer-specific communication buses. When the interconnection status is normal, the preset topology structure of the interconnection of computing accelerators is configured through the default configuration. It should be noted that the preset topology structure also involves the configuration of the port and structure of the switch (Switch), such as the number of lines and the rate corresponding to each link. After configuration, the bandwidth test of the peer-to-peer (P2P) between each computing accelerator and the host-to-device (H2D) between the processor and each computing accelerator is performed to obtain the first test result. If the first test result meets the first preset condition, the subsequent steps are entered (determining the target topology structure according to the preset topology structure, weight parameters and the performance parameters of each computing accelerator corresponding to each collective communication strategy). The purpose of the judgment process of whether the first test result meets the first preset condition here is to detect whether the topology has achieved balance and expected performance. If the first test result does not meet the first preset condition, it is necessary to readjust the interconnection parameters (link configuration information) and retest to determine that the bandwidth is balanced and the performance meets the standard before performing the collective communication test.
[0141] Figure 5 A benchmark test flow chart provided by an embodiment of the present invention, such as Figure 5 As shown, including:
[0142] S31: Check the interconnection status of the computing accelerator;
[0143] S32: Determine whether the topology has been exhausted; if so, proceed to step S33; if not, proceed to step S34;
[0144] S33: output test records and target topology;
[0145] S34: configuring link configuration information corresponding to the interconnection state of the computing accelerator;
[0146] S35: Execute bandwidth tests between computing accelerators in each preset topology structure and between each computing accelerator and the processor to obtain a first test result;
[0147] S36: Determine whether the first test result achieves topological balance; if so, proceed to step S37; if not, proceed to step S34;
[0148] S37: Determine whether the bandwidth is sufficient; if not, proceed to step S38; if yes, proceed to step S39;
[0149] S38: adjusting the link configuration parameters of the interconnection state;
[0150] S39: Perform collective communication test;
[0151] S40: record the test results;
[0152] S41: Evaluate and sort the test results according to the rules; and return to step S32.
[0153] Wherein, step S40 and step S41 correspond to the process of determining and sorting the evaluation scores of each preset topological structure.
[0154] In the benchmark test process provided in this embodiment, the bandwidth test of the corresponding configuration information needs to meet the first preset condition to achieve the balance and expected performance of the preset topology structure, and then the collective communication test can be performed to improve the accuracy of the collective communication test.
[0155] In some embodiments, the first preset condition is a first sub-preset condition or a second sub-preset condition, the first sub-preset condition is a preset condition for bandwidth and delay balance of a topological structure within a connection domain; the second sub-preset condition is a preset condition for bandwidth and delay balance of a topological structure outside a connection domain, and the determination process of the first preset condition includes:
[0156] When each computing accelerator is in a topological structure within the connection domain, obtaining a first bandwidth standard deviation and a first delay standard deviation between the computing accelerators in the topological structure within the connection domain;
[0157] Obtaining a grouping method corresponding to computing accelerators in a topological structure within a connection domain;
[0158] If the number of groups corresponding to the grouping mode is 0, determine whether the first bandwidth standard deviation is less than or equal to the first bandwidth threshold and whether the first delay standard deviation is less than or equal to the first delay threshold; if the first bandwidth standard deviation is less than or equal to the first bandwidth threshold and the first delay standard deviation is less than or equal to the first delay threshold, determine that the first sub-preset condition is met; if the first bandwidth standard deviation is less than or equal to the first bandwidth threshold or the first delay standard deviation is less than or equal to the first delay threshold, determine that the first sub-preset condition is not met;
[0159] If the number of groups corresponding to the grouping mode is greater than 0, determine whether the first bandwidth standard deviation is less than or equal to the second bandwidth threshold and whether the first delay standard deviation is less than or equal to the second delay threshold; if the first bandwidth standard deviation is less than or equal to the second bandwidth threshold and the first delay standard deviation is less than or equal to the second delay threshold, determine that the first sub-preset condition is met; if the first bandwidth standard deviation is less than or equal to the second bandwidth threshold or the first delay standard deviation is less than or equal to the second delay threshold, determine that the first sub-preset condition is not met;
[0160] When each computing accelerator is in a topology structure outside the connection domain, obtaining a second bandwidth standard deviation and a second delay standard deviation between computing accelerators in the topology structure outside the connection domain and across the connection domain;
[0161] Determine whether the second bandwidth standard deviation is less than or equal to the third bandwidth threshold and whether the second delay standard deviation is less than or equal to the third delay threshold; if the second bandwidth standard deviation is less than or equal to the third bandwidth threshold and the second delay standard deviation is less than or equal to the third delay threshold, determine that the second sub-preset condition is met; if the second bandwidth standard deviation is less than or equal to the third bandwidth threshold or the second delay standard deviation is less than or equal to the third delay threshold, determine that the second sub-preset condition is not met.
[0162] Specifically, the distinction between the topology within the connection domain and the topology outside the connection domain is usually defined based on the network topology and communication protocol. In network communication, the topology within the connection domain refers to the devices or nodes in the network that are connected using communication protocols and bandwidth, while the topology outside the connection domain refers to other devices or nodes that are not within this communication range.
[0163] The connection domain is mainly a high-speed connection domain (such as PCIE, etc.), and its judgment is based on:
[0164] Bandwidth balance: Use standard deviation to evaluate the uniformity of P2P bidirectional bandwidth distribution. The standard deviation of all P2P bandwidth values should not exceed the threshold of 5%.
[0165] Delay balancing: Use standard deviation to evaluate the uniformity of P2P delay distribution. The standard deviation of all P2P delays should not exceed the threshold of 10%.
[0166] Group symmetry: GPUs are grouped in even powers of 2, such as 4, 8, or 16 GPUs per group. The standard deviation of GPU P2P bandwidth and latency within the same group should not exceed the threshold of 1%.
[0167] That is, obtain the corresponding first bandwidth standard deviation and first delay standard deviation. How to calculate the standard deviations corresponding to the two performance parameters can be the same as the conventional standard deviation calculation method, or it can be different, and it is not limited here. It is necessary to check the number of groups corresponding to the grouping method corresponding to the current computing accelerator. If the number of groups is 0, you only need to check the bandwidth balance and delay balance, that is, if the first bandwidth standard deviation is less than or equal to the first bandwidth threshold and the first delay standard deviation is less than or equal to the first delay threshold, it is determined that the first sub-preset condition is met. If the first bandwidth standard deviation is less than or equal to the first bandwidth threshold or the first delay standard deviation is less than or equal to the first delay threshold, it is determined that the first sub-preset condition is not met, and the topological balance is not met.
[0168] If the number of groups is greater than 0 and the groups are symmetrical, the number of groups of the computing accelerator is an even power of 2, and it is necessary to determine whether the first bandwidth standard deviation is less than or equal to the second bandwidth threshold, and whether the first delay standard deviation is less than or equal to the second delay threshold. If so, it means that the group symmetry is met. It should be noted that when the second bandwidth threshold and the second delay threshold are met, the first bandwidth threshold and the first delay threshold mentioned above are also met. The limiting condition here is that the first bandwidth threshold is greater than the second bandwidth threshold, and the first delay threshold is greater than the second delay threshold.
[0169] When in a topology outside the connection domain, the balance is determined based on:
[0170] Bandwidth balance: Use standard deviation to evaluate the uniformity of P2P bidirectional bandwidth distribution. The standard deviation of all P2P bandwidth values across high-speed interconnection domains should not exceed the threshold of 10%;
[0171] Latency balancing: The standard deviation is used to evaluate the uniformity of P2P latency distribution. The standard deviation of all P2P latency across high-speed interconnection domains should not exceed a threshold of 20%.
[0172] Specifically, the second bandwidth standard deviation and the second delay standard deviation between the computing accelerators are obtained to determine whether the second bandwidth standard deviation is less than or equal to the third bandwidth threshold and whether the second delay standard deviation is less than or equal to the third delay threshold; if so, the topological balance that satisfies the bandwidth balance and the delay balance is determined; if not, that is, only one balance is satisfied, it means that the topological balance cannot be satisfied, that is, the second sub-preset condition is not satisfied.
[0173] This embodiment provides a determination process for topology balance within a connection domain and topology balance outside a connection domain within a topology structure, so as to achieve a balanced topology structure and meet performance standards, thereby improving the accuracy and authority of subsequent communication tests.
[0174] In some embodiments, determining the parallel training strategy corresponding to the target model of the current business requirement within each initial parallel training strategy includes:
[0175] Obtaining a preset ratio of segmentation of the target model on each computing accelerator;
[0176] Preprocess the target model according to different preset segmentation ratios to determine the performance evaluation parameters corresponding to each initial parallel training strategy;
[0177] The parallel training strategy is determined based on the performance evaluation parameters.
[0178] Specifically, the preset ratio of segmentation corresponding to this embodiment is not limited here, and different ratios can be allocated for tensor parallelism, pipeline parallelism, and data parallelism. For example, the total number of GPUs is 16, and the test will evaluate the acceleration ratio of GPU training when tensor parallelism, pipeline parallelism, and data parallelism (TP|PP|DP) are allocated in different ratios such as 1|2|8, 1|4|4, 4|4|1, and 4|1|4, so as to find the most suitable parameter combination. The performance evaluation parameters here are intended to achieve a balance between training time and model effect. The performance evaluation parameters can be one or more. When there is only one performance evaluation parameter, the corresponding parallel training strategy is determined according to the size of the parameter. If there are multiple performance evaluation parameters, the multiple performance evaluation parameters are weighted and summed to obtain a score, and the parallel training strategy is determined based on the size of the score. A comprehensive evaluation can also be performed in other ways to determine the final parallel training strategy, which is not limited here.
[0179] The present embodiment provides performance evaluation parameters determined by different preset split ratios, and then determines the final parallel training strategy based on the performance evaluation parameters. By traversing the initial parallel training strategy under the performance evaluation parameters determined by different preset split ratios, conventional human experience selection is avoided, and the authority of data selection is improved. At the same time, the final parallel training strategy is determined by the performance evaluation parameters to improve the accuracy of the actual performance of the training process of the topological structure.
[0180] In some embodiments, the performance evaluation parameters include at least training iteration time, model performance, and resource utilization; the target model is preprocessed according to different preset segmentation ratios to determine the corresponding performance evaluation parameters, including:
[0181] The target model is preprocessed according to different preset split ratios to obtain the training iteration time, model performance and resource utilization rate corresponding to the parallel computing accelerator and the single-row computing accelerator in each initial parallel training strategy;
[0182] The model parallel efficiency corresponding to each initial parallel training strategy is determined by the training iteration time and model performance corresponding to the parallel computing accelerator and the single-row computing accelerator in each initial parallel training strategy;
[0183] The performance evaluation parameters are determined according to the model parallel efficiency and resource utilization corresponding to each initial parallel training strategy.
[0184] Specifically, the training iteration time can be determined based on factors such as the complexity of the model, the size of the data set, the hardware resources used, and the optimization algorithm. Model performance, mainly the performance ability on a specific task, is measured by indicators. Resource utilization can be the utilization of the computing accelerator, memory usage, and communication overhead, etc. The type of performance evaluation parameters of this embodiment may also include other parameters, which are not limited here. The target model is preprocessed according to different preset segmentation ratios to obtain performance evaluation parameters. The preprocessing process here can be a training test performed during the operation of the target model, so as to detect the corresponding effects of different initial parallel training strategies through performance evaluation.
[0185] Training iteration time and model performance affect model parallel efficiency, so here we need to use the acceleration ratio of parallel and single-row training in multiple computing accelerators to evaluate model performance. The performance evaluation parameters are obtained by weighting the model parallel efficiency and resource utilization, and taking the corresponding weighted sum.
[0186] The present embodiment provides a method for determining performance evaluation parameters based on three performance parameters, so as to improve accuracy in determining the final parallel training strategy, so as to achieve a balance between training time and model effect.
[0187] In some embodiments, the training iteration time is the time for forward propagation, backward propagation and parameter update of the target model; the process of determining the training iteration time includes:
[0188] Get the time of each iteration and the total number of iterations;
[0189] Determine the training iteration time based on the time of each iteration and the total number of iterations;
[0190] Correspondingly, the model performance is the accuracy and loss performance of the target model in the validation set or training set. The process of determining the model performance includes:
[0191] Get the model parameters and validation set of the target model;
[0192] The model parameters and validation set are evaluated according to the performance evaluation function to obtain the model performance.
[0193] Specifically, regarding the training iteration time, the forward propagation is the propagation of input data through the network layer until the output layer, generating a prediction result. The backward propagation is to use the gradient information of the loss function to calculate the gradient of each parameter through the back propagation algorithm. The parameter update uses the gradient calculated in the backward propagation to update the parameters of the model according to the optimization algorithm. The formula is as follows:
[0194] ;
[0195] in, is the average time per iteration, is the total number of iterations, It is The time taken for the iteration.
[0196] Model performance refers to the accuracy, loss, or other specific indicators (such as F1-score, area under curve (AUC) etc.) of the model on the validation set or test set. It can be tested through the iterative version of the model. The formula is as follows:
[0197] ;
[0198] in, is the model performance, are model parameters, is the validation dataset, is the performance evaluation parameter of the model on the validation set.
[0199] This embodiment provides a process for determining the training iteration time and model performance to facilitate the subsequent calculation of model parallel efficiency and improve the accuracy of data verification.
[0200] In some embodiments, the process of determining the model parallel efficiency includes:
[0201] Obtain a first training iteration time corresponding to the parallel computing accelerator and a second training iteration time corresponding to the single-row computing accelerator;
[0202] Obtaining a first model performance corresponding to a parallel computing accelerator and a second model performance corresponding to a single-row computing accelerator;
[0203] Determine the parallel iteration time according to the first training iteration time and the first model performance corresponding to the parallel computing accelerator;
[0204] Determine a single-row iteration time according to a second training iteration time corresponding to the single-row computing accelerator and a second model performance;
[0205] The model parallel efficiency is determined based on the parallel iteration time and the single row iteration time.
[0206] Specifically, the model parallel efficiency is obtained through the above two formulas, which represents the acceleration ratio of the model on multiple computing accelerators after parallelization compared with single computing accelerator training. It is used to evaluate the performance of the model. The formula is as follows:
[0207] ;
[0208] in, is the model parallel efficiency, is the second training iteration time corresponding to a single-row computing accelerator, is the first training iteration time corresponding to the parallel computing accelerator, is the second model performance corresponding to a single-row computing accelerator, It is the first model performance corresponding to the parallel computing accelerator.
[0209] Automated model testing will configure multiple parallel modes, allocate different split preset ratios, and corresponding acceleration ratios to find the appropriate parameter combination.
[0210] During the parallel training process, resource usage includes computing accelerator utilization, memory usage, and communication overhead. Automated model testing uses a variety of monitoring tools (generally detection tools provided by computing accelerator manufacturers and computing accelerator performance analysis tools, etc.) to monitor the utilization of each computing accelerator in real time. The definition of resource utilization is relatively flexible and is not restricted here. Taking Uresource as the resource utilization, it can be weighted to obtain a comprehensive score with parallel efficiency, which is used to comprehensively evaluate the cost of the training process.
[0211] The model parallel efficiency determination process provided in this embodiment is to facilitate finding a suitable parameter combination corresponding to a parallel training strategy. At the same time, a weighted sum is performed in combination with resource utilization to obtain a comprehensive scoring performance, so as to facilitate determining a parallel training strategy suitable for the current target model and reduce training costs.
[0212] In some embodiments, a computing accelerator based on a parallel strategy inputs training data corresponding to a target number of samples into a target model for training processing to determine model parameters of the target model during the operation of the target model, including:
[0213] Get the current preset sample quantity;
[0214] Input the training data corresponding to the current preset number of samples into the target model;
[0215] During the operation of the target model, determine whether the operation memory exceeds the preset memory;
[0216] If the preset memory is exceeded, the next preset number of samples is obtained, and the process returns to the step of inputting the training data corresponding to the current preset number of samples into the target model, until the running memory corresponding to the training data corresponding to the current preset number of samples during the running of the target model does not exceed the preset memory; wherein the next preset number of samples is an integer multiple of the current preset number of samples;
[0217] If the preset memory is not exceeded, the current preset number of samples is used as the target number of samples, and the loss function of each segment is obtained by training according to the target number of samples; the convergence speed and loss fluctuation of the target model are determined according to the loss value of the loss function of each segment; if the convergence speed and loss fluctuation meet the second preset condition, the initial model parameters of the target model are adjusted in sequence to obtain the final model parameters.
[0218] Specifically, first obtain a current preset number of samples, which is MBS data. Micro-batch size is a concept in deep learning. It refers to the number of data samples passed to the model for forward propagation and backward propagation each time during training. Micro-batch size is a strategy between stochastic gradient descent (SGD) for a single sample and batch gradient descent for the entire training set. In practical applications, micro-batch size is usually a fraction of mini-batch size, which refers to the number of samples used to update weights each iteration in model training. The use of micro-batch size can improve memory usage efficiency, especially in memory-constrained environments. For example, if a model is too large or the data batch is too large to fit into the memory of a single device, it can be streamed to the device by further splitting the mini-batch data into micro-batch data, allowing training with a batch size larger than the device memory capacity. The choice of micro-batch size affects both the training efficiency and the final performance of the model. A smaller micro-batch size can improve the stability of training and help the model escape from the local minimum, but may result in longer training time because only a small amount of data is processed each time. Larger micro-batch sizes can reduce training time because the parallel processing capabilities of computing accelerators can be more effectively utilized, but may increase the risk of the model getting stuck in local optimality and have higher memory requirements.
[0219] Input the training data corresponding to the current preset number of samples into the target model, and check whether there is a memory overflow during the operation of the target model. If so, it means that the current preset number of samples is large, and then get the next preset number of samples. The relationship between the next preset number of samples and the current preset number of samples is an integer multiple relationship. For example, 32, 64, 128, etc. Based on the next preset number of samples as the new current preset number of samples, continue to judge. If it does not exceed the preset memory, the current preset number of samples is used as the target number of samples, and the loss function of each segment is obtained according to the target number of samples. It can be understood that if the memory does not overflow, it means that the training stability evaluation stage of the target model has been entered, and it is necessary to evaluate the convergence and performance stability under the current preset number of samples. The purpose is to ensure that the model is trained with optimal computing performance within the scope allowed by hardware resources.
[0220] The loss value of the loss function determines the convergence speed and loss fluctuation. In some embodiments, the process of determining the convergence speed includes:
[0221] Get the loss value of the current segment and the loss value of the previous segment;
[0222] Performing a difference process on the loss value of the previous segment and the loss value of the current segment to obtain a first loss difference;
[0223] Divide the first loss difference by the loss value of the previous segment to obtain the convergence speed;
[0224] Correspondingly, the process of determining loss fluctuations includes:
[0225] Get the number of measurements and the corresponding loss mean values for all segments;
[0226] The loss value corresponding to the current segment is processed with the loss mean to obtain the second loss difference of the current segment;
[0227] The third loss difference is obtained by summing the squares of the second loss differences of all segments;
[0228] Dividing the third loss difference by the number of measurements to obtain a fourth loss difference;
[0229] The loss fluctuation is obtained by performing square root processing on the fourth loss difference.
[0230] Specifically, its convergence speed formula is as follows:
[0231] ;
[0232] in, is the loss value of the n-1th segment, is the loss value of the nth segment, is the convergence speed. If If it remains close to 0 for multiple segments, it indicates that the target model may have converged.
[0233] The loss volatility formula is as follows:
[0234] ;
[0235] in, is the p-th loss value, is the mean of the loss values (mean loss), is the p-th second loss difference, M is the number of measurements, For loss fluctuations.
[0236] The training stability evaluation will take into account the convergence speed and loss volatility. When the memory does not overflow, the evaluation basis of the automated model test is that when the Micro Batch Size increases exponentially, the convergence speed and loss volatility may increase or decrease.
[0237] The determination process of the corresponding convergence speed and loss fluctuation in this embodiment is used to measure the stability of the corresponding target model under the condition of MBS growth, and provides a basis for judgment to ensure the optimization of the calculation performance of the target model.
[0238] When the convergence speed and loss fluctuation meet the second preset condition, it means that the convergence speed is reduced and the increase of the loss fluctuation is reduced. The second preset condition can set the threshold range corresponding to the convergence speed and loss fluctuation, or set specific parameters, which are not limited here. For example, the decrease of the convergence speed and the increase of the loss volatility should be less than 10%.
[0239] After determining the MBS, the target model test enters the key parameter selection stage. By adjusting different key parameters such as learning rates, optimizer types, regularization methods, and mixed precision, multiple iterative trainings are performed to select the optimal parameter combination.
[0240] It should be noted that the adjustment process in this embodiment is during the operation of the target model, and it is not the adjustment of all model parameters after each iteration. Instead, during the operation of the target model, one model parameter is adjusted first, and test training is carried out. If the test result is qualified, the next model parameter is adjusted until all adjustments are completed.
[0241] In some embodiments, there are multiple types of initial model parameters; adjusting the initial model parameters of the target model in sequence to obtain the final model parameters includes:
[0242] After the current initial model parameters are adjusted, the current initial model parameters are applied to the target model, and applied to the target model to obtain the training and testing results of the target model;
[0243] If the training test results meet the third preset condition, the next initial model parameter is obtained as the current initial model parameter for adjustment, and the process returns to the step of applying the current initial model parameter to the target model until all initial model parameters of the target model are adjusted to obtain the final model parameters.
[0244] Specifically, on the basis of adjusting the current initial model parameters, the training test results of the target model are obtained. If the training test results meet the third preset conditions, it means that the current initial model parameters are adjusted appropriately, and the next initial model parameters need to be adjusted and continued to be adjusted. In other words, the currently adjusted initial model parameters are adjusted after the previous initial model parameters are adjusted.
[0245] This embodiment provides a whole process of sequentially adjusting the initial model parameters of the target model to obtain the final model parameters. The currently adjusted initial model parameters are adjusted after the previous initial model parameters are adjusted, thereby improving the accuracy of model parameter adjustment and training efficiency.
[0246] Figure 6 A flow chart of adjusting model parameters of a target model provided by an embodiment of the present invention is as follows: Figure 6 As shown, including:
[0247] S42: learning rate adjustment evaluation;
[0248] S43: Execute the first training test;
[0249] S44: optimizer adjustment evaluation;
[0250] S45: Execute a second training test;
[0251] S46: Mixed Precision Adjustment Evaluation;
[0252] S47: Execute the third training test;
[0253] S48: Regularization adjustment evaluation;
[0254] S49: Execute the fourth training test;
[0255] S50: perform complete training;
[0256] S51: comprehensive evaluation of iteration time and loss reduction;
[0257] S52: Output test report and evaluation results.
[0258] It should be noted that the model parameters in this embodiment may also include other types of model parameters, which are not limited and can be set according to actual conditions.
[0259] Finally, the automated test will output detailed results of each test step and the optimal parameter configuration. Users can determine the final training parameter configuration based on this information, thereby significantly reducing time and learning costs and obtaining the best model performance with the highest training efficiency.
[0260] The process of determining the model parameters of the target model provided in this embodiment reduces manual parameter adjustment based on human experience and the readjustment of all model parameters after the iterative training of the entire target model. This embodiment achieves the adjustment of the model parameters after the stability test of the target model by selecting the target number of samples while ensuring that there is no memory overflow. At the same time, the model parameter adjustment process, based on the adjustment of the previous model parameter, is completed during the model operation process before adjusting the next model parameter, which improves the efficiency of model parameter adjustment, reduces the training time cost and learning cost, improves the competitiveness of the product, and improves the overall economic benefits.
[0261] The above detailed descriptions of various embodiments corresponding to the service processing method are based on which the present invention further discloses a service processing device corresponding to the above method. Figure 7 FIG. 1 is a structural diagram of a service processing device provided by an embodiment of the present invention. Figure 7 As shown, the service processing device includes:
[0262] An acquisition module 11 is used to acquire a target model corresponding to the current business demand, weight parameters of each set communication strategy, and a preset topology structure composed of each computing accelerator;
[0263] A determination module 12, configured to process the performance parameters of each computing accelerator corresponding to each preset topology structure, weight parameter and each collective communication strategy to determine a target topology structure corresponding to the benchmark test;
[0264] The processing module 13 is used to train the target model based on the target topology structure, the parallel training strategy corresponding to each computing accelerator, and the target sample quantity of the target model to obtain the model parameters of the target model during the operation of the target model to complete the processing of the current business needs.
[0265] In some embodiments, before determining module 12, the method further includes:
[0266] A first determination submodule is used to determine the interconnection status of each computing accelerator in each preset topology structure;
[0267] The first configuration submodule is used to configure the link configuration information corresponding to the preset topology structure when the interconnection state is normal;
[0268] A first test submodule, used to test the bandwidth between each computing accelerator and between each computing accelerator and the processor in each preset topology structure to obtain a first test result;
[0269] A first entry submodule, for entering a step of determining a target topology structure according to each preset topology structure, weight parameter and performance parameter of each computing accelerator corresponding to each collective communication strategy if the first test result satisfies a first preset condition;
[0270] The first return submodule is used to return to the step of configuring the link configuration information corresponding to the preset topology structure if the first test result does not meet the first preset condition, and adjust the link configuration information until the first test result meets the first preset condition.
[0271] Since the embodiments of the device part correspond to the above embodiments, please refer to the description of the embodiments of the method part for the embodiments of the device part, and will not be repeated here.
[0272] For an introduction to a business processing device provided by the present invention, please refer to the above method embodiment, and the present invention will not be repeated here. It has the same beneficial effects as the above business processing method.
[0273] Figure 8 A structural diagram of a test system provided by an embodiment of the present invention, such as Figure 8 As shown, the test system includes:
[0274] A memory 21, used for storing computer programs;
[0275] The processor 22 is used to implement the steps of the business processing method when executing the computer program.
[0276] The test system provided in this embodiment may include but is not limited to a smart phone, a tablet computer, a laptop computer, or a desktop computer.
[0277] Among them, the processor 22 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 22 can be implemented in at least one hardware form of a digital signal processor (DSP), an FPGA, and a programmable logic array (PLA). The processor 22 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU; the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 22 may be integrated with a GPU, which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 22 may also include an artificial intelligence (AI) processor, which is used to process computing operations related to machine learning.
[0278] The memory 21 may include one or more computer-readable storage media, which may be non-transitory. The memory 21 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In this embodiment, the memory 21 is at least used to store the following computer program 211, wherein, after the computer program is loaded and executed by the processor 22, it can implement the relevant steps of the business processing method disclosed in any of the aforementioned embodiments. In addition, the resources stored in the memory 21 may also include an operating system 212 and data 213, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 212 may include Windows, Unix, Linux, etc. The data 213 may include, but is not limited to, data involved in the business processing method, etc.
[0279] In some embodiments, the test system may further include a display screen 23 , an input / output interface 24 , a communication interface 25 , a power supply 26 , and a communication bus 27 .
[0280] Those skilled in the art can understand that Figure 8 The structure shown in the figure does not constitute a limitation of the test system, and may include more or fewer components than shown.
[0281] The processor 22 implements the service processing method provided by any of the above embodiments by calling the instructions stored in the memory 21.
[0282] For an introduction to a test system provided by the present invention, please refer to the above method embodiment, and the present invention will not be repeated here. It has the same beneficial effects as the above business processing method.
[0283] Furthermore, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by the processor 22, the steps of the above-mentioned business processing method are implemented.
[0284] It is understandable that if the method in the above embodiment is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium to execute all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program codes.
[0285] For an introduction to a computer-readable storage medium provided by the present invention, please refer to the above method embodiment, and the present invention will not be repeated here. It has the same beneficial effects as the above business processing method.
[0286] Furthermore, the present invention also provides a computer program product, including a computer program / instruction, which implements the steps of the business processing method when executed by a processor.
[0287] For an introduction to a computer program product provided by the present invention, please refer to the above method embodiment, which will not be described in detail herein. It has the same beneficial effects as the above business processing method.
[0288] The above is a detailed introduction to a business processing method, a test system, a medium and a product provided by the present invention. The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the various embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description. It should be pointed out that for ordinary technicians in this technical field, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the scope of protection of the present invention.
[0289] It should also be noted that, in this specification, relational terms such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises", "comprising" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprising a ..." does not exclude the presence of other identical elements in the process, method, article or device including the element.
Claims
1. A business processing method, characterized in that: include: Obtain the target model corresponding to the current business needs, the weight parameters of each set communication strategy, and the preset topology structure composed of each computing accelerator; Processing the performance parameters of each computing accelerator corresponding to each preset topology structure, the weight parameter and each collective communication strategy to determine a target topology structure corresponding to the benchmark test; The benchmark test is a test in which a target topology is determined by different topologies without combining the actual application scenario of the target model; specifically, it includes: processing each collective communication strategy in each preset topology to obtain the corresponding performance parameter; determining the evaluation score of each preset topology according to the performance parameter and weight parameter of each collective communication strategy; determining the target topology according to the evaluation score of each preset topology; Based on the target topology structure, the target model is trained according to the parallel training strategy corresponding to each computing accelerator and the target number of samples of the target model to obtain the model parameters of the target model during the operation of the target model to complete the processing of the current business needs; wherein the parallel training strategy is to divide a target model based on a preset ratio to multiple computing accelerators corresponding to the corresponding target topology structure for parallel training.
2. The service processing method according to claim 1, characterized in that: The target model is trained based on the target topology structure according to the parallel training strategy corresponding to each computing accelerator and the target sample quantity of the target model to obtain the model parameters of the target model during the operation of the target model, including: Obtaining an initial parallel training strategy corresponding to each of the computing accelerators under the target topology structure; Determining, within each initial parallel training strategy, a parallel training strategy corresponding to the target model of the current business requirement; Based on the computing accelerator of the parallel training strategy, the training data corresponding to the target number of samples is input into the target model for training processing to determine the model parameters of the target model during the operation of the target model.
3. The service processing method according to claim 1, characterized in that: The performance parameters include at least bandwidth and delay, and determining the evaluation score of each of the preset topological structures according to the performance parameters and the weight parameters of each of the collective communication strategies includes: Determine the throughput corresponding to each of the collective communication strategies in the current preset topology structure according to the bandwidth corresponding to each of the collective communication strategies and the delay corresponding to each of the collective communication strategies; The weight parameters and the throughput corresponding to each of the collective communication strategies are used to determine an evaluation score of the current preset topology structure.
4. The service processing method according to claim 1 or 3, characterized in that: Before determining the target topology structure according to each preset topology structure, the weight parameter and the performance parameter of each computing accelerator corresponding to each collective communication strategy, the method further includes: Determining the interconnection status of each of the computing accelerators in each of the preset topological structures; When the interconnection state is normal, configuring link configuration information corresponding to the preset topology structure; Testing the bandwidth between the computing accelerators and between the computing accelerators and the processor in each of the preset topological structures to obtain a first test result; If the first test result satisfies the first preset condition, then proceeding to the step of determining the target topology structure according to each preset topology structure, the weight parameter and the performance parameter of each computing accelerator corresponding to each collective communication strategy; If the first test result does not meet the first preset condition, return to the step of configuring the link configuration information corresponding to the preset topology structure to which the link configuration information belongs, and adjust the link configuration information until the first test result meets the first preset condition.
5. The service processing method according to claim 4, characterized in that: The first preset condition is a first sub-preset condition or a second sub-preset condition, the first sub-preset condition is a preset condition for bandwidth and delay balance of a topological structure within a connection domain; the second sub-preset condition is a preset condition for bandwidth and delay balance of a topological structure outside a connection domain, and a process for determining the first preset condition includes: When each of the computing accelerators is in a connection domain topology structure, obtaining a first bandwidth standard deviation and a first delay standard deviation between the computing accelerators in the connection domain topology structure; Obtaining a grouping method corresponding to computing accelerators in a topological structure within a connection domain; If the number of groups corresponding to the grouping mode is 0, determine whether the first bandwidth standard deviation is less than or equal to the first bandwidth threshold and whether the first delay standard deviation is less than or equal to the first delay threshold; if the first bandwidth standard deviation is less than or equal to the first bandwidth threshold and the first delay standard deviation is less than or equal to the first delay threshold, determine that the first sub-preset condition is met; if the first bandwidth standard deviation is less than or equal to the first bandwidth threshold or the first delay standard deviation is less than or equal to the first delay threshold, determine that the first sub-preset condition is not met; If the number of groups corresponding to the grouping mode is greater than 0, it is determined whether the first bandwidth standard deviation is less than or equal to the second bandwidth threshold and whether the first delay standard deviation is less than or equal to the second delay threshold; if the first bandwidth standard deviation is less than or equal to the second bandwidth threshold and the first delay standard deviation is less than or equal to the second delay threshold, it is determined that the first sub-preset condition is met; if the first bandwidth standard deviation is less than or equal to the second bandwidth threshold or the first delay standard deviation is less than or equal to the second delay threshold, it is determined that the first sub-preset condition is not met; When each of the computing accelerators is in a topological structure outside the connection domain, obtaining a second bandwidth standard deviation and a second delay standard deviation between computing accelerators in the topological structure outside the connection domain and across the connection domain; Determine whether the second bandwidth standard deviation is less than or equal to a third bandwidth threshold and whether the second delay standard deviation is less than or equal to a third delay threshold; if the second bandwidth standard deviation is less than or equal to the third bandwidth threshold and the second delay standard deviation is less than or equal to the third delay threshold, determine that the second sub-preset condition is met; if the second bandwidth standard deviation is less than or equal to the third bandwidth threshold or the second delay standard deviation is less than or equal to the third delay threshold, determine that the second sub-preset condition is not met.
6. The service processing method according to claim 2, characterized in that: The determining of the parallel training strategy corresponding to the target model of the current business requirement within each initial parallel training strategy includes: Obtaining a preset segmentation ratio of the target model on each of the computing accelerators; Preprocessing the target model according to different preset segmentation ratios to determine performance evaluation parameters corresponding to each of the initial parallel training strategies; The parallel training strategy is determined according to the performance evaluation parameters.
7. The service processing method according to claim 6, characterized in that: The performance evaluation parameters include at least training iteration time, model performance and resource utilization; The preprocessing of the target model according to different preset segmentation ratios to determine corresponding performance evaluation parameters includes: Preprocessing the target model according to different preset segmentation ratios to obtain the training iteration time, model performance and resource utilization rate corresponding to the parallel computing accelerator and the single-row computing accelerator in each of the initial parallel training strategies; Determine the model parallel efficiency corresponding to each of the initial parallel training strategies based on the training iteration time and model performance corresponding to the parallel computing accelerator and the single-row computing accelerator in each of the initial parallel training strategies; The performance evaluation parameter is determined according to the model parallel efficiency and the resource utilization rate corresponding to each of the initial parallel training strategies.
8. The service processing method according to claim 7, characterized in that: The training iteration time is the time for forward propagation, backward propagation and parameter update of the target model; The process of determining the training iteration time includes: Get the time of each iteration and the total number of iterations; Determine the training iteration time according to the time of each iteration and the total number of iterations; Correspondingly, the model performance is the accuracy and loss index performance of the target model in the validation set or the training set; the process of determining the model performance includes: Obtaining model parameters and a validation set of the target model; The model parameters and the validation set are evaluated according to a performance evaluation function to obtain the model performance.
9. The service processing method according to claim 7, characterized in that: The process of determining the parallel efficiency of the model includes: Obtain a first training iteration time corresponding to the parallel computing accelerator and a second training iteration time corresponding to the single-row computing accelerator; Obtaining a first model performance corresponding to a parallel computing accelerator and a second model performance corresponding to a single-row computing accelerator; Determining a parallel iteration time according to a first training iteration time and the first model performance corresponding to the parallel computing accelerator; Determine a single-row iteration time according to a second training iteration time and the second model performance corresponding to the single-row computing accelerator; The model parallel efficiency is determined according to the parallel iteration time and the single-row iteration time.
10. The service processing method according to claim 9, characterized in that: The computing accelerator based on the parallel training strategy inputs the training data corresponding to the target sample quantity into the target model for training processing to determine the model parameters of the target model during the operation of the target model, including: Get the current preset sample quantity; Input the training data corresponding to the current preset number of samples into the target model; During the operation of the target model, determining whether the operating memory exceeds the preset memory; If the preset memory is exceeded, the next preset number of samples is obtained, and the process returns to the step of inputting the training data corresponding to the current preset number of samples into the target model, until the running memory corresponding to the training data corresponding to the current preset number of samples during the running of the target model does not exceed the preset memory; wherein the next preset number of samples is an integer multiple of the current preset number of samples; If it does not exceed the preset memory, the current preset number of samples is used as the target number of samples, and the loss function of each segment is obtained by training according to the target number of samples; the convergence speed and loss fluctuation of the target model are determined according to the loss value of the loss function of each segment; if the convergence speed and loss fluctuation meet the second preset condition, the initial model parameters of the target model are adjusted in sequence to obtain the final model parameters.
11. The service processing method according to claim 10, characterized in that: There are multiple types of initial model parameters; The step of sequentially adjusting the initial model parameters of the target model to obtain the final model parameters includes: After the current initial model parameters are adjusted, the current initial model parameters are applied to the target model, and applied to the target model to obtain the training test results of the target model; If the training test result meets the third preset condition, the next initial model parameter is obtained as the current initial model parameter for adjustment, and the step of applying the current initial model parameter to the target model is returned to until the various initial model parameters of the target model are adjusted to obtain the final model parameters.
12. A testing system, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the business processing method according to any one of claims 1 to 11 when executing the computer program.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the business processing method according to any one of claims 1 to 11 are implemented.
14. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps of the business processing method described in any one of claims 1 to 11 are implemented.
Citation Information
Patent Citations
Intelligent computing-oriented pipeline parallel training self-adaptive adjustment system and intelligent computing-oriented pipeline parallel training self-adaptive adjustment method
CN115237580A
Business operation method and device, computer equipment and storage medium
CN117667386A