Policy optimization method, device, equipment, storage medium and computer program product
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI ORIENTAL COMPUTER TECHNOLOGY CO LTD
- Filing Date
- 2026-03-23
- Publication Date
- 2026-08-07
AI Technical Summary
实际性能测试过程需要搭建由多个训练节点构成的大规模硬件集群,并且涉及软硬件协同等多方协作,容易引发训练节点间的通信延迟与报错故障,导致性能测试排查周期拉长,进而导致策略的优化效率不高
首先,获取多个候选训练策略以及已知硬件设备基于每个候选训练策略进行模型训练时的实测性能数据,能够为后续的性能预测提供准确的参考数据。针对每个候选训练策略,基于已知硬件设备基于候选训练策略进行模型训练时的实测性能数据、已知硬件设备的硬件参数以及目标硬件设备的硬件参数,对目标硬件设备基于候选训练策略进行模型训练时的性能数据进行预测,得到预测性能数据,通过利用已知硬件设备的实测性能数据推导目标硬件设备的性能表现,无需在目标硬件设备上实际搭建大规模集群进行测试。接着,基于预测性能数据,对候选训练策略进行优先级评估,得到候选训练策略的优先级,能够量化不同候选训练策略在目标环境下的预期表现的情况。最后,基于多个候选训练策略的优先级,从多个候选训练策略中,筛选出用于应用至目标硬件设备的目标训练策略,能够显著缩短策略性能优化的排查周期,在无需实际部署的前提下快速确定最优的目标训练策略,从而提高策略的优化效率。
Smart Images

Figure CN121900978B_ABST
Abstract
Description
Technical Field
[0001] This application relates to data analysis technology, and more particularly to a strategy optimization method, apparatus, device, storage medium, and computer program product. Background Technology
[0002] In artificial intelligence model training scenarios, in order to improve model training efficiency, it is necessary to pre-design training strategies deployed on hardware and compare the performance of different training strategies to optimize the training strategy.
[0003] To compare the performance of different training strategies, a practical performance test method was adopted, namely, performance testing by directly deploying a large-scale physical hardware cluster. The practical performance test process requires building a large-scale hardware cluster consisting of multiple training nodes and involves multi-party collaboration, including software and hardware coordination. This can easily lead to communication delays and errors between training nodes, resulting in a longer performance test troubleshooting cycle and consequently, low strategy optimization efficiency. Summary of the Invention
[0004] This application provides a strategy optimization method, apparatus, computer-readable storage medium, and computer program product, which can improve the optimization efficiency of strategies.
[0005] The technical solution of this application embodiment is implemented as follows: This application provides a strategy optimization method, the method comprising: Multiple candidate training strategies are obtained, and measured performance data of known hardware devices when training a model based on each of the candidate training strategies are obtained. For each candidate training strategy, based on the measured performance data of the known hardware device when training the model based on the candidate training strategy, the hardware parameters of the known hardware device, and the hardware parameters of the target hardware device, the performance data of the target hardware device when training the model based on the candidate training strategy is predicted to obtain the predicted performance data. Based on the prediction performance data, the candidate training strategies are prioritized to obtain their priorities. Based on the priority of the multiple candidate training strategies, a target training strategy for application to the target hardware device is selected from the multiple candidate training strategies.
[0006] This application provides a strategy optimization device, including: The acquisition module is used to acquire multiple candidate training strategies and acquire measured performance data of known hardware devices when training a model based on each of the candidate training strategies. The prediction module is used to predict the performance data of the target hardware device when training the model based on the candidate training strategy, based on the measured performance data of the known hardware device when training the model based on the candidate training strategy, the hardware parameters of the known hardware device, and the hardware parameters of the target hardware device, for each candidate training strategy, so as to obtain the predicted performance data. An evaluation module is used to evaluate the priority of the candidate training strategies based on the prediction performance data, and obtain the priority of the candidate training strategies. The filtering module is used to filter out the target training strategy for application to the target hardware device from the multiple candidate training strategies based on their priority.
[0007] This application provides an electronic device, the electronic device comprising: Memory is used to store executable instructions or computer programs. When a processor executes computer-executable instructions or computer programs stored in the memory, it implements the strategy optimization method provided in the embodiments of this application.
[0008] This application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing the strategy optimization method provided in this application when executed by a processor.
[0009] This application provides a computer program product, including a computer program or computer executable instructions, which, when executed by a processor, implements the strategy optimization method provided in this application.
[0010] The embodiments of this application have the following beneficial effects: First, acquiring multiple candidate training strategies and measured performance data of known hardware devices training models based on each candidate strategy provides accurate reference data for subsequent performance prediction. For each candidate training strategy, based on the measured performance data of known hardware devices training models based on the candidate strategy, the hardware parameters of the known hardware devices, and the hardware parameters of the target hardware device, the performance data of the target hardware device training models based on the candidate strategy is predicted, yielding predicted performance data. By utilizing the measured performance data of known hardware devices, the performance of the target hardware device is inferred, eliminating the need to actually build a large-scale cluster for testing on the target hardware device. Next, based on the predicted performance data, the candidate training strategies are prioritized, quantifying the expected performance of different candidate training strategies in the target environment. Finally, based on the priorities of multiple candidate training strategies, the target training strategy for application to the target hardware device is selected from among them. This significantly shortens the strategy performance optimization process, quickly determining the optimal target training strategy without actual deployment, thereby improving the optimization efficiency. Attached Figure Description
[0011] Figure 1 This is a schematic diagram of the strategy optimization system architecture provided in the embodiments of this application; Figure 2 This is a schematic diagram of the structure of the device provided in the embodiments of this application; Figure 3 This is a first flowchart illustrating the strategy optimization method provided in this application embodiment; Figure 4 This is a schematic diagram of the second process of the strategy optimization method provided in the embodiments of this application; Figure 5 This is a schematic diagram of the third process of the strategy optimization method provided in the embodiments of this application; Figure 6 This is a schematic diagram of the fourth process of the strategy optimization method provided in the embodiments of this application; Figure 7 This is a schematic diagram of the fifth process of the strategy optimization method provided in the embodiments of this application; Figure 8 This is a flowchart illustrating the application scenario of the strategy optimization method provided in the embodiments of this application; Figure 9 This is a schematic diagram of the technical framework for the strategy optimization method provided in the embodiments of this application under a specific application scenario.
[0012] It should be noted that the terms "first" and "second" mentioned above are only used to distinguish between different options and do not represent the degree of superiority or inferiority of the options or their priority in the implementation process. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0014] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0015] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0016] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0017] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.
[0018] When the technical solutions of this application are applied to real-world scenarios, all data collection and processing processes are subject to compliance with relevant laws and regulations. After obtaining valid authorization from the data subject (such as informed consent or separate consent), the data processing party's data use and processing activities will be limited to the scope defined by law and the valid authorization of the data subject.
[0019] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.
[0020] 1) Tensor Parallelism (TP) is a strategy that splits a single large tensor within a model, such as a weight matrix, across multiple hardware acceleration devices for parallel computation. Each device is used to compute only a portion of the tensor, and the results are aggregated through communication. In the embodiments of this application, it is a core parallel strategy, for example, splitting and computing the model's tensor across eight accelerator cards within a single machine to reduce the memory pressure on a single card.
[0021] 2) Pipeline Parallelism (PP) is a parallel technique that sequentially distributes different layers or blocks of a model to different hardware acceleration devices to form a computational pipeline. Data is processed on one device and then passed to the next. In the embodiments of this application, it is often used for parallelism between multiple machines. Its efficiency is closely related to the cavitation rate and is one of the key strategies that need to be analyzed in performance prediction models.
[0022] 3) Data Parallelism (DP) is a parallel method that divides training data into multiple parts, deploys a complete copy of the model on each hardware acceleration device, processes a portion of the data on each device, and finally synchronizes model parameter updates through communication. In the embodiments of this application, it is usually used in combination with tensor parallelism, pipelined parallelism, etc., to effectively expand the overall training scale.
[0023] 4) Micro-batch size refers to the further subdivision of a global batch of data into smaller data units when using pipelined parallelism. These smaller units are then sequentially fed into the pipeline for processing. In the embodiments of this application, the micro-batch size is a key hyperparameter affecting pipeline efficiency, memory usage, and communication overhead, and is an important parameter that needs to be determined when optimizing parallel strategies.
[0024] 5) Bubble ratio refers to the proportion of idle time in the total training time caused by data dependency waiting during the pipeline startup and emptying phases in pipeline parallelism. In this embodiment, the bubbling ratio is a key negative indicator for measuring pipeline parallelism efficiency. This method incorporates it into the final performance calculation through a formula to achieve more accurate time estimation.
[0025] 6) A performance event refers to a core operational unit that accounts for the majority of the time consumption in the entire model training process, such as computational or communication operations. In the embodiments of this application, model training is broken down into various performance events such as general matrix multiplication, attention mechanism, vector computation, and inter-card communication, so as to facilitate the evaluation of its running efficiency on a single device and its mapping to a new device.
[0026] 7) Hardware utilization rate refers to the ratio of the actual operating performance (such as measured computing power or bandwidth) of hardware to its theoretical peak performance when performing a specific performance event, reflecting the degree of matching between software algorithms and hardware architecture.
[0027] In related technologies, to determine the optimal parallel strategy for training large models on large-scale clusters (e.g., 512 cards), a real-world deployment test method is adopted. This involves directly building a complete hardware cluster, configuring the software environment, writing different parallel strategy code (e.g., strategy A, strategy B), and actually running the training task to test performance metrics (e.g., the number of tokens processed per second).
[0028] The limitations of the related technologies include: 1) When upgrading hardware or migrating training models between different types of hardware devices, it is impossible to use the performance data of the old hardware to predict the performance of the new hardware. This means that each hardware cluster change requires costly field testing from scratch. Therefore, directly deploying and testing the performance of different parallel strategies on a new large-scale hardware cluster would result in huge resource consumption and a long evaluation cycle.
[0029] 2) The inability to quantitatively predict the performance of a strategy before deployment leads to a lack of specific data analysis support for strategy selection, forcing the optimization process to rely solely on repeated trial and error after hardware testing. This application provides a strategy optimization method, apparatus, device, computer-readable storage medium, and computer program product, which can improve the efficiency of strategy optimization. The following describes exemplary applications of the electronic devices provided in this application. These devices can be implemented as various types of terminals such as laptops, tablets, desktop computers, set-top boxes, smartphones, smart speakers, smartwatches, smart TVs, and in-vehicle terminals, or as servers. Exemplary applications when the device is implemented as a terminal or server will be described below.
[0030] See Figure 1 , Figure 1This is a schematic diagram of the architecture of the policy optimization system provided in this application embodiment. To support a policy optimization application, the policy optimization system 100 includes at least a database 500, a terminal 600, a network 300, and a server 200. The terminal 600 is connected to the server 200 through the network 300, which can be a wide area network (WAN), a local area network (LAN), or a combination of both.
[0031] In some embodiments, server 200 may be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server providing basic cloud computing services. Terminals and servers can be connected directly or indirectly via wired or wireless communication, and this embodiment does not impose any limitations.
[0032] In some embodiments, the embodiments of this application can be implemented by the terminal 600 alone. For example, the terminal 600 obtains multiple candidate training strategies from the database 500 through the network 300. The terminal 600 collects measured performance data of known hardware devices training models based on each candidate training strategy on the client side. After receiving the measured performance data of known hardware devices training models based on each candidate training strategy, the terminal 600 predicts the performance data of the target hardware device training models based on the candidate training strategy for each candidate training strategy, based on the measured performance data of known hardware devices training models based on the candidate training strategy, the hardware parameters of known hardware devices, and the hardware parameters of the target hardware device, to obtain predicted performance data. Based on the predicted performance data, the terminal 600 performs priority evaluation on the candidate training strategies to obtain the priority of the candidate training strategies. Based on the priority of multiple candidate training strategies, the terminal 600 selects the target training strategy to be applied to the target hardware device from the multiple candidate training strategies and displays it on the client side.
[0033] In some embodiments, the embodiments of this application can be implemented collaboratively by a server and a terminal. For example, terminal 600 collects measured performance data of known hardware devices training models based on each candidate training strategy on a client side. After receiving the measured performance data, terminal 600 sends the measured performance data to server 200 via network 300. After receiving the audio to be processed, server 200 retrieves multiple candidate training strategies from database 500. For each candidate training strategy, based on the measured performance data of known hardware devices training models based on the candidate training strategies, the hardware parameters of the known hardware devices, and the hardware parameters of the target hardware device, server 200 predicts the performance data of the target hardware device training models based on the candidate training strategies, obtaining predicted performance data. Based on the predicted performance data, server 200 prioritizes the candidate training strategies, obtaining the priority of the candidate training strategies. Based on the priorities of multiple candidate training strategies, server 200 selects the target training strategy to be applied to the target hardware device from the multiple candidate training strategies. Server 200 transmits the target training strategy to terminal 600 via network 300 for display on the client side of terminal 600.
[0034] The strategy optimization method provided in this application can be applied to any scenario where hardware resource configuration optimization is performed for large-scale model training. Specific application scenarios may include: 1) In a resource scheduling scenario for a large model training platform, users upload performance logs (measured performance data) of a benchmark model running on a small-scale test cluster (known hardware devices) via their terminals. After the logs are transmitted to the server, the server begins analysis. The server retrieves multiple parallel training configuration schemes (multiple candidate training strategies) from the database, including different combinations of data parallelism, model parallelism, and pipeline parallelism. Subsequently, the server combines the hardware specifications of the small-scale test cluster (known hardware parameters) and the hardware specifications of the large-scale production cluster to be deployed (target hardware parameters) to calculate the expected training speed and memory usage (predicted performance data) for each configuration scheme on the production cluster using a performance prediction model. Then, the server scores and ranks all configuration schemes based on the expected training speed and whether memory usage overflows (priority evaluation). Finally, the server selects the configuration scheme with the highest score as the optimal deployment scheme (target training strategy) and sends the scheme parameters back to the terminal to guide the user in configuring the production environment.
[0035] 2) In the scenario of heterogeneous computing power recommendation by cloud service providers, users submit performance records (measured performance data) of training AI models on their own old-model GPU servers (known hardware devices) via their terminals. After the records are transmitted to the server, the server identifies the basic characteristics of the user's model. For the new high-performance computing instances (target hardware devices) provided on the cloud platform, the server obtains multiple distributed training strategies (multiple candidate training strategies) suitable for the model. Subsequently, based on the performance records of the old server, the hardware indicators of the old server (hardware parameters of known hardware devices), and the hardware indicators of the new computing instance (hardware parameters of the target hardware device), the server simulates and extrapolates the training time of the model when applying each strategy on the new computing instance (predicted performance data). After that, the server compares the estimated time of each strategy and determines the strategy that can best utilize the performance of the new hardware (priority evaluation). Finally, the server recommends the training strategy with the best predicted performance (target training strategy) and the corresponding expected speedup ratio to the user terminal to assist the user in making computing power upgrade decisions.
[0036] See Figure 2 , Figure 2 This is a schematic diagram of the structure of the electronic device 400 provided in the embodiments of this application. Figure 2 The illustrated electronic device 400 includes at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. The various components in the electronic device 400 are coupled together via a bus system 440. It is understood that the bus system 440 is used to implement communication between these components. In addition to a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 2 The general labeled all buses as Bus System 440.
[0037] Processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0038] User interface 430 includes one or more output devices 431 that enable the presentation of media content, including one or more speakers and / or one or more visual displays. User interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0039] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 450 may optionally include one or more storage devices physically located away from the processor 410.
[0040] The memory 450 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 450 described in this application embodiment is intended to include any suitable type of memory.
[0041] In some embodiments, memory 450 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.
[0042] Operating system 451 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks; The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 including: Bluetooth, WiFi, and Universal Serial Bus (USB), etc. Presentation module 453 is configured to enable the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 431 (e.g., a display screen, a speaker, etc.) associated with user interface 430. The input processing module 454 is used to detect and translate one or more user inputs or interactions from one or more input devices 432.
[0043] In some embodiments, the apparatus provided in this application can be implemented in software. Figure 2 A policy optimization device 455 stored in memory 450 is shown. This device can be software in the form of programs and plug-ins, and includes the following software modules: an acquisition module 4551, a prediction module 4552, an evaluation module 4553, and a screening module 4554. These modules are logically linked and can therefore be arbitrarily combined or further separated according to their implemented functions. The functions of each module will be described below.
[0044] The strategy optimization method provided in the embodiments of this application will be described below. As mentioned above, the electronic device implementing the strategy optimization method of the embodiments of this application can be a terminal, a server, or a combination of both. Therefore, the executing entity of each step will not be described again below.
[0045] See Figure 3 , Figure 3 This is a flowchart illustrating the strategy optimization method provided in the embodiments of this application, which will be combined with... Figure 3 The steps shown are explained.
[0046] In step 101, multiple candidate training strategies are obtained, and measured performance data of known hardware devices when training the model based on each candidate training strategy is obtained.
[0047] Candidate training strategies refer to multiple alternative schemes pre-defined before training a deep learning model to guide the model's computational allocation and data interaction on hardware devices. Candidate training strategies specify the method of splitting model parameters, the method of allocating data, and the communication logic between computing nodes. Candidate training strategies can be a single parallel mode or a combination of multiple parallel modes. For example, candidate training strategies can be data parallelism (DP), tensor parallelism (TP), pipeline parallelism (PP), expert parallelism (EP), sequence parallelism (SP), zero redundancy optimizer (ZeRO) strategies, or any two or more of the above parallel modes combined. Known hardware devices refer to hardware computing units or small computing clusters used for basic performance testing to obtain measured data as a prediction benchmark. Known hardware devices are hardware environments that are easily accessible and have low deployment costs or are small in scale. For example, a known hardware device can be a single Graphics Processing Unit (GPU) card, a single Tensor Processing Unit (TPU) card, a single Neural Processing Unit (NPU) card, a small server node containing a few computing cards (e.g., a single-machine 8-card server), a cluster with the same architecture as the target hardware device but on a smaller scale, or a heterogeneous computing device with a different architecture than the target hardware device but with a mappable relationship. Measured performance data refers to objective data reflecting the hardware's operating status and processing capabilities collected through analysis tools when actually running model training tasks or micro-benchmarking tasks on a known hardware device. Measured performance data is used to quantify the performance of a known hardware device when performing a specific computational or communication task. For example, measured performance data can be the actual execution time of a specific operator (such as matrix multiplication GEMM), peak memory usage, actual transmission rate of the communication interface, utilization of computing units, cache hit rate, power consumption data, and measured values of floating-point operations per second (FLOPS).
[0048] In some embodiments, step 101, "obtaining measured performance data of a known hardware device when training a model based on each candidate training strategy", can be achieved as follows: based on the candidate training strategy, a corresponding simulated training environment is constructed on the known hardware device; a preset model layer is run in the simulated training environment, and the execution data of the underlying operators of the known hardware device when running the preset model layer is collected through a performance analysis interface to obtain measured performance data.
[0049] Here, the simulated training environment refers to the software runtime context configured according to the parallel parameters in the candidate training strategy; the performance analysis interface refers to the programming interface provided by the hardware driver layer or deep learning framework for deriving hardware performance metrics; and the underlying operator execution data refers to the execution details of matrix operations or set communication functions called during model training. The underlying operator execution data can be the time consumption of a general matrix multiplication operator or the bandwidth usage of a fully reduced communication operator.
[0050] For example, configure a runtime environment (simulated training environment) with tensor parallelism of 8 on a single 8-card server (known hardware device), run one layer of model (preset model layer), and export the average time of all matrix multiplication operators of the model layer during the forward propagation process through the performance analysis tool interface (performance analysis interface) (actual performance data).
[0051] In this embodiment, by constructing a simulated training environment on a known hardware device and running a preset model layer, real underlying operator execution data can be obtained at a low resource cost. Obtaining measured performance data based on a performance analysis interface ensures data accuracy and provides a reliable data foundation for subsequent performance prediction across hardware devices.
[0052] In step 102, for each candidate training strategy, based on the measured performance data of the known hardware device when training the model according to the candidate training strategy, the hardware parameters of the known hardware device, and the hardware parameters of the target hardware device, the performance data of the target hardware device when training the model according to the candidate training strategy is predicted to obtain the predicted performance data.
[0053] In this context, hardware parameters refer to static parameters describing the physical attributes, theoretical performance limits, or specifications of a hardware device. These parameters are provided by the hardware manufacturer and do not change with the software's running state. For example, hardware parameters may include theoretical peak computing power for half-precision floating-point numbers (BFLOAT16), theoretical peak computing power for single-precision floating-point numbers (FP32), memory capacity, memory bandwidth, inter-card interconnect bandwidth, inter-machine interconnect bandwidth (such as InfiniBand bandwidth), number of computing cores, clock frequency, and hardware topology information. The target hardware device refers to a large-scale computing environment or specific hardware device cluster that is planned for deploying model training tasks and whose performance needs to be predicted. In this embodiment, the target hardware device has a larger scale, higher cost, or more complex topology than known hardware devices. For example, the target hardware device may be a large-scale high-performance computing cluster containing hundreds or even thousands of computing cards (such as a cluster of 512 GPU cards), a geographically distributed data center computing resource pool, or a virtual hardware environment that has not yet been actually built but has been planned and configured. Predicted performance data refers to data on the expected performance of a representation model running on the target hardware device, derived from measured data and hardware parameters of known hardware devices. Predicted performance data is not obtained through actual testing on the target hardware device, but rather generated through computational logic or model deduction. For example, predicted performance data could include predicted total training time, predicted single-step iteration time, predicted tokens per second (TPS), predicted tokens per card per second (TGS), predicted communication latency, predicted computation operator time, and predicted GPU memory usage. Predicting performance data refers to the action of substituting input variables (measured performance data, hardware parameters) into a pre-defined mathematical or machine learning model for calculation, thereby outputting a prediction result. For example, performance data prediction could include predictions of computation time, communication time, and total throughput.
[0054] In some embodiments, refer to Figure 4 The above step 102 can be achieved through Figure 4 Steps 1021A to 1022A shown are implemented.
[0055] In step 1021A, the effective performance parameters of the known hardware device are determined based on the hardware parameters and utilization rate of the known hardware device, and the effective performance parameters of the target hardware device are determined based on the hardware parameters and utilization rate of the target hardware device.
[0056] Utilization rate refers to the ratio between the actual performance output of a hardware device when performing a specific task and its theoretical peak performance, characterizing the effective utilization of hardware resources. For example, utilization rate can be the computing power utilization rate of a computing core (actual FLOPS / theoretical FLOPS), the memory bandwidth utilization rate (actual transfer rate / theoretical bandwidth), or the bandwidth utilization rate of a communication link. Effective performance parameters refer to parameters that reflect the true processing capability of a hardware device under actual operating conditions, adjusted for utilization rate. Effective performance parameters equal the theoretical hardware parameters of the hardware device multiplied by the corresponding utilization rate. For example, effective performance parameters can be effective computing power (Effective TFLOPS), effective video memory bandwidth, and effective communication bandwidth.
[0057] Here, based on the known hardware parameters and utilization rate of the known hardware device, the effective performance parameters of the known hardware device are determined, and based on the hardware parameters and utilization rate of the target hardware device, the effective performance parameters of the target hardware device are determined, which can be referred to in the following formulas (1) to (2).
[0058] (1) (2) in, Indicates the effective performance parameters of a known hardware device; Indicates the hardware parameters of a known hardware device; This indicates the utilization rate of known hardware devices. Indicates the effective performance parameters of the target hardware device; Indicates the hardware parameters of the target hardware device; This indicates the utilization rate of the target hardware device.
[0059] For example, if the known half-precision floating-point computing power of the hardware device is 148 TFLOPS and the actual measured computing power utilization rate is 0.6, then the calculated effective half-precision floating-point computing power of the known hardware device is 88.8 TFLOPS. If the target hardware device has a video memory bandwidth of 2.039 TB / s and an estimated video memory bandwidth utilization rate of 0.7, then the calculated effective video memory bandwidth of the target hardware device is 1.4273 TB / s.
[0060] In some embodiments, the utilization rate of the known hardware device or the target hardware device can be achieved by performing micro-benchmark tests on the known hardware device or the target hardware device, and determining the utilization rate based on the test time of the micro-benchmark test, the theoretical workload of the micro-benchmark test, and the hardware parameters.
[0061] Micro-benchmarking refers to a single-function test task targeting a specific computational or communication type, such as matrix multiplication testing or fully reduced communication testing. Theoretical workload refers to the number of mathematical operations or data transmissions required to complete a micro-benchmarking task.
[0062] For example, when running a matrix multiplication test task on a known hardware device, the utilization rate is calculated based on the theoretical computational load of the task, the actual measured time of the task, and the theoretical peak computing power of the known hardware device. Here, the computational utilization rate can be referred to the following formula (3).
[0063] (3) in, The number of floating-point operations representing the theoretical computational quantity. Indicates the actual time consumed. Given the theoretical peak computing power of the hardware device, To calculate utilization rate.
[0064] In step 1022A, based on the known effective performance parameters of the hardware device, the effective performance parameters of the target hardware device, and the measured performance data of the known hardware device when training the model based on the candidate training strategy, the performance data of the target hardware device when training the model based on the candidate training strategy is predicted to obtain the predicted performance data.
[0065] In some embodiments, the measured performance data includes measured time consumption data, and the predicted performance data includes predicted time consumption data; the above step 1022A can be implemented in the following way: based on the known effective performance parameters of the hardware device and the known measured time consumption data of the hardware device when training the model based on the candidate training strategy, determine the data load of the target hardware device when training the model based on the candidate training strategy; based on the data load and the effective performance parameters of the target hardware device, determine the predicted time consumption data.
[0066] Measured time consumption data refers to the actual time consumed to complete a specific computation or communication step on a known hardware device. Measured time consumption data is measured in milliseconds (ms), microseconds (µs), or seconds (s). For example, measured time consumption data could be the time to perform a matrix multiplication or the time to complete a gradient synchronization communication. Predicted time consumption data refers to the estimated expected time required for the model to complete the corresponding steps on the target hardware device. Data load refers to the physical quantification of the total workload that the hardware device needs to process in the model training task. For example, data load could be the total number of floating-point operations (Total FLOPs), the total number of data bytes to be transferred (Total Bytes), or the amount of memory data to be read and written.
[0067] Here, the predicted time consumption data is determined based on the data load and the effective performance parameters of the target hardware device, which can be done in the following way (4).
[0068] (4) in, This represents the predicted time consumption data; Indicates data load; This indicates the effective performance parameters of the target hardware device.
[0069] For example, data load The effective computing power of the target hardware device is 2000 TFLOPS. 200 TFLOPS, then the predicted time data is calculated. It lasts for 10 seconds.
[0070] In this embodiment, the measured performance data includes measured time consumption data, and the predicted performance data includes predicted time consumption data, explicitly defining the predicted performance data as a time-dimensional indicator. Based on the known effective performance parameters of the hardware device and the known measured time consumption data of the hardware device when training the model based on the candidate training strategy, the data load of the target hardware device when training the model based on the candidate training strategy is determined. The time consumption and hardware performance are converted into a unified computational load through physical principles. Based on the data load and the effective performance parameters of the target hardware device, the predicted time consumption data is determined, enabling accurate time prediction based on the data load and providing accurate predicted performance data for the evaluation of the training strategy.
[0071] In some embodiments, the above-mentioned "determining the data load of the target hardware device when training the model based on the candidate training strategy, based on the effective performance parameters of the known hardware device and the measured time consumption data of the known hardware device when training the model based on the candidate training strategy" can be achieved in the following way: determining the layer ratio between the known model layer number and the target model layer number, where the known model layer number represents the number of model layers when the model is trained on the known hardware device, and the target model layer number represents the number of model layers when the model is trained on the target hardware device; and determining the data load of the target hardware device when training the model based on the candidate training strategy, based on the layer ratio, the effective performance parameters of the known hardware device, and the measured time consumption data of the known hardware device when training the model based on the candidate training strategy.
[0072] The known model layer count refers to the number of neural network stacks in the model running on the known hardware. For ease of testing, the known model layer count is relatively low. For example, it could be 1, 2, or 4 layers. The target model layer count refers to the number of neural network stacks in the complete large model planned for training on the target hardware. The target model layer count is higher. For example, it could be 32, 64, 96, or even more layers. The layer ratio refers to the ratio between the target model layer count and the known model layer count.
[0073] Here, based on the layer ratio, the known effective performance parameters of the hardware device, and the measured time consumption data of the known hardware device when training the model based on the candidate training strategy, the data load of the target hardware device when training the model based on the candidate training strategy can be determined by referring to the following formula (5).
[0074] (5) in, Indicates data load; This represents the actual measured time consumption data; Indicates the ratio of the number of floors; This represents the effective performance parameters of a known hardware device.
[0075] For example, measured time data For 5 seconds, the ratio of layers The effective computing power of the known hardware device is 64. If the value is 50 TFLOPS, then the calculated data load is... It is 16000 TFLOPS.
[0076] In this embodiment, the layer ratio between the known model layer count and the target model layer count is determined. The known model layer count represents the number of model layers when trained on a known hardware device, while the target model layer count represents the number of model layers when trained on a target hardware device. This incorporates the direct impact of model structure changes on computational load. Based on the layer ratio, the effective performance parameters of the known hardware device, and the measured time consumption data of the known hardware device when training the model using candidate training strategies, the data load of the target hardware device when training the model using candidate training strategies is determined. This allows for the correction of load calculation results when the model structure is adjusted, thereby improving the accuracy of data load quantification.
[0077] In this embodiment, effective performance parameters of the known hardware devices are determined based on their hardware parameters and utilization rates. Similarly, effective performance parameters of the target hardware device are determined based on its hardware parameters and utilization rates. By combining these parameters with the utilization rate variable, the actual capabilities of the hardware in real-world operating scenarios can be more accurately reflected. Based on the effective performance parameters of the known hardware devices, the effective performance parameters of the target hardware device, and the measured performance data of the known hardware devices during model training using candidate training strategies, the performance data of the target hardware device during model training using candidate training strategies is predicted. This predictive performance data corrects prediction biases caused by different hardware device load states, improves the accuracy of cross-hardware device performance prediction, and thus enhances the optimization efficiency of training strategies.
[0078] In some embodiments, refer to Figure 5 The above step 102 can be achieved through Figure 5 Steps 1021B to 1024B shown are implemented.
[0079] In step 1021B, feature extraction is performed on the hardware parameters of the known hardware device to obtain the first feature, feature extraction is performed on the hardware parameters of the target hardware device to obtain the second feature, and feature extraction is performed on the measured performance data to obtain the third feature.
[0080] Feature extraction refers to the process of extracting, transforming, or combining numerical values or vectors that effectively represent the core information of raw data. This can involve directly constructing a vector from a set of hardware parameter values, or it can involve preprocessing such as normalization and standardization. The first feature refers to the data representation obtained after feature extraction of the hardware parameters of a known hardware device. For example, the first feature could be a vector containing values such as the computing power and bandwidth of a known hardware device. The second feature refers to the data representation obtained after feature extraction of the hardware parameters of a target hardware device. For example, the second feature could be a vector containing values such as the computing power and bandwidth of the target hardware device. The third feature refers to the data representation obtained after feature extraction of measured performance data. For example, the third feature could be a vector containing the time consumption or utilization rate of various events measured on a known hardware device.
[0081] For example, extract the three parameters of known hardware devices—peak computing power, memory bandwidth, and interconnect bandwidth—and form a three-dimensional vector as the first feature.
[0082] In step 1022B, based on the first feature and the second feature, a fourth feature is determined, which characterizes the difference between the hardware parameters of the known hardware device and the hardware parameters of the target hardware device.
[0083] The fourth feature refers to the data representation generated by comparing the first and second features, used to characterize the performance differences between the two hardware devices. The target feature refers to the comprehensive feature representation formed by integrating the third feature, which characterizes the workload of the task itself, with the fourth feature, which characterizes the hardware differences, and can be used for final prediction.
[0084] In some embodiments, the above-mentioned "determining a fourth feature based on the first feature and the second feature" can be achieved by calculating the difference between the vector corresponding to the first feature and the vector corresponding to the second feature to obtain the fourth feature.
[0085] Here, the difference refers to the numerical difference between two vectors along the same dimension.
[0086] For example, the first feature is a vector. The second feature is a vector. Then the fourth feature is .
[0087] In step 1023B, the third feature and the fourth feature are fused to obtain the target feature.
[0088] In this context, the target feature refers to the comprehensive feature data that can be used for final prediction, formed by integrating the third feature representing the task's workload and the fourth feature representing hardware differences. Fusion refers to the process of combining two or more feature data into a single feature representation. Fusion can involve concatenating two feature vectors together, or combining them through weighted summation, element-wise multiplication, or other methods.
[0089] In some embodiments, the above-mentioned "fusion of the third feature and the fourth feature to obtain the target feature" can be achieved by concatenating the vector corresponding to the third feature and the vector corresponding to the fourth feature to obtain the target feature.
[0090] Among them, concatenation refers to appending the element sequence of one vector to the element sequence of another vector.
[0091] For example, the third feature is a vector. The fourth feature is a vector. The target features are .
[0092] In step 1024B, a pre-trained neural network model is invoked based on the target features. The performance data of the target hardware device during model training based on the candidate training strategy is predicted using the pre-trained neural network model to obtain the predicted performance data.
[0093] In this context, a pre-trained neural network model refers to a deep learning model that has been trained on a large amount of data and possesses the ability to infer output results from input features. A pre-trained neural network model can be a prediction tool. Calling a pre-trained neural network model refers to the process of providing the target features as input to the pre-trained neural network model and obtaining its output results.
[0094] For example, the concatenated target feature vector is input into a pre-trained multilayer perceptron model, and the output value of the multilayer perceptron model is used as the prediction time data.
[0095] In this embodiment, feature extraction is performed on the hardware parameters of a known hardware device to obtain a first feature; feature extraction is performed on the hardware parameters of a target hardware device to obtain a second feature; and feature extraction is performed on measured performance data to obtain a third feature. Based on the first and second features, a fourth feature is determined. The fourth feature characterizes the difference between the hardware parameters of the known hardware device and the target hardware device, and can characterize the differences between different hardware environments. The third and fourth features are fused to obtain the target feature, which can construct a comprehensive feature representation that includes the differences and measured performance data. Based on the target feature, a pre-trained neural network model is invoked. Through the pre-trained neural network model, the performance data of the target hardware device during model training based on candidate training strategies is predicted to obtain predicted performance data. This allows for performance prediction using a deep learning model, providing reliable prediction results even when there are significant differences between hardware devices, improving the accuracy and efficiency of performance prediction, and thus improving the optimization efficiency of the training strategy.
[0096] In step 103, the candidate training strategies are prioritized based on the prediction performance data to obtain the priority of the candidate training strategies.
[0097] Priority evaluation refers to the process of analyzing and scoring the applicability or performance of each candidate training strategy on the target hardware device based on evaluation criteria or quantitative indicators. The purpose of priority evaluation is to distinguish the superiority or inferiority among multiple candidate training strategies. For example, priority evaluation can involve calculating the estimated throughput of each training strategy and comparing numerical values, calculating a comprehensive score based on a weighted formula, or screening and grading based on preset rules (such as whether there is memory overflow). Priority refers to the order or level in which candidate training strategies are recommended for selection. Priority can be a numerical sorting index (e.g., 1, 2, 3), a score (e.g., 98 points, 85 points), a classification level (e.g., high, medium, low), or a binary label (e.g., recommended, not recommended).
[0098] In some embodiments, refer to Figure 6 The above step 103 can be achieved by, for example Figure 6 Steps 1031 to 1032 shown are implemented.
[0099] In step 1031, based on the predicted performance data and the hardware parameters of the target hardware device, the candidate training strategy is adapted to the target hardware device to obtain the fit degree between the candidate training strategy and the target hardware device.
[0100] Fit refers to a numerical metric that quantifies the efficiency or feasibility of a training strategy on hardware devices. A higher fit indicates that the training strategy is more suitable for the hardware. Fit can be an estimated throughput value or a normalized quantized score of different dimensional metrics. These normalized scores can be used individually as fit, or they can be combined into a single comprehensive score. For example, a dimensional metric could be effective computing power utilization, used to quantify the degree to which the training strategy utilizes the theoretical peak computing power of the hardware; or a dimensional metric could be pipeline cavitation rate, used to quantify the proportion of idle computing units in the training strategy.
[0101] In some embodiments, the predicted performance data includes prediction time data, and the hardware parameters of the target hardware device include the number of hardware modules; step 1031 above can be implemented as follows: based on the number of target samples and the length of a single sequence of the target samples, the total sequence length of the target samples is determined, and the target samples represent the samples when the target hardware device trains the model based on the candidate training strategy; based on the total sequence length of the target samples, the prediction time data, and the number of hardware modules, the throughput of the target hardware device when training the model based on the candidate training strategy is determined; based on the throughput, the fit between the candidate training strategy and the target hardware device is determined, wherein the throughput is positively correlated with the fit.
[0102] The number of hardware modules refers to the total number of physical units involved in computation within the target hardware device. For example, the number of hardware modules could be the total number of GPUs (e.g., 512) or the number of computing nodes. Target samples refer to the data units input into the model in a single training iteration. Target samples consist of text sequences, image data, or feature vectors. The number of target samples refers to the global batch size (GBS), which is the total number of samples processed by all hardware devices in a single iteration. The length of a single sequence refers to the number of basic data units (e.g., tokens, pixels) contained in a single target sample. For example, the length of a single sequence could be 2048, 4096, 8192, or 32k. The total sequence length refers to the sum of data units processed by all hardware devices in a single iteration. The total sequence length equals the number of target samples multiplied by the length of a single sequence. Throughput refers to the amount of data processed by the hardware device per unit of time. Throughput can be expressed as tokens per second (Tokens Per Second) or tokens per second per card (TGS). A positive correlation means that two variables change in the same direction; if one variable increases, the other variable also increases. In this application embodiment, a higher throughput corresponds to a higher fit score.
[0103] Here, the throughput of the target hardware device when training the model based on the candidate training strategy is determined based on the total sequence length of the target sample, the prediction time data and the number of hardware modules. The following formula (6) can be used as a reference.
[0104] (6) in, Indicates throughput; Indicates the number of target samples; Indicates the length of a single sequence; This represents the predicted time consumption data; Indicates the size of the data parallel group in the candidate training strategy; This represents the number of pipeline parallel stages in the candidate training strategy; Indicates the voiding rate; This indicates the number of hardware modules.
[0105] For example, with 64 target samples, a single sequence length of 8192, a prediction time of 10 seconds, a data parallel group size of 8, a pipeline parallelism level of 8, a cavitation rate of 0.25, and 512 hardware modules, the calculated throughput is 524.288 tokens per second.
[0106] In this embodiment, the prediction performance data includes prediction time data, and the hardware parameters of the target hardware device include the number of hardware modules, thus determining the basic parameters required to calculate throughput. Based on the number of target samples and the length of a single sequence of the target samples, the total sequence length of the target samples is determined. Target samples represent the samples used by the target hardware device to train the model based on candidate training strategies, quantifying the total scale of data processing. Based on the total sequence length of the target samples, the prediction time data, and the number of hardware modules, the throughput of the target hardware device during model training based on candidate training strategies is determined, reflecting the hardware's data processing capability per unit time. Based on the throughput, the fit between the candidate training strategy and the target hardware device is determined. Throughput and fit are positively correlated, allowing a standardized throughput metric to be used as the basis for evaluating fit, ensuring the objectivity and accuracy of the evaluation criteria, thereby improving the optimization efficiency of the training strategy.
[0107] In step 1032, the candidate training strategies are prioritized based on their fitness to obtain their priorities.
[0108] In some embodiments, step 1032 above can be implemented as follows: sorting multiple candidate training strategies based on their compatibility with the target hardware device to obtain a sorting result; and determining the priority of the candidate training strategies based on the sorting result.
[0109] The sorting result refers to the list or index sequence of candidate training strategies arranged from high to low or from low to high according to their fitness.
[0110] In this embodiment, multiple candidate training strategies are ranked based on their compatibility with the target hardware device, resulting in a ranking that allows for a direct analysis of the relative merits of each strategy. Based on this ranking, the priority of the candidate training strategies is determined, enabling rapid identification of the optimal strategy and reducing the time cost of manual screening and comparison, thereby improving the optimization efficiency of the training strategy.
[0111] In some embodiments, step 1032 above can be implemented as follows: when the fit meets the preset fit conditions, the first priority is determined as the priority of the candidate training strategy; when the fit result does not meet the fit conditions, the second priority is determined as the priority of the candidate strategy.
[0112] Here, the adaptation condition refers to the pre-defined logical criteria used to determine whether a training strategy is usable; the first priority refers to the priority label indicating that the training strategy's performance meets the standard; and the second priority refers to the priority label indicating that the training strategy's performance does not meet the standard. The adaptation condition can be that the adaptation value is greater than a preset throughput threshold.
[0113] For example, the preset throughput threshold (fitting condition) is 30,000. If a training strategy calculates a throughput (fitness) of 33,554, which is greater than 30,000, then the training strategy is deemed to meet the fitting condition and is marked as recommended (first priority). If the throughput is 25,000, it is marked as not recommended (second priority).
[0114] In this embodiment, by comparing the fit degree with the fit criteria, candidate training strategies can be quickly classified into binary categories. Determining a first or second priority based on the fit criteria can eliminate substandard training strategies, simplifying the priority evaluation process and improving the efficiency of training strategy selection.
[0115] In some embodiments, step 1032 above can be implemented as follows: obtaining a preset set of fitness intervals, the set of fitness intervals includes multiple fitness intervals, each fitness interval corresponds to a priority; matching the fitness with each fitness interval in the set of fitness intervals to determine the target fitness interval to which the fitness belongs; and determining the priority corresponding to the target fitness interval as the priority of the candidate training strategy.
[0116] Here, the fitness interval set refers to a set of continuous and non-overlapping numerical ranges; the target fitness interval refers to the specific numerical range into which the fitness value falls. The fitness interval can be an inefficient interval, a medium-efficient interval, or a high-efficiency interval.
[0117] For example, we can define 0 to 20000 as the inefficient range, corresponding to a low priority; 20000 to 40000 as the medium-efficient range, corresponding to a medium priority; and above 40000 as the efficient range, corresponding to a high priority. If the fit is 33554, then its target fit range is the medium-efficient range, and its priority is determined to be medium.
[0118] In this embodiment, continuous fitness values are mapped to discrete priority levels through a set of fitness intervals, enabling the hierarchical classification of candidate training strategies. Priority is determined based on the target fitness interval, allowing for the rapid selection of high-performance candidate training strategies from a large pool of options, thus improving the efficiency of training strategy selection.
[0119] In this embodiment, based on predicted performance data and the hardware parameters of the target hardware device, candidate training strategies are adapted to the target hardware device to obtain the compatibility degree between the candidate training strategies and the target hardware device. This allows for analysis of the applicability of the training strategies from the perspective of hardware device compatibility. Based on the compatibility degree, candidate training strategies are prioritized to obtain their priorities. This enables the recommendation of training strategies that are more compatible with the characteristics of the hardware device, avoiding the selection of training strategies with high theoretical performance but unstable operation in practice.
[0120] In step 104, based on the priority of multiple candidate training strategies, a target training strategy for application to the target hardware device is selected from the multiple candidate training strategies.
[0121] The target training strategy refers to the optimal training strategy, determined after screening, that will be configured on the target hardware device to perform model training tasks. The target training strategy is the one with the best overall performance or best meets specific conditions among all candidate training strategies.
[0122] In some embodiments, refer to Figure 7 Step 104 above can be achieved in the following way.
[0123] In step 1041, at least one candidate training strategy with a preset priority is determined from multiple candidate training strategies. In step 1042, when there is at least one candidate training strategy, the candidate training strategy is determined as the target training strategy. In step 1043, when there are multiple candidate training strategies, the target training strategy is selected from the multiple candidate training strategies based on the memory usage of the target hardware device after applying the candidate training strategies.
[0124] Among them, preset priority refers to a specific level defined to meet the screening criteria. For example, preset priority could be the top three priorities or a priority with a score greater than 90. Candidate training strategies refer to the set of preferred training strategies that will proceed to the next round of evaluation after the initial screening. GPU memory usage refers to the amount of graphics memory space required for the model to run. GPU memory usage includes model parameter usage, gradient usage, optimizer state usage, and activation value usage.
[0125] For example, select three high-priority (preset priority) training strategies (candidate training strategies) from the training strategy list. Calculate the estimated VRAM usage of each of the three training strategies on the target hardware device. If training strategy one uses 40GB of VRAM, training strategy two uses 60GB, and training strategy three uses 75GB, and assuming the target hardware device has a single-card VRAM capacity of 80GB, select training strategy one with the lowest VRAM usage as the target training strategy.
[0126] In this embodiment, at least one candidate training strategy with a preset priority is determined from multiple candidate training strategies, enabling the selection of high-quality training strategies in the first tier. When there is only one candidate training strategy, it is designated as the target training strategy, allowing for direct selection. Finally, when there are multiple candidate training strategies, the target training strategy is selected based on the memory usage of the target hardware device after applying the candidate training strategies. This allows for further consideration of memory resource limitations while maintaining similar performance, preventing training failures due to memory overflow and ensuring the selected training strategy has practical usability, thereby improving the optimization efficiency of the training strategy.
[0127] The following will describe an exemplary application of the embodiments of this application in a real-world application scenario.
[0128] With the development of artificial intelligence technology, the parameter scale of large models is constantly increasing. To analyze efficient parallel training strategies for large models (i.e., training strategies), direct deployment on physical machines is generally required; however, this approach consumes enormous human and material resources. For example, training a model with hundreds of billions of parameters often requires building a large-scale hardware cluster consisting of hundreds of training nodes, each equipped with multiple high-performance accelerator cards. Before officially starting training, the team may propose several parallel strategies (such as Strategy 1 and Strategy 2). To determine which strategy is superior, the team decides to conduct practical testing. First, the engineering team may spend several weeks configuring the software environment for Strategy 1, dividing the model into slices, and writing the corresponding communication logic. During the initial training startup, frequent errors occur due to excessive inter-node communication latency, requiring joint troubleshooting and resolution by the operations and algorithm teams. Ultimately, obtaining stable performance data from Strategy 1 may take more than a month. Next, the team begins testing Strategy 2, which again requires code modifications and cluster configuration, investing several more weeks of manpower. The entire evaluation process is lengthy, consuming significant human and cluster operating costs, simply to compare the performance of the two strategies.
[0129] The applicant argues that if the performance differences of different training parallel strategies could be quickly predicted before cluster deployment, significant manpower and resources would be saved. The optimal parallel strategy selected based on performance prediction could then be directly adopted during model cluster deployment. In real-world ultra-large clusters (e.g., 512 accelerator cards), multiple parallel strategies are typically used in combination. For example, tensor parallelism (TP) is performed on 8 accelerator cards within a single machine; pipelined parallelism (PP) is performed across multiple machines; and data parallelism (DP) is performed on the remaining dimensions. The difficulty in implementing this in practical applications lies in determining parameters such as the number of tensor parallel partitions, the number of pipelined parallel partitions, and the micro-batch size across the 512 accelerator cards without strategy prediction. Using existing technologies requires writing code for configuration, repartitioning the model, reloading data, and conducting actual tests. If communication congestion is detected during this process, the process must be restarted.
[0130] The main objective of this application is to solve the technical problems of high cost, long time consumption, and complex operation and maintenance in the prior art of directly deploying large-scale hardware acceleration devices for training parallel strategy performance testing.
[0131] To address the aforementioned technical issues, this application proposes a strategy optimization method that can evaluate the training performance of models with the same structure but different parameter sets on other hardware acceleration devices based on model training performance data from known hardware acceleration devices. By analyzing key performance events during model training, the method can quickly obtain model training performance data and metrics on new devices using performance data from existing devices. This allows for the comparison of large model training efficiency under different parallel strategies without deploying a cluster, requiring only a small number of hardware acceleration devices, thus enabling the selection of the final large model training parallel strategy for deployment on the cluster.
[0132] The following is a detailed description of the specific implementation of a strategy optimization method proposed in this application scenario.
[0133] It should be noted that, in the embodiments of this application, the first type of hardware acceleration device is defined as a hardware acceleration device with known model performance training performance data (i.e., known hardware device), and the second type of hardware acceleration device is defined as a hardware acceleration device that needs to estimate model performance (i.e., target hardware device).
[0134] First, the model parameters and parallel strategy (i.e., candidate training strategy) are obtained. This step involves acquiring the model parameters on the first type of hardware acceleration device and the second type. As an example, the model structures deployed on the first and second hardware acceleration devices are identical, both consisting of the same model layers, but the number of layers differs. For instance, the model on the first hardware acceleration device has 1 layer (i.e., the known number of model layers), meaning the first device has 1 model layer; the model on the second hardware acceleration device has 64 layers (i.e., the target number of model layers), meaning the second device has 64 model layers. Furthermore, the total number of batches and the number of micro-batches are set to 1 on the first hardware acceleration device; the number of batches on the second hardware acceleration device is 64, with 1 batch size. The input data length (sequence_length) for both devices is 8192. This step determines the parallel strategy on the first hardware acceleration device and preliminarily determines the parallel strategy on the second hardware acceleration device for which model performance needs to be estimated. For example, the first type of hardware acceleration device uses 8 accelerator cards, employing tensor parallelism with a total of 8 parallelizations. The second type of hardware acceleration device uses 512 accelerator cards, employing tensor parallelism (TP) with 8 parallelizations, pipeline parallelism (PP) with 8 parallelizations, expert parallelism with 8 parallelizations, and data parallelism (DP) with 8 parallelizations. Based on the pipeline parallelism arrangement in the second type of hardware acceleration device, the bubbleratio (br) is determined.
[0135] Next, the hardware parameters are obtained. In this step, the hardware parameters of the first type of hardware acceleration device (i.e., the hardware parameters of the known hardware device) and the hardware parameters of the second type of hardware acceleration device (i.e., the hardware parameters of the target hardware device) are obtained. The hardware parameters include computing power, memory access bandwidth, and various communication bandwidths.
[0136] As examples, the hardware parameters of the first type of hardware acceleration device are as follows: BFLOAT16 computing power is 148 TFLOPS, FP32 computing power is 44 TFLOPS, memory bandwidth is 4 TB / s, inter-card interconnect bandwidth is 900 GB / s, and inter-machine interconnect bandwidth is 50 GB / s. The hardware parameters of the second type of hardware acceleration device are as follows: BFLOAT16 computing power is 312 TFLOPS, FP32 computing power is 19.5 TFLOPS, memory bandwidth is 2.039 TB / s, inter-card interconnect bandwidth is 400 GB / s, and inter-machine interconnect bandwidth is 50 GB / s.
[0137] Next, the performance event decomposition and first device efficiency evaluation are performed. In this step, the model training process on the first type of hardware acceleration device is broken down into several core performance events that account for the end-to-end time consumption, including General Matrix Multiplication (GEMM) execution events, Vector execution events, FlashAttention (FLA) execution events, communication events, and host-to-device / device-to-host (H2D / D2H) events. The applicant believes that these performance events are easier to evaluate in terms of efficiency on a single hardware device, or to obtain performance data for testing more conveniently and quickly, compared to the complete model. In this step, the utilization efficiency of different performance events is evaluated for the first type of hardware acceleration device. Specifically, raw measured data (i.e., measured performance data) on the first type of hardware acceleration device is recorded using a performance analysis tool (Profiler). Among them, the computational power utilization rate of GEMM execution events is r1, the computational power utilization rate of FlashAttention execution events is r2, the utilization rate of Vector execution events is r3, the interconnect bandwidth utilization rate is r4, and the inter-device interconnect bandwidth utilization rate is r5 (i.e., the utilization rate of known hardware devices).
[0138] Subsequently, the mapping and second device efficiency evaluation are performed. In this step, all the split performance events from the first hardware acceleration device are mapped to the second hardware acceleration device, and the utilization efficiency of various performance events is evaluated. The model training on the second hardware acceleration device also includes the same types of performance events as the model on the first hardware acceleration device. Because the model on the second hardware acceleration device contains more layers, the number of different types of performance events will increase, but the types are the same as those on the first hardware acceleration device. The execution utilization of each type of event is evaluated on the second hardware acceleration device: the computational power utilization of GEMM type execution events is a1, the computational power utilization of FlashAttention execution events is a2, the utilization of Vector type execution events is a3, the interconnect bandwidth utilization is a4, and the inter-machine interconnect bandwidth utilization is a5 (i.e., the utilization of the target hardware device).
[0139] Finally, time consumption estimation and overall performance calculation: In this step, based on the hardware parameters on the second type of hardware acceleration device and the various performance events on the second type of hardware acceleration device obtained in step four, the execution efficiency (a1, a2, a3, a4, a5) is used to estimate the time consumption of all performance events, and all events are summarized to obtain an overall performance estimate of model training on the second type of hardware acceleration device.
[0140] The following details the specific estimation method based on the aforementioned parameters.
[0141] In this embodiment of the application, the total computational load or total communication load (i.e., data load) of various events on the second type of hardware acceleration device is calculated, specifically referring to the following formulas (7) to (11).
[0142] (7) in, This represents the total computational load of GEMM-type events for the second type of hardware acceleration device; This indicates the total time taken for GEMM class events of the first type of hardware acceleration device; This indicates the total number of layers in the target large model on the second type of hardware acceleration device; This indicates the number of model layers actually running and being tested on the first type of hardware acceleration device; This represents the GEMM computing power of the first type of hardware acceleration device, BFLOAT16. This represents the computing power utilization rate of GEMM-type events on the first type of hardware acceleration device.
[0143] (8) in, This represents the total computational load of FLA-type events from the second type of hardware acceleration device; This indicates the total time taken for FLA-type events of the first type of hardware acceleration device; This indicates the total number of layers in the target large model on the second type of hardware acceleration device; This indicates the number of model layers actually running and being tested on the first type of hardware acceleration device; This indicates the FLA computing power of the first type of hardware acceleration device, BFLOAT16. This indicates the computing power utilization rate of FLA-type events on the first type of hardware acceleration device.
[0144] (9) in, This represents the total computational cost of Vector-class events for the second type of hardware acceleration device; This represents the total event duration for the Vector class of the first type of hardware acceleration device; This indicates the total number of layers in the target large model on the second type of hardware acceleration device; This indicates the number of model layers actually running and being tested on the first type of hardware acceleration device; This represents the Vector computing power of the first type of hardware acceleration device, BFLOAT16. This represents the computational efficiency of Vector-type events on the first type of hardware acceleration device.
[0145] (10) in, This represents the total communication volume of the second type of hardware acceleration device inter-card communication. This indicates the total communication time between the first type of hardware acceleration device cards; This indicates the inter-card communication bandwidth of the first type of hardware acceleration device; This indicates the utilization rate of interconnect bandwidth on the first type of hardware acceleration device.
[0146] (11) in, This represents the total communication volume of the second type of hardware acceleration device inter-card communication. This indicates the total communication time between the first type of hardware acceleration device cards; This indicates the inter-card communication bandwidth of the first type of hardware acceleration device; This indicates the utilization rate of interconnect bandwidth on the first type of hardware acceleration device.
[0147] In this embodiment of the application, based on the total computing power / communication power calculated above, the time consumption of various events on the second type of hardware acceleration device (i.e., the predicted time consumption data) is calculated, specifically referring to the following formulas (12) to (17).
[0148] (12) in, This indicates the time taken for GEMM class events of the second type of hardware acceleration device; a1 represents the GEMM computing power of the second type of hardware acceleration device BFLOAT16; a1 represents the estimated hardware utilization of GEMM-type events mapped to the second type of hardware acceleration device.
[0149] (13) in, This indicates the FLA event duration for the second type of hardware acceleration device; a1 represents the FLA computing power of the second type of hardware acceleration device BFLOAT16; a2 represents the estimated hardware utilization of FLA-type events mapped to the second type of hardware acceleration device.
[0150] (14) in, This indicates the event duration of the Vector class for the second type of hardware acceleration device; a3 represents the computing power of the second type of hardware acceleration device, Vector; a3 represents the estimated hardware utilization of Vector-type events mapped to the second type of hardware acceleration device.
[0151] (15) in, This indicates the time taken for communication events between the second type of hardware acceleration device cards; a4 represents the inter-card communication bandwidth of the second type of hardware acceleration device; a4 represents the estimated utilization of the interconnect bandwidth mapped to the second type of hardware acceleration device.
[0152] (16) in, This indicates the time taken for inter-device communication events in the second type of hardware acceleration device; a5 represents the inter-device communication bandwidth of the second type of hardware acceleration device; a5 represents the estimated utilization rate of the inter-device interconnect bandwidth mapped to the second type of hardware acceleration device.
[0153] Then, the time taken for all events is summed up to obtain the total system time.
[0154] (17) in, This indicates the total time consumed by the second type of hardware acceleration device.
[0155] In this embodiment of the application, the number of tokens processed per second per card during model training is calculated using the model training parameters, specifically according to the following formula (18).
[0156] (18) Where TGS represents the number of tokens processed per second by each accelerator card (i.e., throughput); gbs represents the global batch size, which is the total number of data samples processed by all accelerator cards in one training iteration (i.e., the number of target samples). This indicates the sequence length, which is the number of tokens (words / characters) contained in a data sample; This represents the total time consumed by the second type of hardware acceleration device as calculated above; Indicates the number of data items processed in parallel; br represents the number of parallel pipelines; br represents the cavitation rate. When using pipeline parallelism (PP), the accelerator card will be idle for a period of time due to waiting for data dependencies. The proportion of this idle time to the total time is the cavitation rate. This indicates the total number of cards (i.e., the number of hardware modules) for the second type of hardware acceleration device.
[0157] This application embodiment measures the basic computing and communication capabilities through small-scale testing (e.g., on a first type of hardware acceleration device), and then simulates the splitting and communication process through mathematical formulas to directly calculate how much slower a large-scale cluster (e.g., 512 cards) would be if not properly combined, thereby saving the cost of physical deployment.
[0158] To more clearly illustrate the technical solutions of the embodiments of this application, refer to... Figure 8 , Figure 8This is a flowchart illustrating an application scenario of a strategy optimization method provided in this application embodiment. In step 801, model parameters and training parameters are set on a first type of device; in step 802, hardware parameters of the first type of device are set; in step 803, model performance test data is collected; and in step 804, various types of performance events are segmented. Simultaneously, in step 805, model parameters and training parameters are set on a second type of device; and in step 806, hardware parameters of the second type of device are set. Next, in step 807, based on the model performance test data and performance events of the first type of device, the test data of performance events of the second type of device are mapped; in step 808, based on the test data of performance events of the second type of device, the execution efficiency of performance events of the second type of device is calculated; finally, in step 809, the model training performance system index is determined based on the execution efficiency of performance events of the second type of device.
[0159] Reference Figure 9 , Figure 9 This diagram illustrates the technical framework of the strategy optimization method provided in this application embodiment. It mainly includes an input layer, a processing layer, and an output layer. The input layer provides basic data, including source device hardware parameters, target device hardware parameters, model parameters, parallel strategies, and measured performance data of the source device. The processing layer is the core of performance prediction, internally performing a series of processes on the input data, including performance event decomposition, cross-device efficiency mapping, target device time estimation, and overall performance calculation. The output layer produces the final prediction results, including prediction time data, prediction throughput, and comprehensive performance indicators of the parallel strategy, thereby providing a decision-making basis for strategy selection (i.e., the target training strategy).
[0160] The strategy optimization method provided in this application has the following beneficial effects: First, the embodiments of this application can evaluate the training performance of models with the same model structure but different parameter amounts on other hardware acceleration devices based on real test data of existing model performance on one hardware acceleration device, without the need to directly deploy large-scale hardware devices.
[0161] Second, the method in this application embodiment can evaluate the performance of a specified hardware device under different parallel strategies in a short time.
[0162] Third, compared with the method of directly deploying hardware devices to test the training performance of the model, the method of this application embodiment significantly reduces the performance evaluation and comparison time, greatly improves efficiency, and the performance trends of different model training parallel strategies are basically consistent with the actual test.
[0163] The following description continues to illustrate the exemplary structure of the strategy optimization device 455 provided in the embodiments of this application as a software module. In some embodiments, such as... Figure 2 As shown, the software modules stored in the strategy optimization device 455 in the memory 450 may include: The acquisition module 4551 is used to acquire multiple candidate training strategies and acquire measured performance data of known hardware devices when training a model based on each of the candidate training strategies. Prediction module 4552 is used to predict the performance data of the target hardware device when training the model based on the candidate training strategy, based on the measured performance data of the known hardware device when training the model based on the candidate training strategy, the hardware parameters of the known hardware device, and the hardware parameters of the target hardware device, for each candidate training strategy, so as to obtain predicted performance data. Evaluation module 4553 is used to evaluate the priority of the candidate training strategies based on the prediction performance data, and obtain the priority of the candidate training strategies. The filtering module 4554 is used to filter out the target training strategy for application to the target hardware device from the multiple candidate training strategies based on the priority of the multiple candidate training strategies.
[0164] In some embodiments, the prediction module 4552 is further configured to determine the effective performance parameters of the known hardware device based on the hardware parameters of the known hardware device and the utilization rate of the known hardware device, and to determine the effective performance parameters of the target hardware device based on the hardware parameters of the target hardware device and the utilization rate of the target hardware device; and to predict the performance data of the target hardware device when training the model based on the candidate training strategy based on the effective performance parameters of the known hardware device, the effective performance parameters of the target hardware device, and the measured performance data of the known hardware device when training the model based on the candidate training strategy, thereby obtaining predicted performance data.
[0165] In some embodiments, the measured performance data includes measured time consumption data, and the predicted performance data includes predicted time consumption data; the prediction module 4552 is further configured to determine the data load of the target hardware device when training the model based on the candidate training strategy, based on the effective performance parameters of the known hardware device and the measured time consumption data of the known hardware device when training the model based on the candidate training strategy; and to determine the predicted time consumption data based on the data load and the effective performance parameters of the target hardware device.
[0166] In some embodiments, the prediction module 4552 is further configured to determine the layer ratio between the known model layer number and the target model layer number, wherein the known model layer number represents the number of model layers when the model is trained on the known hardware device, and the target model layer number represents the number of model layers when the model is trained on the target hardware device; and based on the layer ratio, the effective performance parameters of the known hardware device, and the measured time consumption data of the known hardware device when training the model based on the candidate training strategy, determine the data load of the target hardware device when training the model based on the candidate training strategy.
[0167] In some embodiments, the prediction module 4552 is further configured to extract features from the hardware parameters of the known hardware device to obtain a first feature; extract features from the hardware parameters of the target hardware device to obtain a second feature; and extract features from the measured performance data to obtain a third feature; determine a fourth feature based on the first feature and the second feature, the fourth feature representing the difference between the hardware parameters of the known hardware device and the hardware parameters of the target hardware device; fuse the third feature and the fourth feature to obtain a target feature; and call a pre-trained neural network model based on the target feature, and predict the performance data of the target hardware device during model training based on the candidate training strategy using the pre-trained neural network model to obtain predicted performance data.
[0168] In some embodiments, the evaluation module 4553 is further configured to perform adaptation processing on the candidate training strategy and the target hardware device based on the predicted performance data and the hardware parameters of the target hardware device, to obtain the adaptation degree between the candidate training strategy and the target hardware device; and to perform priority evaluation on the candidate training strategy based on the adaptation degree, to obtain the priority of the candidate training strategy.
[0169] In some embodiments, the prediction performance data includes prediction time data, and the hardware parameters of the target hardware device include the number of hardware modules; the evaluation module 4553 is further configured to determine the total sequence length of the target samples based on the number of target samples and the individual sequence length of the target samples, wherein the target samples represent the samples when the target hardware device performs model training based on the candidate training strategy; determine the throughput of the target hardware device when performing model training based on the candidate training strategy based on the total sequence length of the target samples, the prediction time data, and the number of hardware modules; and determine the fit between the candidate training strategy and the target hardware device based on the throughput, wherein the throughput is positively correlated with the fit.
[0170] In some embodiments, the evaluation module 4553 is further configured to rank the plurality of candidate training strategies based on their compatibility with the target hardware device, and obtain a ranking result; and determine the priority of the candidate training strategies based on the ranking result.
[0171] In some embodiments, the filtering module 4554 is further configured to determine at least one candidate training strategy with a preset priority from the plurality of candidate training strategies; when there is only one candidate training strategy, the candidate training strategy is determined as the target training strategy; when there are multiple candidate training strategies, the target training strategy is selected from the plurality of candidate training strategies based on the memory usage of the target hardware device after applying the candidate training strategy.
[0172] This application provides a computer program product including a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the strategy optimization method described above in this application.
[0173] This application provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the strategy optimization method provided in this application. For example, ... Figure 3 The strategy optimization method is shown.
[0174] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.
[0175] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.
[0176] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).
[0177] As an example, computer-executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.
[0178] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.
Claims
1. A strategy optimization method, characterized in that, The method includes: Multiple candidate training strategies are obtained, and measured performance data of known hardware devices when training a model based on each of the candidate training strategies are obtained. For each candidate training strategy, based on the hardware parameters and utilization rate of the known hardware device, the effective performance parameters of the known hardware device are determined, and based on the hardware parameters and utilization rate of the target hardware device, the effective performance parameters of the target hardware device are determined. Based on the effective performance parameters of the known hardware device, the effective performance parameters of the target hardware device, and the measured performance data of the known hardware device when training the model based on the candidate training strategy, the performance data of the target hardware device when training the model based on the candidate training strategy is predicted to obtain predicted performance data. Based on the prediction performance data, the candidate training strategies are prioritized to obtain their priorities; wherein, the candidate training strategy is a pre-defined alternative scheme to guide the model in computational allocation and data interaction on hardware devices before training the model. Based on the priority of the multiple candidate training strategies, a target training strategy for application to the target hardware device is selected from the multiple candidate training strategies.
2. The method according to claim 1, characterized in that, The measured performance data includes measured time data, and the predicted performance data includes predicted time data; The method involves predicting the performance data of the target hardware device during model training based on the known hardware device's effective performance parameters, the target hardware device's effective performance parameters, and the measured performance data of the known hardware device during model training using the candidate training strategy, to obtain predicted performance data, including: Based on the effective performance parameters of the known hardware device and the measured time consumption data of the known hardware device when training the model based on the candidate training strategy, the data load of the target hardware device when training the model based on the candidate training strategy is determined. The predicted time consumption data is determined based on the data load and the effective performance parameters of the target hardware device.
3. The method according to claim 2, characterized in that, The determination of the data load of the target hardware device during model training based on the candidate training strategy, based on the effective performance parameters of the known hardware device and the measured time consumption data of the known hardware device during model training based on the candidate training strategy, includes: Determine the layer ratio between the known model layer number and the target model layer number, where the known model layer number represents the number of model layers when the model is trained on the known hardware device, and the target model layer number represents the number of model layers when the model is trained on the target hardware device; Based on the layer ratio, the effective performance parameters of the known hardware device, and the measured time consumption data of the known hardware device when training the model based on the candidate training strategy, the data load of the target hardware device when training the model based on the candidate training strategy is determined.
4. The method according to claim 1, characterized in that, The step of prioritizing the candidate training strategies based on the prediction performance data to obtain the priority of the candidate training strategies includes: Based on the predicted performance data and the hardware parameters of the target hardware device, the candidate training strategy is adapted to the target hardware device to obtain the adaptation degree between the candidate training strategy and the target hardware device. Based on the fitness level, the candidate training strategies are prioritized to obtain their priorities.
5. The method according to claim 4, characterized in that, The predicted performance data includes predicted time data, and the hardware parameters of the target hardware device include the number of hardware modules. The step of adapting the candidate training strategy to the target hardware device based on the predicted performance data and the hardware parameters of the target hardware device to obtain the fit degree between the candidate training strategy and the target hardware device includes: The total sequence length of the target samples is determined based on the number of target samples and the individual sequence length of the target samples. The target samples represent the samples used by the target hardware device to train the model based on the candidate training strategy. Based on the total sequence length of the target sample, the prediction time data, and the number of hardware modules, determine the throughput of the target hardware device when training the model based on the candidate training strategy. Based on the throughput, the fit between the candidate training strategy and the target hardware device is determined, wherein the throughput and the fit are positively correlated.
6. The method according to claim 4, characterized in that, The step of prioritizing the candidate training strategies based on the fitness level to obtain the priority of the candidate training strategies includes: Based on the compatibility between the multiple candidate training strategies and the target hardware device, the multiple candidate training strategies are ranked to obtain a ranking result. Based on the ranking results, the priority of the candidate training strategies is determined.
7. The method according to claim 1, characterized in that, The step of selecting a target training strategy for application to the target hardware device from the multiple candidate training strategies based on their priority includes: From the plurality of candidate training strategies, at least one candidate training strategy with a preset priority is determined; When there is only one candidate training strategy, the candidate training strategy is determined as the target training strategy. When there are multiple candidate training strategies, the target training strategy is selected from the multiple candidate training strategies based on the memory usage of the target hardware device after applying the candidate training strategy.
8. A strategy optimization device, characterized in that, The device includes: The acquisition module is used to acquire multiple candidate training strategies and acquire measured performance data of known hardware devices when training a model based on each of the candidate training strategies. The prediction module is used to determine the effective performance parameters of the known hardware device based on the hardware parameters and utilization rate of the known hardware device for each candidate training strategy, and to determine the effective performance parameters of the target hardware device based on the hardware parameters and utilization rate of the target hardware device; based on the effective performance parameters of the known hardware device, the effective performance parameters of the target hardware device, and the measured performance data of the known hardware device when training the model based on the candidate training strategy, the module predicts the performance data of the target hardware device when training the model based on the candidate training strategy, thereby obtaining predicted performance data. An evaluation module is used to evaluate the priority of the candidate training strategies based on the prediction performance data, and obtain the priority of the candidate training strategies; wherein, the candidate training strategies are alternative schemes pre-set before training the model to guide the model in computational allocation and data interaction on the hardware device. The filtering module is used to filter out the target training strategy for application to the target hardware device from the multiple candidate training strategies based on their priority.
9. An electronic device, characterized in that, The electronic device includes: Memory is used to store executable instructions or computer programs. A processor, when executing computer-executable instructions or computer programs stored in the memory, implements the method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, they implement the method described in any one of claims 1 to 7.
11. A computer program product comprising computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, they implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Model deployment method and device, storage medium and electronic equipment
CN119883295A