A model optimization method and related device
By identifying and splitting the computational operations in the neural network model and reconfiguring them to multiple hardware computing units for parallel execution, the problems of waste of hardware resources and inefficiency caused by serial computing operations in the prior art are solved, and more efficient computing and hardware utilization are achieved.
Patent Information
- Application Number
- CN202411649898.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-19
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-11-19
AI Technical Summary
In the prior art, the computational operations of neural network models are usually serial, resulting in waste of hardware resources and low model operation efficiency.
By identifying the calculation operations that meet the splitting conditions in the target model, splitting them into multiple sub-computing operations, and reconfiguring these sub-computing operations on multiple hardware computing units, parallel execution is achieved.
It significantly improves the computing efficiency, hardware utilization, flexibility and real-time performance of the neural network model, and reduces the computational operation delay.
Smart Images

Figure CN119166159B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of artificial intelligence, and particularly relates to a model optimization method and related devices. Background Art
[0002] In the modern field of artificial intelligence, especially in machine learning and deep learning applications, the efficiency and performance of algorithms are crucial. Artificial neural network models, such as the Generative Pretrained Transformer (GPT) model and the Bloom large model in large language models (LLMs), require a large amount of computing resources for data processing and model training. These processes mainly consist of many complex mathematical calculations. Therefore, how to efficiently utilize hardware resources to improve the operation efficiency is of great importance.
[0003] In traditional technologies, the model calculation process is usually serial. After waiting for the previous calculation operation to complete and obtain the calculation result, the next calculation operation is then started. The waiting process will cause other hardware to be in an idle state, resulting in a waste of hardware resources and greatly reducing the model operation efficiency.
[0004] Therefore, there is an urgent need to design a brand-new technical solution to overcome at least one of the above technical problems. Summary of the Invention
[0005] This application provides a model optimization method and related devices, which are used to redistribute the calculation operations of a neural network model to multiple hardware operation units, thereby improving the calculation efficiency, hardware utilization rate, flexibility, and real-time performance of the neural network model, reducing the calculation operation delay, performing hardware acceleration operations on the model, and enhancing the performance of the neural network model.
[0006] In a first aspect, this application provides a model optimization method, which is applied to a hardware computing system deployed with a target model, and the target model belongs to a neural network model; it includes:
[0007] Identify the calculation operations in the target model that meet the splitting conditions;
[0008] Split the calculation operations to obtain multiple sub-calculation operations;
[0009] During the process of loading the target model into the hardware computing system, reconfigure the multiple sub-calculation operations into the corresponding target operation units; the hardware computing system includes at least multiple operation units, and each operation unit is composed of hardware chip resources called based on assembly instructions;
[0010] Execute the respective sub-calculation operations through the target operation units to obtain multiple sub-calculation results;
[0011] Apply the multiple sub-calculation results to the data processing flow of the target model.
[0012] In a second aspect, an embodiment of the present application provides a model optimization device, which is applied to a hardware computing system on which a target model is deployed, and the target model belongs to a neural network model; the device at least includes the following units:
[0013] An identification unit, configured to identify computational operations in the target model that meet the splitting conditions;
[0014] A splitting unit, configured to split the computational operations to obtain a plurality of sub-computational operations;
[0015] A configuration unit, configured to reconfigure the plurality of sub-computational operations into corresponding target operation units during the process of loading the target model into the hardware computing system; the hardware computing system at least includes a plurality of operation units, and each operation unit is composed of hardware chip resources called based on assembly instructions;
[0016] An execution unit, configured to execute the respective corresponding sub-computational operations through the target operation units to obtain a plurality of sub-computation results;
[0017] An application unit, configured to apply the plurality of sub-computation results to the data processing flow of the target model.
[0018] In a third aspect, an embodiment of the present application provides a chip for implementing the model optimization method described in the first aspect.
[0019] In a fourth aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory, and the processor executes the computer program to implement the model optimization method described in the first aspect.
[0020] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed, the model optimization method described in the first aspect is implemented.
[0021] In the embodiments of the present application, it is possible to identify computational operations in a target model that meet the splitting conditions; split the computational operations to obtain multiple sub-computational operations; during the process of loading the target model into the hardware computing system, reconfigure the multiple sub-computational operations into corresponding target computing units; the hardware computing system includes at least multiple computing units, and each computing unit is composed of hardware chip resources called based on assembly instructions; execute the respective corresponding sub-computational operations through the target computing units to obtain multiple sub-computation results; apply the multiple sub-computation results to the data processing flow of the target model. The embodiments of the present application can significantly improve the computing efficiency, hardware utilization, flexibility, and real-time performance, reduce the latency of computational operations, perform hardware acceleration operations for neural network models, enhance the performance of neural network models, and provide strong hardware optimization support for various complex computing tasks by identifying, splitting, reconfiguring, and parallelly executing computational operations in neural network models. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation to the present application. In the drawings:
[0023] Figure 1 is a flowchart of a model optimization method according to an embodiment of the present application;
[0024] Figure 2 A schematic diagram of the principle of a model optimization process according to an embodiment of the present application;
[0025] Figure 3 A schematic diagram of the principle of a model splitting process according to an embodiment of the present application;
[0026] Figure 4 Another schematic diagram of the principle of a model splitting process according to an embodiment of the present application;
[0027] Figure 5 is a structural block diagram of a model optimization device according to an embodiment of the present application;
[0028] Figure 6 is a structural block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0030] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs. The terms used in the description of this application herein are for the purpose of describing specific embodiments only and are not intended to limit this application.
[0031] In the field of artificial intelligence, especially in machine learning and deep learning applications, algorithm efficiency and computing performance are crucial. Currently, most model computations adopt serial computing. Serial computing requires subsequent computing operations to wait until all the dependent data is fully prepared before proceeding, and during the waiting process, the hardware resources required for subsequent operations will be in an idle state.
[0032] In related technologies, for the hardware device on which a neural network model is deployed, if there are multiple computing units in the hardware device (such as computing units for processing matrix addition and matrix multiplication), then tensor parallelism can be used to improve the efficiency of model operation. Currently, under the condition of deploying multiple graphics processing units (GPUs), the method of data parallelism is usually adopted to accelerate model operation. In this method, a unified neural network model is loaded in each GPU, and after each iteration, all neural network models synchronize their respective states with each other. This method can speed up the training speed by investing more GPU resources and solve the problem.
[0033] In related technologies, the premise of data parallelism is that each GPU can accommodate the entire model. If the model is too large and exceeds the memory limit of a single GPU, then the model cannot be loaded on that GPU, and thus training cannot be carried out. This means that for very large models (such as some modern deep learning models like GPT-3), a single GPU cannot support them, and difficulties will be encountered when using data parallelism. In addition, in multi-GPU training, a certain communication time is required for gradient synchronization between each GPU. This is especially obvious when the number of GPUs is very large, which may lead to the idleness of computing nodes. Therefore, in some cases, increasing the number of GPUs will instead introduce additional latency.
[0034] It should also be noted that for tasks with a small amount of data, since full-model training on each GPU may lead to duplicate processing of information and waste of resources, the efficiency of data parallelism may be very low.
[0035] Therefore, there is an urgent need to design a brand-new technical solution to overcome at least one of the above technical problems.
[0036] To this end, the embodiments of the present application provide a model optimization method and related devices, which can identify computational operations in a target model that meet the splitting conditions; split the computational operations to obtain multiple sub-computational operations; during the process of loading the target model into a hardware computing system, reconfigure the multiple sub-computational operations into corresponding target computing units; the hardware computing system includes at least multiple computing units, and each computing unit is composed of hardware chip resources called by assembly instructions; execute the respective corresponding sub-computational operations through the target computing units to obtain multiple sub-computation results; apply the multiple sub-computation results to the data processing flow of the target model. The embodiments of the present application can significantly improve the computing efficiency, hardware utilization, flexibility, and real-time performance, reduce the latency of computational operations, perform hardware acceleration operations for neural network models, improve the performance of neural network models, and provide strong hardware optimization support for various complex computing tasks by identifying, splitting, reconfiguring, and parallelly executing computational operations in neural network models.
[0037] The model optimization solution provided by the embodiments of the present application can also be executed by an electronic device, which can be a server, a server cluster, or a cloud server. The electronic device can also be a terminal device such as a mobile phone, a computer, a tablet computer, a wearable device, or a dedicated device (such as a dedicated terminal device with a model optimization system, etc.). The above-mentioned chips introduced in the above embodiments can also be installed in these electronic devices. Or, these electronic devices can also install a service program for executing the model optimization solution.
[0038] The following describes the model optimization method and related devices provided by the embodiments of the present application with reference to the accompanying drawings. Figure 1 A model optimization method provided by an embodiment of the present application is as Figure 1 shown, and the method includes:
[0039] S101. Identify computational operations in the target model that meet the splitting conditions;
[0040] S102. Split the computational operations to obtain multiple sub-computational operations;
[0041] S103. During the process of loading the target model into the hardware computing system, reconfigure the multiple sub-computational operations into corresponding target computing units;
[0042] S104. Execute the respective corresponding sub-computational operations through the target computing units to obtain multiple sub-computation results;
[0043] S105. Apply the multiple sub-computation results to the data processing flow of the target model.
[0044] In the embodiments of the present application, the hardware computing system and the arithmetic units are key components of the model optimization method. In the embodiments of the present application, the hardware computing system at least includes a plurality of arithmetic units, and each arithmetic unit is composed of hardware chip resources called based on assembly instructions.
[0045] The hardware computing system is an integrated computing platform designed to provide efficient computing power to support complex model operations and data processing tasks. This system is typically composed of multiple arithmetic units, memory, storage, and other supporting components. Specifically, the arithmetic units (Processing Units) are the core part of the computing, which can include different types of processors, such as the Central Processing Unit (CPU) for general computing, suitable for processing complex logic and control tasks. The Graphics Processing Unit (GPU), specially designed for parallel processing of graphics computing, is very suitable for deep learning and large-scale matrix operations. The Tensor Processing Unit (TPU), such as a dedicated chip, is designed to accelerate machine learning tasks. And the Field Programmable Gate Array (FPGA), a programmable logic device, can be efficiently configured for specific tasks. In addition, the hardware computing system also includes memory (Memory), such as Random Access Memory (RAM) and cache, for storing the data and model parameters being processed for quick access and calculation. Storage, including hard disks and solid state drives, for persistent storage of models, data, and calculation results. Buses and interfaces (Bus and Interfaces) for data transfer between different components, including data buses, address buses, and control buses. And a power management module to ensure the stable operation of hardware components and optimize energy consumption.
[0046] The arithmetic units are the basic computing units in the hardware computing system, each responsible for specific computing tasks. The arithmetic units are composed of hardware chip resources and can be called using specific assembly instruction sets. These instruction sets provide descriptions of the basic operations for the arithmetic units, making the execution of instructions more efficient. Each arithmetic unit can be configured to perform different types of operations. For example, the CPU is good at serial processing of logic and complex calculations, while the GPU is suitable for large-scale parallel computing. The FPGA can be reconfigured according to specific needs to implement specific logical operations. The task scheduling and resource allocation of the arithmetic units aim to improve the system operation efficiency. For example, the allocation of sub-computing operations can be dynamically adjusted according to the load and capabilities of the arithmetic units to ensure efficient utilization of hardware resources. There may be a need for efficient data transfer and communication between the arithmetic units to cooperate in completing more complex computing tasks. In the model optimization method, the results of sub-computing operations can be transferred between different arithmetic units to achieve the integration of results.
[0047] Further optionally, the arithmetic unit supports basic arithmetic operations, matrix multiplication, matrix transpose, synchronization between arithmetic units, data transmission between arithmetic units, etc., or can indirectly implement the above functions.
[0048] In the embodiments of the present application, as a whole, the hardware computing system integrates multiple arithmetic units and realizes efficient optimization of the model by identifying, splitting, and reconfiguring computing operations. In specific implementation, the performance of various types of arithmetic units can be fully utilized to ensure efficient processing of complex data and computing tasks. This system architecture provides strong hardware support for fields such as machine learning and deep learning and can adapt to changing computing requirements.
[0049] Exemplarily, the hardware computing system can be a chip, a system composed of multiple chips, or a system composed of other hardware resources.
[0050] In practical applications, on a hardware architecture with multiple arithmetic units, through the above steps, complex data calculations can also be divided into multiple data streams to achieve distributed stream computing. While fully utilizing hardware resources, the model operation efficiency can be improved through parallel computing.
[0051] Specifically, on a hardware architecture with multiple arithmetic units, by dividing complex data calculations into multiple data streams to achieve distributed stream computing, the operation efficiency of the model can be significantly improved while fully utilizing hardware resources. In practical applications, complex data computing tasks can be divided into multiple small data streams. Each data stream contains one or more sub-computing operations, and these operations are logically independent or have data dependencies at specific points. The data streams can be divided according to the computing graph of the model, data dependencies, computing complexity, and characteristics of hardware resources. For example, the computing tasks of different layers or modules in a deep learning model can be divided into different data streams, and each data stream is responsible for processing a specific part of the data or computing tasks. Each data stream is assigned to a different arithmetic unit for execution. The system performs task allocation and scheduling according to the computing power of each arithmetic unit, the current load situation, and the priority of the data stream. The load situation of each arithmetic unit can be dynamically monitored, and task reallocation can be performed according to the real-time load to ensure full utilization of resources. For example, when an arithmetic unit is idle, the system can automatically assign a new data stream task to it or migrate some tasks from an arithmetic unit with a higher load.
[0052] On the hardware architecture of multiple computing units, multiple data streams can be executed in parallel, making full use of the parallel computing power of the hardware. Each data stream is executed in parallel on its respective computing unit, and data interaction between different data streams can perform synchronous or asynchronous communication at specific points. For example, during the forward propagation of a deep learning model, the computational tasks of different layers can be executed in parallel on different computing units, thereby accelerating the overall computational speed.
[0053] In some cases, multiple data streams need to perform data integration or synchronization at specific points to ensure the correctness and integrity of the calculation results. When a certain data stream requires the calculation results of other data streams, the system will perform data integration at a specific point, using the outputs of other data streams as inputs for continued calculation. For example, during the backpropagation of a deep learning model, the calculation of gradients requires the intermediate results of the forward propagation, and data integration will be performed at an appropriate time point. Synchronization operations will be carried out at key points (such as the splittable computational operations in the embodiments of the present application) to ensure that all data streams have completed the tasks of the current stage before continuing the calculation. This synchronization mechanism can be an explicit waiting operation or an implicit dependency handling. Finally, the calculation results of all data streams will be integrated into the overall data processing flow of the model, driving the model to perform forward calculation or backward propagation. The integrated calculation results are passed as input data to the next layer or the next module of the model to continue subsequent computational tasks. For example, the integrated feature map or hidden state will be passed to the next layer of the model for further calculation.
[0054] During the training process of the model, the calculation results will also be used for backpropagation to calculate gradients and update parameters. This feedback mechanism ensures that the model can be continuously optimized and improved.
[0055] By dividing complex data calculations into multiple data streams and implementing distributed stream computing on the hardware architecture of multiple computing units, while making full use of hardware resources, the operating efficiency of the model can be significantly improved. This approach not only increases the parallelism and resource utilization of the calculation, but also enhances the flexibility and response speed of the system, enabling complex computational tasks to run efficiently on high-performance hardware.
[0056] The following introduces in combination with specific embodiments Figure 1 The specific implementation methods of the various steps shown.
[0057] As an optional embodiment, in S101, identifying the computational operations that meet the splitting conditions in the target model can be implemented as:
[0058] 201. Obtain the calculation types to which the various computational operations in the target model belong, the data interaction relationships between the computational operations, and / or the model structure of the target model;
[0059] 202. Determine the computational operations in the target model that meet the splitting conditions based on the computational type and / or the data interaction relationship.
[0060] In the embodiments of the present application, the splitting conditions are preset based on the model type to which the target model belongs and the model structure.
[0061] In step 201, first, relevant information of the target model needs to be collected and analyzed. Specifically, the computational type (Computation Types), that is, each computational operation (such as convolution, matrix multiplication, activation function, etc.) has its specific computational type. For example, in a deep learning model, convolution operations and matrix multiplications are usually computationally intensive tasks, while activation functions may be relatively simple. By identifying the type of computational operation, it is possible to better judge which operations are suitable for splitting. The data interaction relationship (Data Interaction Relationships) between computational operations, analyze the dependency relationship and data flow between computational operations. For example, the output of some computational operations is the input of other operations, and this dependency relationship determines whether these operations can be split and executed in parallel on different computing units. The data interaction relationship is crucial for ensuring that the sub-computational operations after splitting can be executed independently or executed in parallel to a limited extent. Analyzing the way and frequency of data interaction between computational operations through the data interaction relationship helps to optimize the data flow and reduce unnecessary communication overhead. The model structure (Model Architecture) of the target model, such as convolutional neural network (CNN), recurrent neural network (RNN), Transformer, etc., different model structures have different requirements for the dependency and execution order of computational operations. The model structure helps to judge which parts can be split and parallelized, identify the key nodes and bottlenecks in the model, so as to better split the computational operations. For example, in a CNN, the convolutional layer may contain a large number of parallel computing opportunities, while in an RNN, the dependency between time steps is relatively strong, and the splitting strategy will be different.
[0062] In step 202, after obtaining information such as the computational type and the data interaction relationship, next, based on this information, determine which computational operations meet the splitting conditions. Specifically, according to the computational type and the data interaction relationship, filter out the computational operations suitable for splitting. For example, some computational types (such as matrix multiplication) perform well in parallel computing and have less data interaction, and these operations are more likely to be split. Among the filtered computational operations, determine the specific splitting points. The selection of the splitting points should consider the granularity of the computational operations and the continuity of the data flow to ensure that the sub-computational operations after splitting can still be executed efficiently.
[0063] It can be understood that the following characteristics of the splitting conditions are related to the computing operations. First, the independence of computing operations. If some computing operations are relatively independent, that is, their execution does not strongly depend on the results of other operations, then these operations are ideal candidates for splitting. For example, in the forward propagation of a neural network, the computations of different layers can be carried out relatively independently, so they are suitable for splitting. Second, the computational intensity. Computation-intensive operations (such as matrix multiplication, convolution, etc.) are usually performance bottlenecks. Splitting these operations and executing them on hardware with stronger parallel computing capabilities (such as GPUs, TPUs) can significantly improve performance. Third, data dependence. If the data dependence between computing operations is weak, or the dependence relationship can be resolved through the reconfiguration of the data flow, then these operations can also be split. For example, in some network structures, the operations of non-adjacent layers can be executed in parallel because there is no direct data dependence between them. Fourth, the adaptability of the model type and model structure. According to the type and structure of the target model, certain specific types of models are naturally suitable for operation splitting. For example, in a deep convolutional network, the convolutional layers can be split and processed in parallel on different GPUs, and in a Transformer model, the multi-head computations in the self-attention mechanism can also be split.
[0064] Exemplarily, the splitting conditions are preset based on the model type and model structure to which the target model belongs. For example, complex computing operations (such as large-scale matrix multiplication or convolution operations) are split to utilize the parallel computing capabilities. For example, computing operations with local data interactions are split into multiple sub-operations to reduce the overhead of global data transmission. For example, in computationally intensive modules (such as certain layers in the model), key computing operations are identified and split to improve the overall computing efficiency. For example, computing operations that can be executed in parallel on hardware resources are identified and split to make full use of the advantages of multi-core or multi-device.
[0065] By applying the above steps and splitting conditions, by identifying and splitting complex computing operations and allocating them to multiple arithmetic units for parallel execution, the computing speed and efficiency are significantly improved. According to the computing type and data interaction relationship, the computing tasks are dynamically allocated to the most suitable hardware resources to improve the resource utilization rate of the overall computing system. By optimizing the data interaction relationship, unnecessary communication between computing operations is reduced, and the communication overhead is lowered, further enhancing the system performance. The splitting conditions are preset, flexibly adapting to different types of models and their structures, making the optimization method have wide applicability and scalability.
[0066] In summary, by analyzing and identifying the computing operations of the target model and their relationships in detail and optimizing based on the preset splitting conditions, the execution efficiency and performance of the model on the hardware computing system can be significantly improved.
[0067] Further optionally, in the above steps, determining the computational operations in the target model that meet the splitting conditions based on the computational type and / or the data interaction relationship can be implemented as follows:
[0068] If the target model is a large general model, determine the computational operations in the target model with the computational type of matrix multiplication as the computational operations that meet the splitting conditions. And / or, determine the computational operations in the target model with the data interaction structure obtained by splicing matrix multiplication and activation functions as the computational operations that meet the splitting conditions.
[0069] In the embodiments of the present application, identifying the computational operations in the target model that meet the splitting conditions can be determined by a specific model type (such as a large general model) and a specific data interaction structure. Specifically, if the target model is a large general model (also known as a large pre-trained model, such as BERT, GPT, etc.), then determine the computational operations in the target model with the computational type of matrix multiplication as the computational operations that meet the splitting conditions. Exemplarily, assume that the target model is a BERT model, which contains multiple self-attention mechanism layers. In these layers, matrix multiplication operations account for most of the computational workload. By identifying these operations with the computational type of matrix multiplication, these operations can be further split to improve computational efficiency. Matrix multiplication operations have a high degree of parallelism and can be efficiently executed on multiple GPUs or TPUs, thus significantly improving the computational speed. Distributing matrix multiplication operations to multiple computing units can make full use of hardware resources, reduce the load on a single computing unit, and improve the efficiency of the overall system.
[0070] Determine the computational operations in the target model with the data interaction structure obtained by splicing matrix multiplication and activation functions as the computational operations that meet the splitting conditions.
[0071] Exemplarily, in the self-attention mechanism layer of the BERT model, a matrix multiplication operation is usually followed by an activation function (such as ReLU, GELU, etc.). There is a close data interaction relationship between these computational operations. By identifying this data interaction structure, matrix multiplication and the activation function can be spliced into a sub-computational operation and split and executed on multiple computing units. In this way, combining closely related computational operations (such as matrix multiplication and activation functions) reduces the transmission overhead of intermediate data and improves the efficiency of data processing. By optimizing the data interaction structure, unnecessary memory access and data rearrangement can be reduced, further improving the performance of the computational system.
[0072] Assume that the target model is a BERT model, and assume that the model contains the following sequence of computational operations: Matrix multiplication operation: Calculate attention scores; Activation function operation: Apply the GELU activation function. Through the above splitting conditions, the following computational operations that meet the splitting conditions can be identified and determined: Split and execute the matrix multiplication operation as a separate sub-computational operation in parallel. Or, take the operation composed of the concatenation of the matrix multiplication and the activation function as a composite sub-computational operation and split and execute it in parallel.
[0073] Thus, the matrix multiplication operation is executed in parallel on multiple GPUs, significantly shortening the calculation time. The concatenation operation of the matrix multiplication and the activation function is executed in parallel on multiple computing units, further reducing the calculation latency. By reasonably allocating the computational tasks to multiple computing units, the hardware resources are fully utilized, avoiding resource idleness or overload. Closely related computational operations (such as matrix multiplication and activation function) are executed on the same computing unit, reducing the data transfer overhead and communication latency. By identifying and splitting the key computational operations, the data processing flow is optimized, improving the computational efficiency and performance of the overall system. In summary, through the splitting conditions based on the specific model type and data interaction structure, the system can effectively identify and optimize the computational operations, improving the execution efficiency and performance of the model on the hardware computing system.
[0074] As an optional embodiment, in S102, splitting the computational operation to obtain multiple sub-computational operations can be implemented as the following steps:
[0075] 301. Determine the candidate splitting methods for the computational operation. The candidate splitting methods are determined based on the computational type of the computational operation.
[0076] In this step, for the computational operation to be split, multiple candidate splitting methods are determined. The definition of the candidate splitting methods is usually closely related to the type of the computational operation, such as matrix multiplication, convolution, etc. According to the specific computational operation type (such as matrix multiplication, convolution, activation function, etc.), the corresponding splitting strategy is determined. For matrix multiplication, common splitting methods include row splitting, column splitting, or dividing the entire operation into multiple sub-matrix multiplications. For each computational type, analyze in which dimensions splitting can be performed. For example, in the convolution operation, splitting the input channels, output channels, or spatial dimensions can be considered. List all possible splitting methods. For example, a two-dimensional matrix multiplication can be split into multiple one-dimensional operations, or block partitioning of the input and output can also be considered.
[0077] 302. Based on the data interaction relationship between the computational operation and other computational operations in the target model, obtain the operation efficiency corresponding to the candidate splitting methods.
[0078] In this step, the operation efficiency of each candidate splitting method is analyzed, and its data interaction relationship with other computing operations is evaluated. For each candidate splitting method, performance metrics such as the time required for its execution, the occupancy of memory bandwidth, and the overhead of data transmission are calculated. These efficiency data can be obtained through simulation or benchmark testing. Combining the arrangements of other computing operations in the target model, the impact of the candidate splitting method on the overall computing process is evaluated. If a splitting method can effectively reduce data transmission, increase the computing speed, or better utilize the resources of the computing unit, then its operation efficiency is higher. The operation efficiencies of all candidate splitting methods are compared to find the optimal splitting strategy. Metrics such as time complexity, memory utilization, and overall computing efficiency can be used for quantitative evaluation in this process.
[0079] One optional embodiment is that in step 302, first, the number of other computing operations associated with the computing operation in the target model is obtained. Among them, the other computing operations associated with the computing operation are used to indicate the computing operations that apply some or all of the computing results in the computing operation during the computing process. Furthermore, based on the number of other computing operations associated with the computing operation, the operation efficiency corresponding to the candidate splitting method is determined. Among them, the fewer the number of other computing operations associated with the computing operation, the higher the operation efficiency corresponding to the candidate splitting method.
[0080] Specifically, the fewer the associated computing operations, it means that the output result of this computing operation is used less frequently by subsequent computing operations, so a more efficient splitting method can be selected. The core principle of this method is to reduce data dependence, thereby improving the efficiency of parallel computing.
[0081] Another optional embodiment is that in the above step 302, first, the number of data aggregations to be performed after the computing operation in the target model is obtained. Furthermore, based on the number of data aggregations to be performed after the computing operation, the operation efficiency corresponding to the candidate splitting method is determined. Among them, the fewer the number of data aggregations to be performed after the computing operation, the higher the operation efficiency corresponding to the candidate splitting method.
[0082] In this optional embodiment, step 302 evaluates the operation efficiencies of different candidate splitting methods by analyzing the number of other computing operations associated with the current computing operation. Specifically, the fewer the associated computing operations, it means that the output result of this computing operation is used less frequently by subsequent computing operations, so a more efficient splitting method can be selected. The core principle of this method is to reduce data dependence, thereby improving the efficiency of parallel computing.
[0083] In deep learning models or other complex computational models, there are often dependencies between computational operations. For example, the output result of computational operation A may be used by multiple operations such as computational operations B, C, and D. If the output result of a certain computational operation is only used by a small number of other operations, it means that there is less data interaction, and the degree of parallelism after splitting is higher.
[0084] For a computational operation, there can be multiple splitting methods (such as row splitting, column splitting, etc.). The system evaluates the computational efficiency of each splitting method by analyzing the number of other computational operations associated with this computational operation. If the number of associated computational operations corresponding to a certain splitting method is small, it means that this method can reduce data transmission and dependencies, thereby improving the overall computational efficiency.
[0085] The computational efficiency of the splitting method is inversely proportional to the number of associated computational operations. The fewer the number of other associated computational operations, it means that after splitting, the sub-computational operations can be executed more independently, reducing the overhead of synchronization and data transmission, thereby improving the degree of parallelism and computational efficiency.
[0086] Suppose there is a deep learning model that includes a computational operation of large-scale matrix multiplication (MatMul), and its output result is used for subsequent other computational operations. The output result of the matrix multiplication operation is used by multiple subsequent operations in the model, such as 5 different convolutional layers and 2 activation function layers, a total of 7 associated computational operations. In the first case, candidate splitting method 1 is to split the matrix multiplication by rows, and candidate splitting method 2 is to split the matrix multiplication by columns. Since the result of this matrix multiplication operation needs to be used by 7 other computational operations, it means that no matter which splitting method is selected, there will be more data transmission and synchronization operations. Therefore, the computational efficiency of method 1 and method 2 is relatively low.
[0087] In the second case, there are fewer associated computational operations. The output result of the matrix multiplication operation is only used by 1 subsequent activation function layer. Then, candidate splitting method 1 is to split the matrix multiplication by rows, and candidate splitting method 2 is to split the matrix multiplication by columns. Since the result of the matrix multiplication operation is only used by 1 activation function layer, the data interaction and dependency are low. Column splitting may be more suitable because it can better utilize computational resources, reduce unnecessary data transmission, and thereby improve the computational efficiency.
[0088] For the first case, due to the large number of associated computational operations, choosing the row splitting method may slightly optimize the computational efficiency. For the second case, due to the small number of associated computational operations, the column splitting method has higher computational efficiency because it reduces the complexity of data interaction and can better utilize parallel computational resources.
[0089] In the above embodiments, by selecting a splitting method with a smaller number of associated calculation operations, the data dependency between calculation operations can be effectively reduced, enabling each sub-calculation operation to be executed more independently, thereby improving the parallelism. In addition, the splitting method with fewer associated calculation operations means that the data transmission and synchronization operations required after splitting will also be correspondingly reduced, which will significantly reduce the communication overhead and improve the calculation efficiency. Further, reducing data dependency and data transmission enables each sub-calculation operation to be executed in parallel on multiple arithmetic units more efficiently, thus greatly improving the overall calculation performance. At the same time, selecting an appropriate splitting method can not only improve the operation efficiency but also help to better utilize hardware resources (such as multi-core CPUs, GPUs, TPUs, etc.), avoiding resource idleness or overload. When dealing with large-scale deep learning models or complex calculation tasks, this method can flexibly adapt to different calculation requirements and optimize the overall performance.
[0090] Further optionally, after step 302, the storage space occupancy rate corresponding to the calculation operation under each candidate splitting method can also be obtained. Then, based on the sorting relationship of the storage space occupancy rates, the numerical values of the operation efficiency corresponding to each candidate splitting method are re-determined.
[0091] In this embodiment, after step 302, the influence of the storage space occupancy rate on the operation efficiency of the candidate splitting methods is additionally evaluated. Specifically, the storage space required during the execution of each candidate splitting method is calculated, and by comparing these storage space occupancy rates, the numerical values of the operation efficiency of each splitting method are re-adjusted and determined.
[0092] For each candidate splitting method, calculate the memory space required during the execution. This includes the storage of input data, the storage of intermediate calculation results, and the storage of output results. Compare the storage space required for each splitting method with the overall available storage space, and calculate the storage space occupancy rate for each splitting method.
[0093] Sort the storage space occupancy rates of each splitting method. The splitting method with a lower occupancy rate has a higher priority. Based on the original operation efficiency numerical values, adjust in combination with the storage space occupancy rate. For the splitting method with a lower storage space occupancy rate, its operation efficiency numerical value will be correspondingly increased because the lower storage requirement can reduce the memory bottleneck and improve the overall calculation efficiency.
[0094] Exemplarily, assume there is a deep learning model that includes a computational operation of large-scale matrix multiplication (MatMul). Its candidate splitting methods include row splitting, column splitting, and block splitting. For these three methods, the original operation efficiency evaluation results are as follows: Method 1 (row splitting) has an operation efficiency value of 0.7, Method 2 (column splitting) is 0.8, and Method 3 (block splitting) is 0.9. At the same time, the evaluation results of the storage space occupancy rate show that the storage space occupancy rate of Method 1 is 20%, Method 2 is 30%, and Method 3 is 15%. When re-determining the operation efficiency, according to the sorting of the storage space occupancy rate, Method 3 (15%) has the highest priority, followed by Method 1 (20%) and Method 2 (30%). Assume that for every 10% reduction in the storage space occupancy rate, the operation efficiency value increases by 0.1. The adjusted operation efficiency values are: Method 3 is adjusted from the original 0.9 to 1.1, Method 1 is adjusted from the original 0.7 to 0.8, and Method 2 is adjusted from the original 0.8 to 0.7. In this way, not only is the storage space utilization rate improved, but also the overall computational performance of the matrix multiplication operation is optimized.
[0095] In this way, by considering the storage space occupancy rate, the system can select a splitting method with lower memory requirements, thereby reducing the memory bottleneck and improving the computational efficiency. By combining the storage space occupancy rate and the operation efficiency, the system can better balance the utilization of memory and computational resources and avoid performance degradation caused by insufficient memory. For memory-constrained hardware environments (such as embedded devices, mobile devices, etc.), it is particularly important to select a splitting method with a low storage space occupancy rate. This method can ensure that the model will not be performance-limited due to memory problems when running on these hardware. By optimizing the use of storage space, the system can reduce memory access latency and data transfer overhead, further improving the overall computational performance.
[0096] By further evaluating the storage space occupancy rate after step 302 and making adjustments in combination with the original operation efficiency, this method can more comprehensively evaluate and select the best splitting method. It not only considers the computational efficiency but also takes into account the memory utilization rate, thus achieving better computational performance in different hardware environments. This method is particularly suitable for application scenarios with limited memory resources or strict requirements for memory optimization.
[0097] 303. Adopt the candidate splitting method with the highest operation efficiency to split the said computational operation to obtain multiple sub-computational operations.
[0098] After evaluating all candidate splitting methods, the strategy with the highest operation efficiency will be selected to split the original computational operation to generate multiple sub-computational operations. Based on the evaluation results of step 302, the candidate splitting method with the highest operation efficiency is selected, and the original computational operation is split using this method. For matrix multiplication, it can be divided into multiple smaller matrix multiplication operations. The selected splitting method is used for actual code implementation to generate corresponding sub-computational operations and allocate computing units for these sub-operations. Ensure that each sub-computational operation remains correct and the data flow is smooth after splitting. The generated multiple sub-computational operations are scheduled to available computing units to ensure load balancing, so as to optimize resource utilization and improve the overall computing efficiency.
[0099] In one optional embodiment, in step 303, if the target model is a pan-light large model, the weight matrix is vertically split into multiple first weight matrices. Then, the input matrix is vertically split into multiple first sub-input matrices. Furthermore, the multiplication calculation between multiple first weight matrices and multiple first sub-input matrices is used as the corresponding multiple first sub-computational operations.
[0100] In this way, by splitting a large matrix into multiple small matrices, each sub-computational operation can be executed in parallel on independent computing units (such as multi-core CPUs or GPUs). This can significantly improve the computing efficiency and reduce the overall computing time. When splitting, the sizes of each sub-weight matrix and sub-input matrix are smaller, thus reducing the memory occupation during the computing process and avoiding performance bottlenecks caused by insufficient memory, especially suitable for hardware environments with limited resources. After matrix splitting, the required data transfer volume is reduced, which can reduce data transfer latency and bandwidth occupation and improve the overall system throughput. In this way, it helps to efficiently utilize network resources in a distributed environment. For large neural networks, using this splitting method, the model can be conveniently trained or inferred in a distributed manner, enabling it to be processed in parallel on multiple machines, which not only improves the training speed but also meets the computing requirements of large models. For different tasks or datasets, using the split weight matrices and input matrices, various experiments can be flexibly and quickly implemented without having to reconstruct the entire model, which helps to rapidly iterate research and development.
[0101] By vertically splitting the weight matrix and input matrix in the pan-light large model, the computational complexity can be reduced while improving the operation efficiency and memory utilization. This strategy not only improves the computing performance but also helps to flexibly cope with different application scenarios, greatly enhancing the adaptability and scalability of the model.
[0102] Further optionally, in the above steps, applying the multiple sub-computation results to the data processing flow of the target model can be implemented as: adding the multiple sub-computation results obtained from the multiple first sub-computation operations to obtain the final computation result corresponding to the computation operation.
[0103] In the above steps, applying the multiple sub-computation results to the data processing flow of the target model is specifically implemented as adding the sub-computation results obtained from the multiple first sub-computation operations to obtain the final computation result corresponding to the computation operation.
[0104] For a large-scale matrix multiplication operation Y = W · X, it is split into multiple sub-computation operations Yi = Wi · Xi, where Wi is a block of the weight matrix and Xi is the corresponding block of the input matrix. The results Yi of the multiple sub-computation operations are added together to obtain the final computation result Y. Exemplarily, in a large neural network, assume the dimension of the weight matrix W is 1000×4000 and the dimension of the input matrix X is 4000×500. First, the weight matrix W is vertically split into 4 first weight matrices W1, W2, W3, W4, each with a dimension of 1000×1000. Then, the input matrix X is also vertically split into 4 first sub-input matrices X1, X2, X3, X4, each with a dimension of 1000×500. Then, multiplication operations Y1 = W1·X1, Y2 = W2·X2, Y3 = W3·X3, Y4 = W4·X4 are performed respectively to obtain the sub-computation results Y1, Y2, Y3, Y4. Finally, these sub-computation results are added together to obtain the final computation result Y = Y1 + Y2 + Y3 + Y4, thus completing the entire computation process. This method not only improves the parallel computing ability, but also optimizes the memory usage and data transmission, simplifies the computation process, and enhances the flexibility and scalability of the model.
[0105] In this way, by splitting and parallelly executing sub-computation operations, multiple computing units (such as multi-core CPUs or GPUs) can simultaneously process different sub-computation tasks, greatly improving the computing efficiency and reducing the overall computing time. The matrix sizes involved in each sub-computation operation are smaller, thus reducing the memory occupancy and avoiding memory bottlenecks, especially in an environment with limited memory resources, the effect is particularly significant. The amount of data involved in the split computation operations is less, which can reduce the data transmission latency and bandwidth occupancy, and improve the system throughput, especially in a distributed computing environment, the optimization effect is obvious. The simple addition operation of the sub-computation results avoids complex merging operations, simplifies the computation process, reduces the implementation complexity, and also reduces the computing overhead. By splitting and combining the sub-computation results, the model can be flexibly deployed in different hardware environments, adapt to different computing resource configurations, and improve the scalability and adaptability of the system.
[0106] Assume that during the calculation process, the result of the first sub-calculation operation is Y1 = W1·X1 = [1, 2, 3], the result of the second sub-calculation operation is Y2 = W2·X2 = [4, 5, 6], the result of the third sub-calculation operation is Y3 = W3·X3 = [7, 8, 9], and the result of the fourth sub-calculation operation is Y4 = W4·X4 = [10, 11, 12]. Adding these sub-calculation results together, the final calculation result YY can be obtained: Y = Y1 + Y2 + Y3 + Y4 = [1 + 4 + 7 + 10, 2 + 5 + 8 + 11, 3 + 6 + 9 + 12] = [22, 26, 30].
[0107] In this way, by aggregating the results of multiple sub-calculations, a comprehensive final result is obtained, demonstrating the effectiveness of achieving performance optimization through effective splitting and combining calculations.
[0108] By adding the results of multiple sub-calculations, the results of each sub-calculation operation can be effectively integrated to obtain the final calculation result. This method not only improves the calculation efficiency, optimizes memory usage and data transmission, but also simplifies the calculation process, enhances the flexibility and scalability of the model. In scenarios dealing with large-scale data processing, this method can significantly improve the overall performance and adaptability of the system.
[0109] Another alternative embodiment is that in step 303 above, if the target model is a large-scale model, the weight matrix is horizontally split into multiple second weight matrices. Furthermore, the multiplication calculations between each of the multiple second weight matrices and the input matrix are used as the corresponding multiple second sub-calculation operations.
[0110] For large matrix operations, through horizontal splitting and parallel computing, the parallel computing capabilities of multi-core processors or GPUs can be fully utilized. Each sub-calculation operation Yi can be performed independently, significantly reducing the overall calculation time. After horizontal splitting, the size of each second weight matrix Wi is reduced, reducing memory usage. In the case of limited memory resources, memory overflow can be avoided, and computing resources can be used more efficiently. Horizontally splitting the weight matrix makes the execution of the model more flexible on different hardware platforms. For example, the same model can be executed on GPUs with different video memory sizes without reconstructing the overall structure of the model. By using the entire input matrix X, multiple data transmissions in the calculation can be avoided, improving data utilization rate and reducing unnecessary data copying operations, thereby enhancing the transmission efficiency of the overall calculation.
[0111] This splitting method allows for dynamic adjustment of the results of each sub-calculation. For example, at runtime, the participation of each component can be adjusted according to the calculation amount or memory load, so as to flexibly respond to different calculation requirements.
[0112] Accordingly, in the above steps, the process of applying multiple sub-computation results to the target model can be implemented as follows: concatenate the multiple sub-computation results obtained from multiple second sub-computation operations to obtain the final computation result corresponding to the computation operation.
[0113] In practical applications, for a large input matrix XX (such as image data containing many samples), through the above splitting and computation, the following sub-computation results may be obtained:
[0114] Y1 = [1, 2, 3,..., 1000]
[0115] Y2 = [4, 5, 6,..., 2000]
[0116] Y3 = [7, 8, 9,..., 3000]
[0117] Y4 = [10, 11, 12,..., 4000]
[0118] Finally, the concatenation operation combines these results, and the result is: Y = [1, 2, 3,..., 1000; 4, 5, 6,..., 2000; 7, 8, 9,..., 3000; 10, 11, 12,..., 4000]
[0119] This processing method ensures the efficiency of computation and the optimized use of memory, and also helps to improve the performance of the model in practical applications.
[0120] Another optional embodiment is that in step 303 above, if the target model is a large-scale generalization model, split the input matrix, and use the GeLU operation associated with the split input matrix as the corresponding sub-computation operation.
[0121] In this way, by splitting the input matrix and independently applying the GeLU operation, parallel computing resources such as multi-core CPUs or GPUs can be fully utilized, thereby significantly improving the computation efficiency. Each sub-computation operation can be executed simultaneously, reducing the overall computation time. After splitting the input matrix, the size of each sub-input matrix is reduced, which can effectively reduce memory occupancy. Especially in an environment with limited memory resources, it can avoid memory bottlenecks and make more effective use of available memory resources. For large GeLU operations, splitting them into multiple sub-operations can reduce the computational complexity of each operation. The GeLU function involves non-linear computation. After splitting, the amount of operation for each sub-matrix is reduced, thereby reducing the complexity in the entire computation process.
[0122] During the training process of a neural network, especially in models such as Transformer that use the GeLU activation function, by splitting and parallelizing the GeLU operation, the forward and backward propagation speeds can be significantly accelerated, improving the model's training speed.
[0123] This splitting method allows for dynamically adjusting the size of the sub-matrices according to the data volume in practical applications, thus flexibly coping with input data of different scales. This flexibility is particularly important when dealing with large amounts of data and can improve the adaptability and practicality of the model.
[0124] Exemplarily, assume that in an input matrix containing multiple samples, each sample is represented as a row vector, and the size of the input matrix X is 4000×500. It is split into four blocks:
[0125] X1 = [x1,1 x1,2 … x1,500 x1000,1 x1000,2 … x1000,500]
[0126] X2 = [x1001,1 x1001,2 … x1001,500 x2000,1 x2000,2 … x2000,500]
[0127] X3 = [x2001,1 x2001,2 … x2001,500 x3000,1 x3000,2 … x3000,500]
[0128] X4 = [x3001,1 x3001,2 … x3001,500 x4000,1 x4000,2 … x4000,500]
[0129] Then, apply the GeLU operation to each sub-input matrix. For example, Y1 = GeLU(X1), Y2 = GeLU(X2), Y3 = GeLU(X3), Y4 = GeLU(X4). The final concatenated result is: Y = [Y1; Y2; Y3; Y4].
[0130] This processing method can not only effectively improve the computing efficiency and optimize the memory usage, but also significantly improve the running speed and data processing ability of large neural networks without adding additional hardware resources.
[0131] By executing these three steps, the best splitting method can be determined according to the nature and performance diagnosis of the calculation operation to be split, and the calculation operation can be optimized into multiple sub-calculation operations that are efficiently executed. This not only improves the computing performance but also optimizes the resource usage to meet the requirements of different hardware environments. Such a process ensures that in complex computing tasks, the model can maximize the potential of computing resources and improve the overall execution efficiency.
[0132] S103. During the process of loading the target model into the hardware computing system, reconfigure multiple sub-computation operations into corresponding target computing units.
[0133] In an optional example, before reconfiguring multiple sub-computation operations into corresponding target computing units during the process of loading the target model into the hardware computing system, it is also possible to obtain the first data transfer efficiency between the candidate computing units to be configured for each of the multiple sub-computation operations, and the second data transfer efficiency between the candidate computing units and the surrounding computing units. Furthermore, use the candidate computing unit corresponding to the optimal solution of the first data transfer efficiency and the second data transfer efficiency as the target computing unit corresponding to each of the multiple sub-computation operations.
[0134] In the above optional example, during the process of loading the target model into the hardware computing system, by optimizing the configuration of sub-computation operations on the hardware computing units, the computing efficiency and system performance can be significantly improved. Specifically, before reconfiguring multiple sub-computation operations into the corresponding computing units, the system first obtains the first data transfer efficiency between the candidate computing units of each sub-computation operation, and the second data transfer efficiency between these candidate computing units and other surrounding computing units. Then, by solving the optimal combination of the two, select the most suitable computing unit to execute each sub-computation operation.
[0135] By selecting the computing unit with the highest data transfer efficiency, the data transfer latency between sub-computation operations can be reduced, thereby accelerating the overall computing speed. Especially when processing large-scale neural network models, this optimization effect is particularly significant. By comprehensively considering the first and second data transfer efficiencies, the system can allocate computing resources more evenly, avoiding the situation where some computing units are overloaded while others are idle, thereby improving the overall utilization rate of hardware resources. Efficient data transfer and balanced resource utilization can reduce unnecessary data movement and redundant calculations, thereby reducing the energy consumption of the system, which is particularly important for large-scale computing tasks and high-performance computing environments. By optimizing the selection of computing units, system congestion or failures caused by data transfer bottlenecks or resource contention can be reduced, thereby enhancing the stability and reliability of the system. This method allows the system to dynamically adjust the configuration of computing units according to the current load situation during operation, ensuring the best performance under different computing tasks and load conditions.
[0136] Exemplarily, assume that sub-computation operations Y1, Y2, Y3, Y4 of a neural network model need to be configured into four arithmetic units U1, U2, U3, U4 in a hardware system. The system first measures the first data transfer efficiency between each sub-computation operation and candidate arithmetic units (for example, the efficiency from Y1 to U1 is 90% and to U2 is 85%, etc.), and the second data transfer efficiency between these arithmetic units and other surrounding units (for example, the efficiency between U1 and U2 is 95%, etc.). By comprehensively analyzing these efficiency data, the system selects the optimal combination of arithmetic units. For example, Y1 is configured into U2, Y2 is configured into U3, etc. This optimized configuration not only speeds up data transfer but also balances the computational load, thus significantly improving the overall computational performance.
[0137] In another alternative example, during the process of loading the target model into the hardware computing system, before reconfiguring multiple sub-computation operations into corresponding target arithmetic units, the distance between the candidate arithmetic units to be configured for each of the multiple sub-computation operations can also be obtained based on the data interaction relationship between the computation operations and other computation operations in the target model. Furthermore, the candidate arithmetic units with the distance within a preset range are used as the target arithmetic units corresponding to each of the multiple sub-computation operations.
[0138] In this example, the goal is to optimize the configuration of sub-computation operations on the hardware arithmetic units based on the data interaction relationship between the computation operations during the process of loading the target model into the hardware computing system. Specifically, before selecting the target arithmetic units, the data interaction relationship between each sub-computation operation and other computation operations in the target model is analyzed, and based on this, the distance between the candidate arithmetic units to be configured for each of the multiple sub-computation operations is obtained. Then, the arithmetic units with the distance within the preset range are selected as the target arithmetic units corresponding to each sub-computation operation.
[0139] Specifically, first, the system analyzes the data interaction relationships between each sub-computation operation in the target model and other computation operations within the model. This includes data dependencies, the direction of data flow, and the data transfer requirements between computation operations. Based on the data interaction relationships, the system obtains the physical distance or logical distance between the candidate computing units to be configured for each sub-computation operation. Here, the distance can be the physical distance (such as the location on a hardware chip) or the logical communication delay (such as the transmission time in a network). Further optionally, a preset range is defined, which can be based on empirical values or the optimal performance region obtained through experiments. For example, a suitable distance threshold is set, indicating that the computing units within this distance range are considered to have a higher data transfer efficiency. The candidate computing units whose distances are within the preset range are screened out and used as the target computing units corresponding to each sub-computation operation. For example, for the candidate computing units of the sub-computation operation Yi, the system calculates the distance between the candidate computing units and other computing units. If the distance meets the distance threshold, the candidate computing unit is selected as the target computing unit of Yi.
[0140] In this way, by selecting computing units with shorter distances, the time delay of data transfer can be significantly reduced. Especially in computation operations that require frequent data interaction, this optimization can greatly improve the overall computing efficiency. This approach makes the geographical or logical distribution of computing resources more reasonable, avoiding bottleneck problems that may be caused by long-distance data transfer, thereby improving the overall utilization rate of hardware resources. Reducing the need for long-distance data transfer can reduce the load on the network or hardware transmission link, thereby enhancing the stability and reliability of the system. By reducing unnecessary long-distance data transfer, the energy consumption of the system can be reduced, which is particularly important for large-scale computing tasks and high-performance computing environments. This approach allows the system to dynamically adjust the configuration of computing units according to the current data interaction mode during operation, thus flexibly coping with different computing tasks and load conditions.
[0141] Suppose a neural network model contains four sub-computation operations and needs to be executed on four computing units of a hardware system. After analyzing the data interaction relationships, the system determines the data dependencies between each sub-computation operation and other operations, and then measures the distances between the candidate computing units. Based on this data and the data interaction frequency between the computing units, the target computing unit for the sub-computation operation Y1 is selected. Suppose the distances between this target computing unit and the computing units U2 and U4 of the sub-computation operations Y2 and Y4 are 5 and 10 respectively, and the distance between U2 and U4 is 3, all within the preset range.
[0142] Through this optimized configuration, the data transfer delay can be greatly reduced, the computing efficiency and system performance can be improved, and at the same time, improvements are also achieved in terms of hardware resource utilization and energy consumption efficiency.
[0143] S104. Execute respective corresponding sub-computation operations through the target computing units to obtain multiple sub-computation results.
[0144] In this step, the most suitable target computing units have been assigned to each sub-computation operation. Next, the system will execute the respective sub-computation operations through these computing units to obtain the corresponding sub-computation results. According to the previous optimization configuration, each sub-computation operation is assigned to the corresponding target computing unit. These computing units may be different processor cores, GPU cores, TPUs, or other computing resources in the hardware system. The target computing units start to execute the sub-computation operations assigned to them. These operations can be common deep learning computing tasks such as matrix multiplication, convolution, activation functions (such as ReLU, GeLU), pooling, etc. The computing units execute the corresponding computing tasks according to their specific computing instruction sets and architectures. These tasks are usually highly parallelized and can make full use of the computing power of the hardware.
[0145] After each computing unit completes the calculation, it generates the corresponding sub-computation result. These results may be the outputs of the intermediate layers, feature maps, hidden states, etc. These sub-computation results are usually stored in the form of tensors (Tensor), and the dimensions of the tensors depend on the specific computing operations and the architecture of the model.
[0146] S105. Apply the multiple sub-computation results to the data processing flow of the target model.
[0147] After generating multiple sub-computation results, the system needs to re-integrate these results into the overall data processing flow of the model to drive the model to compute forward or backward propagate. This step ensures that the sub-computation results can correctly participate in the subsequent calculations, thus completing the entire model's computing tasks.
[0148] After obtaining multiple sub-computation results, they are recombined according to the computational graph or data flow of the model. This may involve operations such as concatenating, adding, and element-wise multiplying multiple sub-computation results. According to the architecture of the model and data dependencies, the sub-computation results are integrated into an input format suitable for processing by the next layer or module. The integrated sub-computation results are passed as input data to the next layer or module in the model. These results may serve as the input for the next layer's computation or as part of a feedback mechanism to participate in the backpropagation process during training. In the forward propagation of a neural network, the sub-computation results are passed to subsequent layers for further computation. In backpropagation, these results are used for gradient calculation and parameter update. The application of sub-computation results drives the entire model's computational process forward until the forward or backward computation of the entire model is completed. According to the computational graph of the model, the integrated results are sequentially passed to subsequent arithmetic units to ensure that each computational step can be correctly executed, thereby generating the final model output or completing the backpropagation during training.
[0149] By performing steps S104 and S105, the system can efficiently allocate computational tasks to different arithmetic units and ensure that the computational results can be correctly integrated into the overall data processing flow of the model. This approach improves computational efficiency, fully utilizes hardware resources, and ensures the accuracy and integrity of model computations.
[0150] In the embodiments of this application, in the Bloom large model, there are many structures where a matrix multiplication is followed by an activation function. Suppose there is an input matrix X, an output matrix Y, and a weight matrix A. Then, the structure of matrix multiplication followed by an activation function in the Bloom large model is mainly split for calculation to improve the model's running efficiency. Specifically, when there is a splittable matrix multiplication in the model, the weight matrix can be split by columns into multiple arithmetic units. Each arithmetic unit independently performs part of the matrix multiplication, and the obtained results can independently enter the activation function to participate in subsequent calculations. The advantage of this splitting method is that it allows parallel processing, thus significantly improving the computational speed.
[0151] When choosing a splitting method, the impact of subsequent computational operations needs to be considered to select a method that is more convenient for subsequent calculations and generates fewer data aggregation operations. For example, splitting left and right usually has higher computational efficiency because it only requires splitting the weight matrix, and the result matrix can be directly concatenated to obtain the complete result. In contrast, although splitting up and down requires less computational space, it may require additional data aggregation operations, which will increase the computational overhead.
[0152] In addition, it can also be applied to other divisible computational operations in the Pantheon large model, even without matrix multiplication. For example, for certain calculations that can be performed without complete data (such as the GeLU activation function), the input matrix can also be split to enable multiple computing units to execute in parallel. In this case, the choice of the splitting method should be mainly based on the principle of facilitating subsequent calculations and reducing the number of aggregations.
[0153] Generally speaking, by reasonably splitting the computational operations in the model, efficient parallel computing on multiple computing units is achieved, improving the overall operating efficiency of the model, while having good generality and flexibility.
[0154] Exemplarily, assume that the divisible part in the Pantheon large model is matrix multiplication. Figure 2 Two splitting methods of the weight matrix are shown. If the weight matrix A is split by column into n computing units, and then XA1, XA2, …, XAn are executed, finally n calculation results Y1, Y2, …, Yn will be obtained. These results can independently enter the activation function for calculation. The choice of the splitting method is mainly affected by subsequent computational operations. The selected splitting method should be more convenient for subsequent calculations and generate fewer data aggregation operations in subsequent calculations.
[0155] Figure 3 in represents splitting matrix A vertically, represents splitting matrix B vertically. In the forward propagation, f means the data passing through here remains unchanged, and g means the data of each computing unit needs to be combined into complete data here. In the backward propagation, the meanings are opposite.
[0156] In the multi-layer perceptron module, first, the input matrix X and matrix A need to perform matrix multiplication. Here, matrix A is split horizontally. If it is split vertically, the resulting data cannot continue with the subsequent GeLU and needs to be aggregated immediately, which will obviously consume more resources than horizontal splitting. After horizontal splitting and calculation through GeLU, Y is obtained. Here, Y is the result at different positions. Then, it needs to perform matrix multiplication with matrix B. Here, matrix B is split vertically, and it can only be split vertically because the input matrix has already been split horizontally at this time. After that, the subsequent dropout calculation requires complete data, so aggregation needs to be performed first. It is not easy to clarify the choice of the splitting method and its impact on the subsequent operations. It can be understood by referring to the description of the two splitting methods in the first figure.
[0157] It can be stated that the calculation process of the self-attention module is similar. The self-attention module only requires two splits and one aggregation data processing flow, which will not be elaborated here.
[0158] The main modules in the Fanguang large model are only the multi-layer perceptron module and the self-attention mechanism module. The remaining parts can also be split, but the improvement in efficiency is relatively small. For the self-attention module and multi-layer perceptron module in other neural network models, if the structure remains unchanged, this splitting method can be directly adopted. If the structure changes, certain modifications may be required, which should be specifically set according to the actual application scenario.
[0159] It can be seen in Figure 2 that the final result is still obtained after the left-right splitting calculation. Or rather, each computing unit obtains the final result at different positions. As long as these results are concatenated or combined together, the same result as that without splitting can be obtained. If it is split up and down, partial results with the same calculation size as that without splitting can be obtained, and these results need to be added together to obtain the final result. Most calculations can continue with the results at different positions, while some partial results cannot continue most of the operations. The inability to continue means that aggregation is required, that is, adding the results in each computing unit and then synchronizing them to each computing unit. Aggregation will bring a lot of overhead. For higher model running efficiency, the choice of splitting method should generate as few aggregation operations as possible, which needs to be specifically selected according to the model structure and the required calculations.
[0160] In most cases of matrix multiplication, the left-right splitting method has higher computing efficiency in terms of weight splitting, while the up-down splitting method requires less computing space. The specific reasons can be roughly seen from Figure 2 that for left-right splitting, only the weights need to be split, and after obtaining the results, only concatenation is required to get the complete result. However, each computing unit needs to store 14 values (X, A1, Y1) simultaneously. For up-down splitting, not only the weights need to be split but also the input matrix needs to be split, and the result matrix needs to be added together to obtain the final result. But this splitting method only needs to store 12 values (X1, A1, Y1) simultaneously.
[0161] If a certain module does not have matrix multiplication, but there are some calculations that can be performed without all the data, splitting can also be attempted. For example, in the second figure, if there is no matrix multiplication of XA and the input directly participates in the GeLU calculation, since this calculation does not require complete data, the input matrix can also be split to allow multiple computing units to run the GeLU calculation simultaneously. Here, the splitting method does not affect the computing efficiency of GeLU. As for which method to choose, it is still selected according to facilitating subsequent calculations and reducing the number of aggregations. The partial results obtained are equivalent to splitting the complete result column by column. In the subsequent matrix multiplication, the weight matrix can be split row by row, so that the partial results obtained previously can be directly used to participate in the calculation without additional communication overhead.
[0162] According to the technical solution given in the above steps, a similar computing structure can be split into multiple computing units, and then the data is aggregated only when synchronization is required. See Figure 3 and Figure 4 As shown, the data flow diagrams of the Self-Attention module and the Multi-Layer Perceptron (MLP) module in the Bloom large model after applying this method. Figure 3 and Figure 4 In, f is the identity operator for forward propagation, indicating that the data remains unchanged after passing through here during forward propagation. At the same time, it is also the aggregation operator for backpropagation, that is, the data needs to be synchronized here, while g is the opposite of f. Since the Self-Attention module in the Bloom large model adopts the Multi-Head Attention mechanism, the calculation of Multi-Head Attention itself is parallel because each head is independent. Therefore, the calculations in the Self-Attention module are very suitable for split calculations.
[0163] In the embodiments of the present application, it is possible to identify the computing operations in the target model that meet the splitting conditions; split the computing operations to obtain multiple sub-computing operations; during the process of loading the target model into the hardware computing system, reconfigure the multiple sub-computing operations into the corresponding target computing units; the hardware computing system includes at least multiple computing units, and each computing unit is composed of the hardware chip resources called by assembly instructions; execute the respective corresponding sub-computing operations through the target computing units to obtain multiple sub-computation results; and apply the multiple sub-computation results to the data processing flow of the target model. The embodiments of the present application can improve the computing efficiency, hardware utilization rate, flexibility, and real-time performance, reduce the latency of computing operations, perform hardware acceleration operations for the model, and improve the performance of the neural network model by identifying, splitting, reconfiguring, and parallelly executing the computing operations in the neural network model.
[0164] In particular, the embodiments of the present application can split the complex mathematical calculations in the neural network model onto different computing units, reduce the waiting time required between computing operations, and this method can improve the utilization rate of hardware resources per unit time, give full play to the hardware performance, help avoid waste caused by resource idleness, and enable the neural network model (such as the Bloom large model, etc.) to be more efficiently implemented and deployed on the hardware with multiple computing units.
[0165] Based on the same implementation principle, the embodiments of the present application also provide a model optimization device for implementing the model optimization method in the above embodiments.
[0166] Figure 5The block diagram of the model optimization device provided by the embodiment of the present application. This device is applied to a hardware computing system where a target model is deployed, and the target model belongs to a neural network model. As Figure 5 shown, the device at least includes the following units:
[0167] An identification unit, configured to identify the computing operations in the target model that meet the splitting conditions;
[0168] A splitting unit, configured to split the computing operations to obtain a plurality of sub-computing operations;
[0169] A configuration unit, configured to reconfigure the plurality of sub-computing operations into corresponding target operation units during the process of loading the target model into the hardware computing system; the hardware computing system at least includes a plurality of operation units, and each operation unit is composed of hardware chip resources called based on assembly instructions;
[0170] An execution unit, configured to execute their respective corresponding sub-computing operations through the target operation units to obtain a plurality of sub-computing results;
[0171] An application unit, configured to apply the plurality of sub-computing results to the data processing flow of the target model.
[0172] In an optional embodiment, when the identification unit identifies the computing operations in the target model that meet the splitting conditions, it is configured to:
[0173] Obtain the computing types to which the respective computing operations in the target model belong, the data interaction relationships between the computing operations, and / or the model structure of the target model;
[0174] Based on the computing type and / or the data interaction relationship, determine the computing operations in the target model that meet the splitting conditions; the splitting conditions are preset based on the model type to which the target model belongs and the model structure.
[0175] In an optional embodiment, when the identification unit determines the computing operations in the target model that meet the splitting conditions based on the computing type and / or the data interaction relationship, it is configured to:
[0176] If the target model is a large-scale model, determine the computing operations in the target model whose computing type is matrix multiplication as the computing operations that meet the splitting conditions; and / or,
[0177] Determine the computing operations in the target model whose data interaction structure is obtained by splicing matrix multiplication and an activation function as the computing operations that meet the splitting conditions.
[0178] In an alternative embodiment, when the splitting unit splits the computing operation to obtain a plurality of sub-computing operations, it is configured to:
[0179] Determine candidate splitting methods for the computing operation; the candidate splitting methods are determined based on the computing type of the computing operation;
[0180] Based on the data interaction relationship between the computing operation and other computing operations in the target model, obtain the operation efficiency corresponding to the candidate splitting method;
[0181] Use the candidate splitting method with the highest operation efficiency to split the computing operation to obtain a plurality of sub-computing operations.
[0182] In an alternative embodiment, when the splitting unit obtains the operation efficiency corresponding to the candidate splitting method based on the data interaction relationship between the computing operation and other computing operations in the target model, it is configured to:
[0183] Obtain the number of other computing operations associated with the computing operation in the target model; where the other computing operations associated with the computing operation are used to indicate the computing operations that apply some or all of the computing results in the computing operation during the computing process;
[0184] Determine the operation efficiency corresponding to the candidate splitting method based on the number of other computing operations associated with the computing operation; where the fewer the number of other computing operations associated with the computing operation, the higher the operation efficiency corresponding to the candidate splitting method.
[0185] In an alternative embodiment, when the splitting unit obtains the operation efficiency corresponding to the candidate splitting method based on the data interaction relationship between the computing operation and other computing operations in the target model, it is configured to:
[0186] Obtain the number of data aggregation times to be performed after the computing operation in the target model;
[0187] Determine the operation efficiency corresponding to the candidate splitting method based on the number of data aggregation times to be performed after the computing operation; where the fewer the number of data aggregation times to be performed after the computing operation, the higher the operation efficiency corresponding to the candidate splitting method.
[0188] In an alternative embodiment, when the splitting unit uses the candidate splitting method with the highest operation efficiency to split the computing operation to obtain a plurality of sub-computing operations, it is configured to:
[0189] If the target model is a large-scale model, split the weight matrix vertically into a plurality of first weight matrices;
[0190] Vertically split the input matrix into multiple first sub-input matrices;
[0191] Use the multiplication calculation between multiple first weight matrices and multiple first sub-input matrices as the corresponding multiple first sub-calculation operations.
[0192] In an alternative embodiment, when applying the multiple sub-calculation results to the data processing flow of the target model, the application unit is configured to:
[0193] Add the multiple sub-calculation results obtained from the multiple first sub-calculation operations to obtain the final calculation result corresponding to the calculation operation.
[0194] In an alternative embodiment, when the splitting unit splits the calculation operation into multiple sub-calculation operations by using the candidate splitting method with the highest operation efficiency, it is configured to:
[0195] If the target model is a large-scale model, horizontally split the weight matrix into multiple second weight matrices;
[0196] Use the multiplication calculation between each of the multiple second weight matrices and the input matrix as the corresponding multiple second sub-calculation operations.
[0197] In an alternative embodiment, when applying the multiple sub-calculation results to the data processing flow of the target model, the application unit is configured to:
[0198] Concatenate the multiple sub-calculation results obtained from the multiple second sub-calculation operations to obtain the final calculation result corresponding to the calculation operation.
[0199] In an alternative embodiment, when the splitting unit splits the calculation operation into multiple sub-calculation operations by using the candidate splitting method with the highest operation efficiency, it is configured to:
[0200] If the target model is a large-scale model, split the input matrix and use the GeLU operation associated with the split input matrix as the corresponding sub-calculation operation.
[0201] In an alternative embodiment, after obtaining the operation efficiency corresponding to the candidate splitting method based on the data interaction relationship between the calculation operation and other calculation operations in the target model, the splitting unit is further configured to:
[0202] Obtain the storage space occupancy rate corresponding to the calculation operation under each candidate splitting method;
[0203] Based on the sorting relationship of the storage space occupancy rates, re-determine the numerical values of the operation efficiencies corresponding to each candidate splitting method.
[0204] In an alternative embodiment, an optimization unit is further included and is further configured to: before the configuration unit reconfigures a plurality of sub-computation operations into corresponding target computing units during the process of loading the target model into the hardware computing system, obtain the first data transfer efficiency between the candidate computing units to be configured for each of the plurality of sub-computation operations and the second data transfer efficiency between the candidate computing units and the surrounding computing units;
[0205] Use the candidate computing unit corresponding to the optimal solution of the first data transfer efficiency and the second data transfer efficiency as the target computing unit corresponding to each of the plurality of sub-computation operations.
[0206] In an alternative embodiment, an optimization unit is further included and is further configured to: before the configuration unit reconfigures a plurality of sub-computation operations into corresponding target computing units during the process of loading the target model into the hardware computing system, obtain the distance between the candidate computing units to be configured for each of the plurality of sub-computation operations based on the data interaction relationship between the computation operations and other computation operations in the target model;
[0207] Use the candidate computing unit whose distance is within a preset range as the target computing unit corresponding to each of the plurality of sub-computation operations.
[0208] It should be noted that for the model optimization device provided in the embodiments of the present application, the specific functions and implementation details of each module have the same implementation principle as the implementation process of the corresponding steps in the foregoing method embodiments. Specifically, reference can be made to the description of the corresponding parts in the foregoing method embodiments, which will not be elaborated herein.
[0209] Based on the same implementation principle, an embodiment of the present application further provides an electronic device, Figure 6 which is a structural block diagram of the electronic device 400, as shown in Figure 6 The electronic device 400 includes: a processor 401, a memory 402, a communication interface 403, a communication bus 404, and a controller 405; wherein, the processor 401, the memory 402, and the communication interface 403 communicate with each other through the communication bus 404; the memory 402 is used to store a computer program; the processor 401 is used to execute the program stored in the memory 402 to implement the corresponding processing function; the communication interface 404 is used for communication between the electronic device 400 and other devices; the controller 405 is used to implement the model optimization method described in the foregoing method embodiments.
[0210] In the embodiments of the present application, the communication bus 404 may be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus 404 may be divided into an address bus, a data bus, a control bus, etc., and the specific form is not limited. For the sake of convenience of representation, Figure 6 only a thick line is used to represent it in Figure 6 , but it does not mean that there is only one bus or one type of bus.
[0211] The memory 402 may include a Random Access Memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor 401.
[0212] The processor 401 may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and the specific form may be determined according to actual requirements.
[0213] Correspondingly, the embodiments of the present application also provide a computer-readable storage medium storing a computer program, and when the computer program is executed, it can implement the steps executable by the electronic device in the above method embodiments.
[0214] It should be noted that although one or more embodiments of the present application provide the method operation steps as described in the embodiments or flowcharts, based on routine or non-creative labor, there may be more or fewer operation steps. The step order listed in the embodiments is only one way among the execution orders of numerous steps, and does not represent the only execution order. When the actual device or client product is executed, it may be executed in the order of the embodiments or the method shown in the drawings, or executed in parallel (such as in an environment of parallel processors or multi-threaded processing).
[0215] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, apparatuses (systems), or computer program products. Therefore, the embodiments of the present application can take the form of all-hardware embodiments, all-software embodiments, or embodiments combining software and hardware aspects. Moreover, one or more embodiments of the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0216] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of one or more embodiments of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of one or more embodiments of the present application, and they should all be covered by the scope of the claims and the description of one or more embodiments of the present application.
[0217] The above has described one or more embodiments of the present application in combination with optional embodiments, but these embodiments are only exemplary and only play an illustrative role. On this basis, various replacements and improvements can be made to one or more embodiments of the present application, and these all fall within the protection scope of one or more embodiments of the present application.
Claims
1. A model optimization method, characterized in that: Applied to a hardware computing system deployed with a target model, the target model being a neural network model; the method comprises: Identifying computing operations in the target model that meet splitting conditions; the splitting conditions are pre-set based on the model type and model structure to which the target model belongs; Splitting the computing operation to obtain multiple sub-computing operations; In the process of loading the target model into the hardware computing system, multiple sub-computing operations are reconfigured to corresponding target computing units; the hardware computing system includes at least multiple computing units, each of which is composed of hardware chip resources called based on assembly instructions; Execute the corresponding sub-computation operations through the target computing unit to obtain multiple sub-computation results; Applying the multiple sub-calculation results to the data processing flow of the target model; Wherein, in the process of loading the target model into the hardware computing system, before reconfiguring the plurality of sub-computing operations to the corresponding target computing units, the method further includes: Acquire a first data transmission efficiency between candidate computing units to be configured for each of the plurality of sub-computing operations, and a second data transmission efficiency between the candidate computing units and surrounding computing units; taking the candidate computing unit corresponding to the optimal solution of the first data transmission efficiency and the second data transmission efficiency as the target computing unit corresponding to each of the plurality of sub-computing operations; Wherein, in the process of loading the target model into the hardware computing system, before reconfiguring the plurality of sub-computing operations to the corresponding target computing units, the method further includes: Based on the data interaction relationship between the computing operation and other computing operations in the target model, obtaining distances between candidate computing units to be configured for each of the plurality of sub-computing operations; The candidate computing units whose distances are within the preset range are used as target computing units corresponding to the multiple sub-computing operations.
2. The method according to claim 1, characterized in that The identifying the computing operations in the target model that meet the splitting conditions includes: Acquire the calculation type to which each calculation operation in the target model belongs, the data interaction relationship between the calculation operations, and / or the model structure of the target model; Based on the calculation type and / or the data interaction relationship, the calculation operations in the target model that meet the splitting condition are determined; the splitting condition is pre-set based on the model type to which the target model belongs and the model structure.
3. The method according to claim 2, characterized in that The determining, based on the calculation type and / or the data interaction relationship, the calculation operation in the target model that meets the splitting condition includes: If the target model is a large floodlight model, determining a calculation operation whose calculation type in the target model is matrix multiplication as a calculation operation that meets the splitting condition; and / or, Determine that the data interaction structure in the target model is a calculation operation obtained by concatenating matrix multiplication and activation function as a calculation operation that meets the splitting condition.
4. The method according to claim 1, characterized in that: The splitting of the computing operation to obtain multiple sub-computing operations includes: Determine a candidate splitting method for the computing operation; the candidate splitting method is determined based on a computing type of the computing operation; Based on the data interaction relationship between the computing operation and other computing operations in the target model, obtaining the computing efficiency corresponding to the candidate splitting method; The computing operation is split using a candidate splitting method with the highest computational efficiency to obtain a plurality of sub-computing operations.
5. The method according to claim 4, characterized in that The obtaining, based on the data interaction relationship between the computing operation and other computing operations in the target model, the computing efficiency corresponding to the candidate splitting mode includes: Acquire the number of other computing operations associated with the computing operation in the target model; wherein the other computing operations associated with the computing operation are used to indicate computing operations applied to part or all of the computing results in the computing operation during the computing process; The computing efficiency corresponding to the candidate splitting method is determined by the number of other computing operations associated with the computing operation; wherein, the smaller the number of other computing operations associated with the computing operation, the higher the computing efficiency corresponding to the candidate splitting method.
6. The method according to claim 4, characterized in that The obtaining, based on the data interaction relationship between the computing operation and other computing operations in the target model, the computing efficiency corresponding to the candidate splitting mode includes: Obtaining the number of data aggregations to be performed after the calculation operation in the target model; The computing efficiency corresponding to the candidate splitting method is determined based on the number of data aggregation times to be performed after the computing operation; wherein, the smaller the number of data aggregation times to be performed after the computing operation, the higher the computing efficiency corresponding to the candidate splitting method.
7. The method according to claim 4, characterized in that The candidate splitting method with the highest computational efficiency is adopted to split the computing operation to obtain multiple sub-computing operations, including: If the target model is a large floodlight model, the weight matrix is vertically split into a plurality of first weight matrices; Splitting the input matrix vertically into a plurality of first sub-input matrices; The multiplication calculations between the plurality of first weight matrices and the plurality of first sub-input matrices are used as the corresponding plurality of first sub-calculation operations.
8. The method according to claim 7, characterized in that The step of applying the plurality of sub-calculation results to the data processing flow of the target model includes: A plurality of sub-calculation results obtained by the plurality of first sub-calculation operations are added together to obtain a final calculation result corresponding to the calculation operation.
9. The method according to claim 4, characterized in that The candidate splitting method with the highest computational efficiency is adopted to split the computing operation to obtain multiple sub-computing operations, including: If the target model is a large floodlight model, the weight matrix is split horizontally into a plurality of second weight matrices; The multiplication calculation between each of the plurality of second weight matrices and the input matrix is performed as the corresponding plurality of second sub-calculation operations.
10. The method according to claim 9, characterized in that The step of applying the plurality of sub-calculation results to the data processing flow of the target model includes: The multiple sub-computation results obtained by the multiple second sub-computation operations are concatenated to obtain a final calculation result corresponding to the calculation operation.
11. The method according to claim 4, characterized in that The candidate splitting method with the highest computational efficiency is adopted to split the computing operation to obtain multiple sub-computing operations, including: If the target model is a large floodlight model, the input matrix is split, and the GeLU operations associated with the split input matrix are used as corresponding sub-computation operations.
12. The method according to claim 4, characterized in that After obtaining the computational efficiency corresponding to the candidate splitting method based on the data interaction relationship between the computational operation and other computational operations in the target model, the method further includes: Obtaining the storage space occupancy rate corresponding to the calculation operation under each candidate splitting mode; Based on the ranking relationship of the storage space occupancy rate, the numerical value of the computational efficiency corresponding to each candidate splitting method is re-determined.
13. A model optimization device, characterized in that: The device is applied to a hardware computing system deployed with a target model, wherein the target model is a neural network model; the device comprises at least the following units: an identification unit configured to identify computing operations in the target model that meet a splitting condition; the splitting condition is pre-set based on a model type and a model structure to which the target model belongs; a splitting unit, configured to split the computing operation into a plurality of sub-computing operations; a configuration unit configured to reconfigure the plurality of sub-computing operations into corresponding target computing units during the process of loading the target model into the hardware computing system; The hardware computing system comprises at least a plurality of computing units, each of which is composed of hardware chip resources called by assembly instructions; An execution unit is configured to execute the corresponding sub-computation operations through the target operation unit to obtain multiple sub-computation results; an application unit, configured to apply the plurality of sub-calculation results to a data processing flow of the target model; The optimization unit is further configured to: obtain, during the process of loading the target model into the hardware computing system, a first data transmission efficiency between candidate computing units to be configured for each of the multiple sub-computing operations, and a second data transmission efficiency between the candidate computing units and surrounding computing units, before the configuration unit reconfigures the multiple sub-computing operations into corresponding target computing units; and use the candidate computing unit corresponding to the optimal solution of the first data transmission efficiency and the second data transmission efficiency as the target computing unit corresponding to each of the multiple sub-computing operations; In which, the optimization unit is also configured to: during the process of loading the target model into the hardware computing system, before the configuration unit reconfigures multiple sub-computing operations to corresponding target computing units, obtain the distances between the candidate computing units to be configured for each of the multiple sub-computing operations based on the data interaction relationship between the computing operations and other computing operations in the target model; and use the candidate computing units whose distances are within a preset range as the target computing units corresponding to each of the multiple sub-computing operations.
14. A chip, characterized in that: The chip includes a processor coupled to a transceiver, and is used to execute the model optimization method according to any one of claims 1 to 12.
15. An electronic device, characterized in that: The system comprises a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the model optimization method according to any one of claims 1 to 12.
16. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program instructions, which, when executed, implement the model optimization method described in any one of claims 1 to 12.
Citation Information
Patent Citations
Model scheduling method and system, computing equipment and storage medium
CN118468926A