A model optimization method and related device
By pre-calculating and replacing the constant computing unit of the large language model in the chip hardware simulation environment, the redundant computing problem caused by constant computing is solved, and the execution performance of the model is significantly improved.
Patent Information
- Application Number
- CN202411549279.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-01
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-11-01
AI Technical Summary
In large language models, constant calculation leads to redundant calculations, increases computational burden, and reduces model operation efficiency, especially in high-parallel computing scenarios.
By obtaining the target model to be compiled in the chip hardware simulation environment, traversing its calculation units, filtering out the target calculation units that do not rely on the dynamic data of the runtime, and pre-calculating them during the compilation stage to replace the constant calculation operations in the model.
It reduces the computational resource calls during the model runtime, reduces unnecessary computational overhead, improves the utilization of computing resources, and improves the overall execution performance of the model.
Smart Images

Figure CN119088401B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of machine learning, and in particular to a model optimization method and related devices. Background Art
[0002] In the field of artificial intelligence, especially in machine learning and deep learning applications, algorithm efficiency and computing performance are crucial. For example, in the Bloom model of the Large Language Model (LLM), the reasoning or training process requires a large amount of computing, so the computing efficiency and the utilization of computing resources will directly affect the model operation efficiency.
[0003] In related technologies, the model structure and its calculation formula are usually optimized. However, if the model contains constant calculations, this optimization method will generate a large number of redundant calculations, resulting in a significant increase in the amount of calculations and reduced model operation efficiency. In particular, since constant calculations occupy storage resources and computing units, the utilization rate of various resources is high in computing scenarios with high parallelism, and the occupation of redundant resources will also cause other computing operations to be delayed, further reducing the model operation efficiency.
[0004] Therefore, it is urgent to design a new technical solution to overcome at least one of the above technical problems. Summary of the invention
[0005] The present application provides a model optimization method and related devices for pre-completing target computing operations in a chip hardware simulation environment, reducing the computing resources required to call when the model is running, effectively reducing unnecessary computing overhead, improving computing resource utilization, and improving the overall execution performance of the model.
[0006] In a first aspect, the present application provides a model optimization method, comprising:
[0007] Obtain the target model to be compiled from the chip hardware simulation environment;
[0008] Traversing each computing unit in the target model, and screening out a target computing unit containing a target computing operation; the target computing operation does not depend on dynamic data at runtime, and the execution logic of the target computing operation is not associated with the runtime operation;
[0009] Pre-calculating the target computing unit to obtain a pre-calculation result of the target computing unit;
[0010] The pre-calculation result is used to replace the constant calculation operation contained in the target calculation unit to obtain an optimized target model.
[0011] In a second aspect, an embodiment of the present application provides a model optimization device, which includes at least the following units:
[0012] An acquisition unit is configured to acquire a target model to be compiled from a chip hardware simulation environment;
[0013] A screening unit is configured to traverse each computing unit in the target model and screen out a target computing unit containing a target computing operation; the target computing operation does not depend on dynamic data at runtime, and the execution logic of the target computing operation is not associated with the runtime operation;
[0014] a pre-calculation unit, configured to pre-calculate the target calculation unit to obtain a pre-calculation result of the target calculation unit;
[0015] The replacement unit is configured to use the pre-calculation result to replace the constant calculation operation contained in the target calculation unit to obtain an optimized target model.
[0016] In a third aspect, an embodiment of the present application provides a chip for implementing the model optimization method described in the first aspect.
[0017] In a fourth aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the model optimization method described in the first aspect.
[0018] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed, the model optimization method described in the first aspect is implemented.
[0019] In an embodiment of the present application, the target model to be compiled is obtained from the chip hardware simulation environment; each computing unit in the target model is traversed to screen out the target computing unit containing the target computing operation; the target computing operation does not depend on the dynamic data at runtime, and the execution logic of the target computing operation is not associated with the runtime operation; the target computing unit is pre-calculated to obtain the pre-calculation result of the target computing unit; the pre-calculation result is used to replace the constant computing operation contained in the target computing unit to obtain the optimized target model. In an embodiment of the present application, firstly, by screening out the target computing units that do not depend on dynamic data in the model, and pre-calculating these constant calculations in the compilation stage, it is avoided to repeat these calculations each time the model is run, thereby reducing a large number of unnecessary computing operations. Secondly, the pre-calculation result can replace the constant computing operation in the target computing unit, reducing the workload of real-time computing, and thus reducing the computing burden when the model is running. This is particularly critical for scenarios that require processing a large number of parallel calculations, which helps to significantly improve the overall computing efficiency. Thirdly, the storage and computing resources occupied by the constant computing operation are optimized and released, reducing the invalid occupation of computing units and storage space, so that other key computing tasks can be executed more quickly, thereby improving the resource utilization of the system. In addition, by removing unnecessary constant calculations, resource competition in high-parallel computing scenarios is reduced, the probability of delayed execution of tasks is reduced, and parallel operations can be completed more efficiently, thereby improving the execution efficiency of the model in parallel computing. The optimized model structure is more concise, reducing the complexity caused by constant calculations, and enhancing the scalability and stability of the model in large data sets and complex scenarios. In short, the target calculation operation can be completed in advance through the embodiment of the present application, thereby reducing the computing resources required to call when the model is running, effectively reducing unnecessary computing overhead, and improving the utilization of computing resources. , And then improve the overall execution performance of the model, especially in the reasoning or training scenarios involving large-scale models. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0021] Figure 1 is a flow chart of a model optimization method according to an embodiment of the present application;
[0022] Figure 2 A schematic diagram of the principle of a model optimization method in an embodiment of the present application;
[0023] Figure 3 is a structural block diagram of a model optimization device according to an embodiment of the present application;
[0024] Figure 4 It is a structural block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0025] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0026] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application.
[0027] In the field of artificial intelligence, especially in machine learning and deep learning applications, algorithm efficiency and computing performance are crucial. For example, in large language models (such as the Bloom model), both the reasoning and training processes require a lot of computing resources, so the computing efficiency and the utilization of computing resources will directly affect the overall performance of the model. In order to improve computing efficiency, the model structure and calculation formula are often optimized. However, when a large number of constant calculations are involved in the model, this optimization method may cause redundant calculations, significantly increase the computing burden, and thus reduce the operating efficiency of the model.
[0028] Specifically, constant calculations not only occupy valuable storage resources and computing units, but also affect the rational use of resources in a highly parallel computing environment. Due to the redundant occupation of constant calculations, other computing tasks may be delayed, further reducing the overall computing efficiency. This is particularly evident in parallel computing scenarios with high resource utilization. Therefore, it is urgent to design a new technical solution to effectively solve the problems of resource waste and performance degradation caused by constant calculations, thereby improving the operating efficiency of large models. Therefore, it is urgent to design a new technical solution to overcome at least one of the above technical problems.
[0029] To this end, the embodiment of the present application provides a model optimization method and related devices, by obtaining the target model to be compiled from the chip hardware simulation environment, traversing each computing unit in the target model, screening out the target computing unit that does not depend on the runtime dynamic data and whose execution logic is not associated with the runtime operation, and pre-calculating these constant calculations in the compilation stage, obtaining the pre-calculation results, which are used to replace the constant calculation operations in the target computing unit, thereby obtaining the optimized target model. First, by pre-calculating the constant operation in the compilation stage, it is avoided to repeat these calculations each time the model is run, and a large number of unnecessary calculation operations are reduced; secondly, the pre-calculation results replace the real-time calculations, which reduces the calculation burden of the model when it is running, especially for the scene that needs to process a large number of parallel calculations, and can significantly improve the overall calculation efficiency. In addition, the storage and computing resources occupied by the constant calculation operation are optimized and released, reducing the invalid occupation of the computing unit and storage space, so that other key computing tasks can be executed more quickly, and the resource utilization of the system is improved; at the same time, by removing unnecessary constant calculations, the resource competition in the high parallel computing scene is reduced, the delay execution probability of the task is reduced, and the parallel operation is more efficient, which further improves the execution efficiency of the model in parallel computing. Finally, the optimized model structure is more concise, reducing the complexity caused by constant calculations, and enhancing the scalability and stability of the model in large data sets and complex scenarios. In general, the embodiment of the present application can complete the target calculation operation in advance before running, thereby reducing the computing resource calls during model runtime, effectively reducing computing overhead, and improving computing resource utilization, thereby improving the overall execution performance of the model, which is particularly significant in the reasoning and training scenarios of large-scale models.
[0030] The model optimization scheme provided in the embodiment of the present application can also be executed by an electronic device, which can be a server, a server cluster, or a cloud server. The electronic device can also be a terminal device such as a mobile phone, a computer, a tablet computer, a wearable device, or a dedicated device (such as a dedicated terminal device with a model optimization system). These electronic devices can also be equipped with the chips introduced in the above embodiments. Alternatively, these electronic devices can also be installed with a service program for executing the model optimization scheme.
[0031] The model optimization method and related devices provided in the embodiments of the present application are described below in conjunction with the accompanying drawings. Figure 1 A model optimization method provided in an embodiment of the present application, such as Figure 1 As shown, the method includes:
[0032] S101, obtaining a target model to be compiled from a chip hardware simulation environment;
[0033] S102, traversing each computing unit in the target model, and screening out the target computing unit containing the target computing operation;
[0034] S103, pre-calculating the target computing unit to obtain a pre-calculation result of the target computing unit;
[0035] S104: Using the pre-calculation result to replace the constant calculation operation included in the target calculation unit to obtain an optimized target model.
[0036] In the embodiment of the present application, the chip hardware simulation environment is a computing tool for simulating and testing integrated circuit design and semiconductor chip behavior. It allows engineers to perform functional verification, performance testing, and behavior analysis on the chip before physical chip manufacturing by creating a virtual model of the chip hardware in software. This environment can provide a detailed observation of the chip design and help identify and solve potential problems.
[0037] Specifically, the hardware simulation environment can simulate various behaviors of the chip under real operating conditions, including signal transmission, logic operation, and power consumption. At the same time, it can support different workflows, such as design space exploration, design verification, debugging, and power consumption analysis. Since hardware design is usually very complex, the use of a simulation environment can greatly reduce the development cycle and cost, reduce the risks brought by design defects, and make full preparations for subsequent manufacturing and deployment. In this context, the environment is used to obtain the target model for further optimization processing.
[0038] In the chip hardware simulation environment, when compiling the target model, first, the simulation tool will parse the input target model (such as Verilog or VHDL code written in hardware description language HDL) to understand its structure and design intent; then, it will perform syntax and semantic checks to ensure that there are no writing errors or logical errors in the model code, so as to avoid fatal problems during the compilation process. Then, the parsed model will be converted into an intermediate representation (IR), which is convenient for subsequent optimization and analysis. In the intermediate representation stage, the simulation tool will perform various optimization operations, such as constant folding, dead code elimination, and logic simplification, to improve the execution performance of the model. Next, the optimized intermediate representation will be mapped to specific hardware resources to determine how to implement the logical function on the hardware, including allocating registers, adjusting timing, and optimizing data paths. Subsequently, the optimized representation will be generated into binary or executable code to run on the simulator and perform further simulation. After the compilation is completed, the environment will prepare the necessary simulation settings, such as initializing signals, setting clock signals, and configuring input and output parameters. Through these steps, the chip hardware simulation environment is able to convert high-level models into detailed versions that can be verified in the simulation environment, helping engineers discover potential problems and optimize designs before actually manufacturing the chip.
[0039] In S101, a target model to be compiled is obtained from a chip hardware simulation environment.
[0040] In the embodiment of the present application, the target model is an abstract representation of the chip design, which includes multiple computing units, which are logic modules that can perform specific operations. Each computing unit may correspond to different hardware resources, such as adders, multipliers, logic gates, registers, etc.
[0041] In S102, each computing unit in the target model is traversed to filter out the target computing unit that contains the target computing operation.
[0042] As an optional embodiment, in S102, each computing unit in the target model is traversed to select the target computing unit containing the target computing operation, see Figure 2 As shown, the implementation is as follows:
[0043] S201, obtaining a calculation graph of the target model, and traversing each node in the calculation graph;
[0044] S202, determining whether the traversed node is used to perform a constant calculation operation;
[0045] S203: If the traversed node is used to perform a constant calculation operation, mark the traversed node as the target calculation unit.
[0046] In the embodiments of the present application, a computational graph is a graphical representation used to describe the computational process in a target model. In a computational graph, nodes represent computational operations, while edges represent data flows or dependencies between operations. A computational graph is an abstraction of the model execution process that can clearly show the order and dependencies of computations.
[0047] In neural networks, computation graphs are mainly used to represent the computational flow of the entire model. Nodes represent single operations or functions, and edges represent the transfer of data between operations. For a neural network model, such as the Floodlight model (assuming it is a complex deep learning model), the nodes in the computation graph are designed to represent different types of neural network layers or operations.
[0048] In the embodiment of the present application, each node in the computation graph corresponds to each computing unit in the target model. In the context of chip hardware simulation and optimization, each node in the computation graph corresponds to each computing unit in the target model, which means that each node in the computation graph represents a specific computing operation or computing module in the target model.
[0049] For example, in the context of neural networks, a computing unit generally refers to the operation of a neuron or network layer. For example, the weighted summation and activation operations in the fully connected layer can be considered as a computing unit. In neural networks, such as convolutional neural networks (CNN) or deep neural networks (DNN), matrix multiplication and addition are used extensively. These operations will become independent nodes in the computational graph, corresponding to specific hardware units or resources that process these calculations. There will be nodes in the computational graph specifically for activation function operations, such as ReLU, Sigmoid, Tanh, etc. These nodes can process data through nonlinear functions, and each activation function node corresponds to a computing unit. In convolutional neural networks, convolution operations are independent nodes, and each convolution node represents a group of specific computing units, such as the calculation of convolution kernels. These operations are also represented by independent nodes in the computational graph. The batch normalization operation node corresponds to the data standardization processing computing unit, and the pooling operation node corresponds to the data dimension reduction or feature extraction computing unit.
[0050] When traversing the computation graph, it is necessary to identify constant computation operation nodes, that is, nodes that do not depend on changes in input data during the operation. For example, certain initialization operations, fixed weight updates, pre-computed parameters, etc. may be constant computations, and these nodes can be marked as target computation units for optimization and further processing.
[0051] In an embodiment of the present application, a constant calculation operation refers to a calculation in which the value of an operand or parameter does not change with the input data during the calculation process. These calculations can be performed in advance during the compilation or initialization phase without repeating the calculation each time it is run. For example, in neural networks and other computing models, some parameters are assigned fixed constant values when the model is initialized. For example, the bias in some network layers may be initialized to zero or a specific constant. Linear transformation with fixed coefficients. If the operation contains fixed weighting coefficients, such as some layers that do not require training, their calculations can be performed in advance. Some network structures may have inherent, deterministic calculation steps, such as scale-invariant transformations or arbitrary angle rotations. In this way, by identifying these constant calculation operations, pre-calculations can be performed during the model deployment or optimization phase, thereby reducing the burden of real-time calculations and improving operating efficiency. This is particularly beneficial for hardware deployment with limited resources because it can reduce unnecessary computing overhead.
[0052] In this way, by identifying and optimizing constant computing nodes, we can improve efficiency when deploying neural network models and reduce the computational burden during actual runtime. This is especially important in hardware implementation and inference optimization, allowing models to run faster and more efficiently in resource-limited environments.
[0053] Through the mapping recognition mechanism introduced above, the implementation of neural network models in hardware environments can be more effectively designed and optimized. This abstraction and mapping is particularly important in situations where a large amount of parallel computing and memory optimization are required.
[0054] As an optional embodiment, in S201, a computation graph of the target model is obtained, and each node in the computation graph is traversed.
[0055] In step S201, the computation graph of the target model is obtained and each node in the computation graph is traversed, mainly to analyze the computation structure in the neural network one by one, identify and mark specific computation nodes (such as constant computation operation nodes), and thus optimize them.
[0056] A computation graph is a structured representation of a target model (such as a neural network model) that describes the data flow and dependencies between operations. Each layer, operation, or function in a neural network can be represented by a node in the computation graph, and the edges between these nodes represent the transfer of data between operations. The computation graph is part of the model building process, and modern frameworks (such as TensorFlow and PyTorch) automatically generate computation graphs when defining the model. Each layer's matrix multiplication, activation function, pooling layer, etc. are mapped to a node. Nodes are connected by edges, indicating that data flows from one layer to another, or from one operation to another. Traversing the computation graph means sequentially or recursively accessing each node in the computation graph to analyze the operation types of these nodes one by one. Get the computation graph object of the target model. This computation graph contains all nodes and the dependencies between them. Start traversing from the input node or the root node, and advance layer by layer according to the connection relationship between the nodes. For example, using depth-first traversal (DFS), you can start from a node, follow a path down until you reach a leaf node, and then backtrack. For example, using breadth-first traversal (BFS), you can start from the root node, first visit all nodes in one layer, and then visit the nodes in the next layer.
[0057] During the traversal, the calculation type of each node is analyzed. Constant calculation nodes perform fixed operations that are independent of the input, such as weight initialization and fixed bias addition. Variable calculation nodes depend on the input data and produce different results each time they are executed, such as convolutional layers or activation functions. If a node is identified as a constant calculation operation, it is marked as a target calculation unit for optimization in subsequent steps (such as pre-calculation).
[0058] Taking the neural network model as an example, it is very important to optimize the execution efficiency of the model by traversing each node of the computational graph, identifying and marking constant and variable calculations. Taking a neural network containing an input layer, a convolutional layer, a ReLU activation function, a pooling layer, and a fully connected layer as an example, each node of the convolutional layer, ReLU, and pooling layer is judged as a variable calculation because its operation depends on the input data or the output of the previous layer. The bias initialization node is identified as a constant calculation, and its value is fixed and does not change with the input, so it is marked as a target calculation unit. Through this traversal and analysis, constant nodes can be calculated offline before inference, thereby reducing the amount of real-time calculations, and pruning and fusion can be implemented to reduce the model size. In addition, through node type analysis, specific calculations can be mapped to specialized hardware acceleration units such as FPGA and TPU to further optimize performance.
[0059] Thus, in step S201, by traversing the computation graph of the target model and analyzing the node types in the computation graph one by one, it is possible to identify which nodes are constant computation operations and mark them as target computation units. This analysis and marking process not only helps optimize the execution efficiency of the model, but also lays the foundation for subsequent pruning, parameter fusion, and hardware acceleration.
[0060] Further optionally, in S202, the traversed nodes may be converted into corresponding instruction sequences. Then, it is determined whether the instruction sequence contains pure functions. Alternatively, it is also determined whether the instruction sequence contains constant values. If the instruction sequence contains pure functions and / or constant values, it is determined that the traversed nodes are used to perform constant calculation operations.
[0061] In step S202, pure functions refer to functions that always produce the same output and have no side effects under the same input. After combining constant values, their calculation results can be pre-calculated and cached before model reasoning. In the reasoning process, the pre-calculated results are used directly, which reduces repeated calculations and improves execution efficiency, especially in scenarios with high reasoning time requirements (such as real-time applications). By identifying constant calculation operations, the memory and processor resource usage during model runtime can be effectively reduced. Constant calculation operations do not need to be repeated, avoiding unnecessary memory allocation and processor scheduling. Especially on embedded devices or edge computing devices, this optimization can significantly reduce energy consumption and resource overhead. For nodes containing pure functions or constant values, they can be merged or pruned in the model optimization stage to simplify the model structure and further reduce the computational burden. This optimization provides a basis for model compression (such as neural network pruning, quantization and other technologies), and can also accelerate model operation by reducing unnecessary calculation instructions. The pre-calculation and caching of constant calculation operations make the calculations performed during reasoning more predictable and reduce potential floating or uncertainty factors. This is particularly critical for applications that need to ensure high reliability and stability (such as autonomous driving, medical image analysis, etc.). The identified constant calculation operations can be mapped to dedicated accelerator units in the hardware (such as FPGA, TPU, etc.), which can process fixed-mode calculations more efficiently, thereby making full use of hardware resources and further improving the execution speed of the overall model. Through these beneficial effects, the model is not only more efficient during execution, but also optimizes overall performance by reducing computational complexity and memory usage, making the deployment of neural networks in various practical scenarios more lightweight, fast and stable.
[0062] In the above steps, an optional embodiment of determining whether the instruction sequence contains a pure function may be to detect whether the output result of the calculation expression in the instruction sequence is determined by the input parameters. Among them, the same input parameter in the calculation expression of the pure function corresponds to the same output result. Furthermore, for the candidate calculation expression whose output result is determined by the input parameter, detect whether the candidate calculation expression contains global variable modification instructions, static variable modification instructions, I / O operations, and non-pure function call instructions. Detect whether the candidate calculation expression has a dependency relationship with an external variable state parameter. Detect whether the candidate calculation expression does not have referential transparency. If the detection results of the candidate calculation expression are all negative, the candidate calculation expression is a pure function.
[0063] In the above steps, the calculation expression in the instruction sequence is detected to determine whether it is a pure function. This process can accurately identify pure functions by detecting whether the output of the calculation expression is completely determined by the input parameters. The result of a pure function depends only on the input and does not depend on the external environment or state. Therefore, the result can be pre-calculated and cached before reasoning. This avoids repeated calculations in the reasoning stage and significantly improves the calculation efficiency, especially in the scenario where the pure function is called multiple times. Pure functions do not modify global variables or static variables, nor do they perform I / O operations or call other non-pure functions, so real-time dynamic calculations are not required. By analyzing the candidate calculation expressions to ensure that there is no external state dependency and side effects, these calculations can be executed or optimized in advance, reducing the calculation burden in real-time reasoning and thus improving the execution speed. Since the results of a pure function are always the same under the same input, its calculation process has referential transparency, that is, the calculation result is independent of the external state, ensuring the stability and predictability of the reasoning process. This is extremely important in applications with high stability requirements (such as financial transactions and medical systems) and can reduce the potential risk of calculation errors. Detecting whether there are global or static variable modification instructions, I / O operations, or non-pure function calls can effectively avoid computations that introduce side effects. Side effects may lead to unpredictable behaviors or errors. By excluding expressions with side effects, the reliability of the program can be improved, especially in multi-threaded or parallel computing scenarios, potential race conditions or resource competition problems can be avoided. The combination of pure function identification and constant calculation optimization can not only reduce the amount of computation at runtime, but also simplify the deployment of models on different platforms. For example, on resource-constrained embedded devices, pure function calculation results can be pre-stored to reduce the computing requirements and energy consumption of the device and optimize resource management. After determining the pure function, the compiler can optimize these calculations during the compilation phase, such as constant folding and code inlining, to further speed up the running speed of the model. In addition, the determination of pure functions provides a basis for hardware acceleration. Dedicated hardware accelerators (such as FPGAs and TPUs) can efficiently process pure function calculations and improve hardware utilization. Since pure functions do not rely on external states and do not modify global or static variables, they have good parallelization characteristics. Multiple pure function calls can be safely executed in parallel without worrying about thread safety issues. This opens up the possibility for further multi-core and multi-thread optimization, improving the performance of the model in high-performance computing environments.
[0064] In this way, by detecting whether the instruction sequence contains pure functions and their related dependencies, the above method can improve the computational efficiency of the reasoning stage, reduce side effects, optimize resource management, and lay the foundation for further compression and hardware acceleration of the model.
[0065] Further optionally, for a candidate calculation expression whose output result is determined by an input parameter, if the candidate calculation expression is a combinatorial function, the nested calculation expression of the combinatorial function is expanded, and all pure function detection operations are performed on the nested calculation expressions one by one.
[0066] It is understandable that, in the case where the candidate calculation expression is not a completely pure function calculation, the pure function part in the candidate calculation expression can also be accelerated to further accelerate the calculation efficiency of the target model.
[0067] For example, for the calculation expression F(x) = G(H(x)) + K(x), first expand the combinatorial function, parse G(H(x)) into the output of H(x) as the input of G, and then check the pure function properties one by one. After confirming that H(x) is a pure function, we can pre-calculate and cache the result of H(x), for example, use a lookup table to save the result of H(x) corresponding to each input x. When executing F(x), the result of H(x) can be directly obtained from the cache, thereby reducing the runtime calculation. At the same time, the G part of G(H(x)) is internally analyzed to find pure conversions and optimize them. In addition, if G(y) can be processed in parallel between different inputs, and the calculation of G uses a parallel mechanism for the output of H(x), then parallelization acceleration can be applied to these parts to improve the overall execution efficiency. These optimization strategies can effectively improve computing efficiency and program response speed, while optimizing resource utilization.
[0068] In this way, by caching the calculation results of pure functions, redundant calculations during real-time execution can be reduced, so that performance can be significantly optimized and computing efficiency can be improved when the same input appears multiple times. At the same time, pure function detection and caching avoid repeated execution of sub-calculation expressions, saving processor resources and time, and reducing unnecessary computing overhead. This not only enhances the responsiveness of the program, enabling it to respond more quickly when processing a large number of inputs or high concurrent access, but also optimizes resource utilization, especially in resource-limited environments (such as embedded systems), by calculating static results in advance, reducing memory reading and writing and computing resource consumption. In addition, decomposing computing tasks and identifying parallelizable parts can optimize pure functions that can be executed in parallel, so that performance can be further improved in multi-core or multi-threaded environments. Through these optimization methods, not only the overall execution efficiency of the model is effectively improved, but also a more flexible and scalable computing optimization solution is provided for complex systems.
[0069] Further optionally, after marking the traversed node as the target computing unit, the target computing unit can also be graded according to the constant utilization. Then, a dynamic cache area is established in the memory. The dynamic cache area is used to detect and update the target computing unit whose utilization is lower than the set threshold. Finally, based on the constant utilization grading result, the target computing unit is stored in the storage space corresponding to the dynamic cache area.
[0070] For example, in a large numerical computing program, in order to improve computing efficiency, the complex mathematical operations and function calls in the program are first marked, and the call frequency of each computing unit is recorded during the code execution. According to these frequencies, the computing units are graded according to the constant utilization rate, divided into high-frequency use (more than 70%), medium-frequency use (20%-70%), and low-frequency use (less than 20%). Then, a dynamic cache area is established in the memory for these different levels of computing units, and more cache resources are allocated to high-frequency units first. For computing units below the set utilization threshold, the system will detect and update in real time, and reallocate their cache resources to higher-frequency units. Finally, based on the grading results, the results or states of the computing units are stored in the storage space corresponding to the dynamic cache area to ensure that the cache can be quickly accessed and utilized the next time it is called, thereby reducing repeated calculations and improving the overall efficiency of the program.
[0071] In this way, by introducing a hierarchical and dynamic caching mechanism for the target computing unit, the computing efficiency can be significantly improved. The results of frequent calculations are graded and cached, which reduces the number of repeated calculations, thereby improving the overall execution efficiency. In addition, cache resources are dynamically allocated according to utilization, so that limited memory resources can be reasonably utilized and the cost-effectiveness of the system is optimized. The cache of high-frequency functions speeds up the processing of common operations and reduces response time, especially significantly improving the user experience in real-time or near-real-time systems. For complex computing units, such as large-scale matrix operations or recursive calculations, the cache avoids redundant operations and simplifies the processing flow. The dynamic cache area can also flexibly adapt to changes in input data or usage patterns to ensure that high-frequency computing units are optimally allocated resources. In short, this optimization method significantly saves computing resources and time in high-load or resource-constrained environments, and has great practical significance.
[0072] For a target computing unit with a higher constant utilization level, it is possible to further determine whether the target computing unit is reused. If so, the dependency of the target computing unit is optimized based on the actual reuse situation.
[0073] For example, in a neural network model, target computing units with high constant utilization levels are deeply analyzed to determine whether these units are reused. Suppose there is a program for data analysis, and one of the frequently called computing units is a large-scale matrix multiplication operation. Through analysis, it is found that in some cases, this matrix operation is reused in different algorithms. By identifying and optimizing these reuse relationships, it is decided to extract the matrix operation, perform the calculation only once when the corresponding data does not change, and store the result for reuse by other modules. This optimization reduces the repeated calls to matrix operations and significantly reduces the computational overhead in compute-intensive tasks. At the same time, the demand for memory bandwidth is reduced because there is no need to frequently re-acquire and load large amounts of data, optimizing the overall resource usage of the program. Ultimately, this optimization of dependency relationships not only improves program execution efficiency, but also makes the code more concise and easy to maintain because the reuse and dependency links between modules are clarified. This improvement is particularly prominent in high-load environments, improving the system's responsiveness and throughput, and improving the user experience.
[0074] For example, in the Bloom model, the activation function used is the Gaussian Error Linear Unit (GELU). Assume that the expression of the target calculation unit in the Bloom model is as follows:
[0075]
[0076] The instruction sequence obtained by using the compilation method in the related art is as follows:
[0077] R1 = 2
[0078] R2 = 1 / π
[0079] R3 = mulR1, R2
[0080] R4 = sqrt(R3)
[0081] R5 = mulR4, x
[0082] R6 = mulR4, 0.044715
[0083] R7 = mulR6, x
[0084] R8 = mulR7, x
[0085] R9 = mulR8, x
[0086] R10 = addR5, R9
[0087] R11 = tanh(R10)
[0088] R12 = addR11, 1
[0089] R13 = mulR12, x
[0090] R14 = mulR13, 0.5
[0091] The instruction sequence obtained after optimization using the technical solution provided in the embodiment of the present application is as follows:
[0092] R1 = sqrt(x / π)
[0093] R2 = mulR1, x
[0094] R3 = mulR1, 0.044715
[0095] R4 = mulR6, x
[0096] R5 = mulR7, x
[0097] R6 = mulR8, x
[0098] R7 = add R2, R6
[0099] R8 = tanh(R7)
[0100] R9 = addR8, 1
[0101] R10 = mulR9, x
[0102] R11 = mulR10, 0.5
[0103] Take the Gaussian error linear unit activation function as an example, where is a constant expression.
[0104] If the traditional calculation method in the related art is used to implement this expression, 4 instructions will be occupied, and 4 registers will be used in this process. However, by using the technical solution provided in the embodiment of the present application to optimize the above instruction sequence, the constant expression can be directly pre-calculated and the result can be passed to the register, so only 1 instruction and one register will be occupied. In this way, not only the constant calculation operation during the model execution process is avoided, but also the occupancy rate of hardware computing resources is greatly reduced, and the computing efficiency is improved.
[0105] By identifying and optimizing the fixed parts of the target model, constant folding can be used to reduce the amount of computation at runtime, thereby improving execution efficiency. For example, in the Bloom large model, in addition to the Gaussian error linear unit activation function, the embedding matrix, regularization parameters, mask matrix, layer normalization parameters, attention mask, and transformation matrix can all be optimized through constant folding. These optimization measures can significantly improve the inference performance of large models and make more efficient use of computing resources.
[0106] As an optional embodiment, after obtaining the target model to be compiled from the chip hardware simulation environment, the computing resource occupancy of each computing unit can also be predicted in the chip hardware simulation environment to obtain the predicted occupancy of each computing unit. Then, it is identified whether the computing unit whose predicted occupancy is higher than the set threshold contains a computing operation that meets the pre-computation condition. The pre-computation condition is that the expression corresponding to the computing operation can be converted into a constant expression. If it contains a computing operation that meets the pre-computation condition, the expression conversion is performed on the computing unit whose predicted occupancy is higher than the set threshold. Finally, the constant expression obtained by the conversion is pre-calculated, and the pre-calculation result is replaced in the computing unit whose predicted occupancy is higher than the set threshold.
[0107] In the chip hardware simulation environment, the target model to be compiled can be carefully analyzed to predict the computing resource usage of each computing unit. For example, when designing a neural network model loaded by a digital signal processing chip, it was found that a certain signal processing algorithm contains a large number of Fourier transform operations. Through the prediction of computing resource usage, it was found that the usage of some computing units in these operations was higher than the set threshold. Further analysis of these computing units identified that some computing operations can actually be converted into constant expressions, such as certain constant coefficients in Fourier transform or repeatedly used fixed input parameters. By converting these operations into constant expressions and pre-calculating them, we can significantly reduce the complexity of real-time calculations.
[0108] Thus, after conversion into constant expressions and pre-calculation, a large amount of unnecessary real-time computing load is reduced, thereby effectively reducing the power consumption and heat generation of the chip. Secondly, this expression pre-calculation and replacement improves the speed of overall signal processing, allowing the chip to respond more quickly in high-frequency operations, which is especially important for large data streams that need to be processed in real time, such as audio and video signal processing. Ultimately, this optimization not only improves the utilization efficiency of hardware resources, but also extends the life and reliability of the chip. At the same time, performing such optimization in the design phase can reduce production costs because a simpler hardware design can be selected to achieve the same functional performance.
[0109] In the above optional embodiment, further optionally, after the identification of whether the computing unit whose predicted occupancy is higher than the set threshold includes computing operations that meet the pre-computation conditions, if it contains computing operations that meet the pre-computation conditions, the dependency of the computing operations that meet the pre-computation conditions is adjusted to isolate the dependency between the computing operations that meet the pre-computation conditions and the external variable state.
[0110] For example, in large-scale flood models, computational tasks are often highly complex and resource-intensive, and optimizing them can significantly improve performance. In such models, adjusting the dependencies of computational units can effectively isolate the dependencies between computational operations that meet pre-computation conditions and external variable states. During the forward propagation of the model, identify computational blocks that do not change with changes in input data. For example, in some layers, some activation functions or fixed-parameter matrix multiplications may be constant, especially when the network structure is fixed and some weights are constant during the learning process. Extract these constant computational blocks and reconstruct the computational graph so that they are only calculated once after the model is initialized or the weights are fixed. For example, if a part of the model is always multiplied with a fixed word embedding matrix regardless of the input data, this multiplication operation can be reconstructed as an independent process to avoid repeated calculations each time data is input. Store the extracted constant calculation results in a cache to ensure that these results are directly used in each forward propagation without being recalculated. This requires the introduction of an intermediate data cache mechanism in the model to store the fixed results of the computational expressions. During the refactoring process, ensure that these independent constant calculation modules no longer rely on any input data or external mutable state. For operations such as matrix multiplication, this means that it is already completed during the model import phase and is separated from the dynamic input. In addition, when performing a large number of model validations, ensure that the refactoring and cache introduction do not affect the accuracy of training or reasoning. In particular, verify whether the cache mechanism for constant calculation is reliable and effective under different input conditions.
[0111] Through the above steps, the calculation part that is not related to the external variable state can be isolated in the large-scale flood light model, and the pre-calculation and cache storage of this part can be realized. This method can reduce unnecessary computing consumption and improve the execution efficiency of the model. Especially in online reasoning scenarios, this optimization can significantly improve the overall performance and response speed by reducing computing costs and delays.
[0112] Further optionally, after the dependency adjustment is performed on the computing operations that meet the pre-computation conditions, it is also possible to verify whether the dependency adjustment process of the computing operations that meet the pre-computation conditions is successful. If it is verified that the dependency adjustment fails, idle resources in the chip hardware simulation environment are called to execute the computing operations that meet the pre-computation conditions to achieve computational acceleration of the target model.
[0113] In the floodlight model, for computational operations that meet the pre-computation conditions, the success or failure of dependency adjustment has an important impact on the overall performance and resource utilization of the model. After completing the dependency adjustment of the pre-computation conditions, the model's dependencies are first verified. This can be done through dependency analysis tools or custom test frameworks, which simulate and test the adjusted modules to ensure that these modules no longer rely on external mutable states. Verification methods can include testing the stability of the module under different input conditions, or using automated testing to check whether the pre-computation module maintains the same output results when the input changes.
[0114] If the dependency adjustment fails during the verification process, it means that the pre-calculation module still depends on external variable states. At this time, identify these modules that have not been successfully adjusted, and record their locations and dependencies for subsequent processing. For modules that have failed to adjust dependencies, you can use the idle computing resources of the chip hardware simulation environment to accelerate the calculation. This accelerated calculation can be to dynamically allocate computing tasks to the current idle resource pool to ensure that the pre-calculation operations that have not been successfully adjusted can also be executed in parallel in multiple idle units. Create a task scheduling queue in the simulation environment, put these operations that have failed to adjust dependencies into the queue, and use idle resources to reduce the overall execution time of the model. During the execution process, cache the calculation results and synchronize them to the model to ensure that even if the dependency adjustment fails, the system can still use idle resources to achieve higher execution efficiency. In the simulation environment, monitor the execution effect of the modules that have not been successfully adjusted, and continuously optimize the scheduling algorithm to improve the model operation efficiency. For example, by preferentially allocating high-frequency tasks to low-latency idle resources, the impact of unadjusted modules on overall performance can be further reduced.
[0115] Through the above processing, the floodlight model can continue to use the idle resources in the simulation environment to accelerate the calculation when the dependency adjustment fails, thereby achieving the purpose of performance optimization and efficient resource utilization.
[0116] Further optionally, in the above optional embodiment, predicting the computing resource occupancy of each computing unit in a chip hardware simulation environment to obtain the predicted occupancy of each computing unit can be implemented as the following steps:
[0117] 301. Calling a resource prediction model loaded in a chip hardware simulation environment; the resource prediction model is pre-trained based on the model structures of multiple models and historical model operation data;
[0118] 302. Construct a virtual mapping model to be analyzed based on the model structure of the target model; the virtual mapping model and the target model have the same computing units and connection relationships between computing units;
[0119] 303. Collect current environmental parameters and construct a virtual input data set in combination with the historical model operation data;
[0120] 304. Input the virtual input data set into the virtual mapping model to perform multi-modal operation monitoring, and obtain the operation flow data of each computing unit in the virtual mapping model under different modes;
[0121] 305. Predict resource occupancy based on the operating traffic data of each computing unit in different modes to obtain predicted occupancy of each computing unit.
[0122] Taking the Pan-Light Big Model as an example, the various steps of predicting computing resource usage in the chip hardware simulation environment are further explained. Pan-Light Big Models usually refer to deep learning models with a large number of parameters, such as language models. Predicting the usage of computing resources is crucial to optimizing resource utilization. In the resource prediction of the Pan-Light Big Model, the resource prediction model that has been loaded in the chip hardware simulation environment must be called first. This resource prediction model is pre-trained based on multiple model structures and historical model operation data, and is used to predict the resource requirements of the target model. This model can make accurate predictions based on different model structures (for example, the number of layers, the number of parameters, the connection method, etc.) and historical operation data (for example, the usage of computing, storage, bandwidth and other resources when the model was previously running).
[0123] Next, a virtual mapping model is constructed based on the specific model structure of the omni-optical model (including information such as the arrangement of layers, the number of parameters, and activation functions). This virtual mapping model has exactly the same computing units (for example, neurons, convolution operations, fully connected layers, etc. in each layer) and the connection relationship between computing units as the target model. The purpose of constructing a virtual mapping model is to map the model in a chip simulation environment for subsequent resource occupancy prediction. The role of the virtual mapping model is to simulate the operation of the real model on the chip, but without actual data processing, it is predicted based on the model structure. This step collects parameters in the current operating environment, such as the current hardware configuration, memory bandwidth, network communication rate, etc. These parameters are very important for determining the resource occupancy of the computing unit. Then, a virtual input data set is constructed in combination with the operating data of the historical model (for example, the actual operation of the omni-optical model with a similar structure in the same hardware environment). The virtual input data set is used to simulate the operation of the target model in the current environment. It can contain simulated input data samples and resource consumption information based on historical operation.
[0124] Then, the constructed virtual input data set is input into the virtual mapping model, and multimodal operation monitoring is performed. During this process, the simulation environment simulates the execution of the target model in different operation modes (or modes). For example, the resource usage difference between the model in inference mode and training mode can be simulated. Through this multimodal monitoring, the operation traffic data of each computing unit (such as each layer, each computing block) in the virtual mapping model in different modes can be obtained. These traffic data include the amount of calculation required to process each input sample, memory usage, bandwidth requirements, etc.
[0125] Further optionally, in inference mode, the model processes the input data without parameter updates to generate predictions. For large flood models, this typically involves large matrix multiplications. Multimodal operation monitoring focuses on monitoring computational load, memory usage, and bandwidth requirements to help identify computational bottlenecks and memory reuse opportunities during inference. For example, monitoring the resource consumption of the self-attention mechanism under inputs of different lengths can help optimize inference efficiency.
[0126] In training mode, the model continuously adjusts parameters through forward propagation and back propagation to optimize the model's predictive ability. Monitoring during the training phase focuses on the balance of forward and backward computing loads, memory usage, and communication overhead in distributed training. For example, when training large-scale language models, the amount of back propagation is often greater than that of forward propagation. Monitoring resource consumption in these stages helps optimize training performance and avoid memory overflow or communication bottlenecks.
[0127] The mixed precision training mode accelerates model training by using floating point numbers of different precisions while reducing memory usage. In this mode, multimodal monitoring focuses on the balance between the numerical accuracy and computational efficiency of the model. By monitoring the accuracy and acceleration effect of the computing units, it ensures that the model maintains accuracy while improving performance. For example, mixed precision training can significantly increase the speed of model training, but it is necessary to monitor whether it will cause loss of model accuracy.
[0128] Pruning mode is used to reduce redundant parameters in the model and speed up the model operation by removing useless computing units. Monitoring the computing unit utilization after pruning, the accuracy change of the model, and the acceleration effect are key. For example, monitoring the performance of the pruned model on different computing devices (such as GPU, TPU) to ensure that the prediction accuracy of the model is not significantly lost while the inference speed is improved.
[0129] Finally, based on the running traffic data of each computing unit in different modes obtained in the previous step, the resource occupancy is predicted. Specifically, the collected traffic data can be used to model the occupancy of computing resources, including the computing cycles required by each unit, memory consumption, and communication overhead.
[0130] This prediction not only covers the resource usage of a single computing unit, but also takes into account the parallel computing, communication delays and dependencies between computing units. This can accurately predict the resource requirements of each computing unit under different conditions and generate a detailed prediction result to guide resource allocation and optimize model deployment.
[0131] Through the above steps, we can accurately predict the resource usage of each computing unit in the large flood model, thereby optimizing the resource utilization in the chip hardware simulation environment. This method provides strong support for the operation efficiency of the model, can better schedule computing resources, reduce computing bottlenecks, and significantly improve performance, especially when facing reasoning and training tasks of large models.
[0132] In the embodiments of the present application, the target computing operations can be completed in advance, thereby reducing the computing resources required to be called when the model is running, effectively reducing unnecessary computing overhead, improving the utilization of computing resources, and thus improving the overall execution performance of the model, especially in reasoning or training scenarios involving large-scale models.
[0133] Based on the same implementation principle, an embodiment of the present application also provides a model optimization device for implementing the model optimization method in the above embodiment.
[0134] Figure 2 The structural block diagram of the model optimization device provided in the embodiment of the present application is as follows: Figure 3 As shown, the device comprises at least the following units:
[0135] An acquisition unit is configured to acquire a target model to be compiled from a chip hardware simulation environment;
[0136] A screening unit is configured to traverse each computing unit in the target model and screen out a target computing unit containing a target computing operation; the target computing operation does not depend on dynamic data at runtime, and the execution logic of the target computing operation is not associated with the runtime operation;
[0137] a pre-calculation unit, configured to pre-calculate the target calculation unit to obtain a pre-calculation result of the target calculation unit;
[0138] The replacement unit is configured to use the pre-calculation result to replace the constant calculation operation contained in the target calculation unit to obtain an optimized target model.
[0139] In an optional embodiment, when the screening unit traverses each computing unit in the target model and screens out the target computing unit containing the target computing operation, it is configured to: obtain a computing graph of the target model and traverse each node in the computing graph; each node in the computing graph corresponds to each computing unit in the target model;
[0140] Determine whether the traversed node is used to perform constant calculation operations;
[0141] If the traversed node is used to perform a constant calculation operation, the traversed node is marked as the target calculation unit.
[0142] In an optional embodiment, when the screening unit determines whether the traversed node is used to perform a constant calculation operation, it is configured to:
[0143] Convert the traversed nodes into corresponding instruction sequences;
[0144] Determining whether the instruction sequence contains a pure function, and / or determining whether the instruction sequence contains a constant value;
[0145] If the instruction sequence includes pure functions and / or constant values, it is determined that the traversed nodes are used to perform constant calculation operations.
[0146] In an optional embodiment, when the screening unit determines whether the instruction sequence includes a pure function, it is configured to:
[0147] Detecting whether an output result of a calculation expression in the instruction sequence is determined by an input parameter; wherein the same input parameter in a calculation expression of a pure function corresponds to the same output result;
[0148] For a candidate calculation expression whose output result is determined by an input parameter, detecting whether the candidate calculation expression contains a global variable modification instruction, a static variable modification instruction, an I / O operation, or a non-pure function call instruction;
[0149] Detecting whether the candidate calculation expression has a dependency relationship with an external variable state parameter;
[0150] Detecting whether the candidate computation expression does not have referential transparency;
[0151] If the detection results of the candidate calculation expression are all negative, then the candidate calculation expression is a pure function.
[0152] In an optional embodiment, for a candidate calculation expression whose output result is determined by input parameters, if the candidate calculation expression is a combinatorial function, the screening unit is further configured to expand the nested calculation expression of the combinatorial function and perform all pure function detection operations on the nested calculation expressions one by one.
[0153] In an optional embodiment, after the screening unit marks the traversed node as the target computing unit, it is further configured to: perform constant utilization grading on the target computing unit;
[0154] Establishing a dynamic buffer in the memory; the dynamic buffer is used to detect and update the target computing unit whose utilization rate is lower than a set threshold;
[0155] The target computing unit is stored in a storage space corresponding to the dynamic cache area based on the constant utilization classification result.
[0156] In an optional embodiment, it further includes a prediction unit configured to predict the computing resource occupancy of each computing unit in the chip hardware simulation environment after the acquisition unit acquires the target model to be compiled from the chip hardware simulation environment, so as to obtain the predicted occupancy of each computing unit;
[0157] Identify whether the computing unit whose predicted occupancy is higher than the set threshold contains a computing operation that meets the pre-computation condition; the pre-computation condition is that the expression corresponding to the computing operation can be converted into a constant expression;
[0158] If it contains computing operations that meet the pre-computation conditions, then the expression conversion is performed on the computing units whose predicted occupancy is higher than the set threshold;
[0159] The converted constant expressions are pre-calculated, and the pre-calculated results are replaced into the computing units whose predicted occupancy is higher than the set threshold.
[0160] In an optional embodiment, the prediction unit is further configured to: after identifying whether the computing unit whose predicted occupancy is higher than a set threshold includes computing operations that meet the pre-computation conditions, if it includes computing operations that meet the pre-computation conditions, adjust the dependencies of the computing operations that meet the pre-computation conditions to isolate the dependencies between the computing operations that meet the pre-computation conditions and the external variable state.
[0161] In an optional embodiment, the prediction unit is further configured to: after adjusting the dependencies of the computing operations that meet the pre-computation conditions, verify whether the dependency adjustment processing of the computing operations that meet the pre-computation conditions is successful; if it is verified that the dependency adjustment fails, call the idle resources in the chip hardware simulation environment to execute the computing operations that meet the pre-computation conditions, so as to achieve computational acceleration of the target model.
[0162] In an optional embodiment, the prediction unit predicts the computing resource occupancy of each computing unit in the chip hardware simulation environment, and when obtaining the predicted occupancy of each computing unit, is configured as follows:
[0163] Calling a resource prediction model loaded in a chip hardware simulation environment; the resource prediction model is pre-trained based on the model structures of multiple models and historical model operation data;
[0164] Based on the model structure of the target model, a virtual mapping model to be analyzed is constructed; the virtual mapping model and the target model have the same computing units and connection relationships between computing units;
[0165] Collect current environmental parameters, combine with the historical model operation data, and construct a virtual input data set;
[0166] Inputting the virtual input data set into the virtual mapping model for multi-modal operation monitoring, and obtaining the operation flow data of each computing unit in the virtual mapping model under different modes;
[0167] The resource occupancy is predicted based on the operating traffic data of each computing unit in different modes to obtain the predicted occupancy of each computing unit.
[0168] It should be noted that, regarding the model optimization device provided by the implementation of this application, the specific functions and implementation details of each module therein have the same implementation principles as the implementation process of the corresponding steps in the aforementioned method embodiment. For details, please refer to the description of the corresponding parts in the aforementioned method embodiment, which will not be repeated here.
[0169] Based on the same implementation principle, the embodiment of the present application also provides an electronic device, Figure 4 is a structural block diagram of the electronic device 400, such as Figure 4 As shown, the electronic device 400 includes: a processor 401, a memory 402, a communication interface 403, a communication bus 404 and a controller 405; wherein the processor 401, the memory 402, and the communication interface 403 communicate with each other via the communication bus 404; the memory 402 is used to store computer programs; the processor 401 is used to execute the programs stored in the memory 402 to implement corresponding processing functions; the communication interface 404 is used for communication between the electronic device 400 and other devices; and the controller 405 is used to implement the model optimization method described in the above method embodiment.
[0170] In the embodiment of the present application, the communication bus 404 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus 404 may be divided into an address bus, a data bus, a control bus, etc., and the specific form is not limited. For ease of representation, Figure 3Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0171] The memory 402 may include a random access memory (RAM) or a non-volatile memory (non-volatile memory), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far away from the processor 401 .
[0172] Processor 401 may be a general purpose processor, including a central processing unit (Central Processor-
[0173] It can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other editable logic devices, discrete gate or transistor logic devices, discrete hardware components. The specific form can be determined according to actual needs.
[0174] Accordingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, which, when executed, can implement each step that can be executed by an electronic device in the above method embodiment.
[0175] It should be noted that although one or more embodiments of the present application provide method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative labor. The order of steps listed in the embodiments is only one way of executing the steps among many orders, and does not represent the only order of execution. When the actual device or client product is executed, it can be executed in the order of the method shown in the embodiments or the drawings or in parallel (for example, in a parallel processor or multi-threaded processing environment).
[0176] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, devices (systems) or computer program products. Therefore, the embodiments of the present application can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, one or more embodiments of the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0177] One or more embodiments of the present application are described with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to one or more embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0178] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0179] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0180] Each embodiment in this application is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0181] In this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. For those of ordinary skill in the art, the specific meaning of the above terms in one or more embodiments of the present application can be understood according to the specific circumstances.
[0182] It should be noted that, in the absence of conflict, the features in one or more embodiments of the present application and the embodiments may be combined with each other. The one or more embodiments of the present application are not limited to any single aspect, nor to any single embodiment, nor to any combination and / or replacement of these aspects and / or embodiments. Moreover, each aspect and / or embodiment of one or more embodiments of the present application may be used alone or in combination with one or more other aspects and / or embodiments thereof.
[0183] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of one or more embodiments of the present application, rather than to limit them. Although one or more embodiments of the present application have been described in detail with reference to the aforementioned embodiments, a person of ordinary skill in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some or all of the technical features may be replaced by equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of one or more embodiments of the present application, and they should all be included in the scope of the claims and specifications of one or more embodiments of the present application.
[0184] One or more embodiments of the present application are described above in conjunction with optional implementation methods, but these implementation methods are only exemplary and serve only as an illustration. On this basis, multiple replacements and improvements can be made to one or more embodiments of the present application, all of which fall within the scope of protection of one or more embodiments of the present application.
Claims
1. A model optimization method, characterized in that: The method comprises: Obtain the target model to be compiled from the chip hardware simulation environment; Traversing each computing unit in the target model, and screening out a target computing unit containing a target computing operation; the target computing operation does not depend on dynamic data at runtime, and the execution logic of the target computing operation is not associated with the runtime operation; Pre-calculating the target computing unit to obtain a pre-calculation result of the target computing unit; The pre-calculation result is used to replace the constant calculation operation contained in the target calculation unit to obtain an optimized target model; After obtaining the target model to be compiled from the chip hardware simulation environment, the method further includes: Predicting the computing resource usage of each computing unit in a chip hardware simulation environment to obtain the predicted usage of each computing unit; Identify whether the computing unit whose predicted occupancy is higher than the set threshold contains a computing operation that meets the pre-computation condition; the pre-computation condition is that the expression corresponding to the computing operation can be converted into a constant expression; If it contains computing operations that meet the pre-computation conditions, then the expression conversion is performed on the computing units whose predicted occupancy is higher than the set threshold; The converted constant expressions are pre-calculated, and the pre-calculated results are replaced into the computing units whose predicted occupancy is higher than the set threshold.
2. The method according to claim 1, characterized in that The traversing each computing unit in the target model and filtering out the target computing unit containing the target computing operation includes: Obtaining a computation graph of the target model, and traversing each node in the computation graph; each node in the computation graph corresponds to each computation unit in the target model; Determine whether the traversed node is used to perform constant calculation operations; If the traversed node is used to perform a constant calculation operation, the traversed node is marked as the target calculation unit.
3. The method according to claim 2, characterized in that The determining whether the traversed node is used to perform a constant calculation operation includes: Convert the traversed nodes into corresponding instruction sequences; Determining whether the instruction sequence contains a pure function, and / or determining whether the instruction sequence contains a constant value; If the instruction sequence includes pure functions and / or constant values, it is determined that the traversed nodes are used to perform constant calculation operations.
4. The method according to claim 3, characterized in that The determining whether the instruction sequence includes a pure function comprises: Detecting whether an output result of a calculation expression in the instruction sequence is determined by an input parameter; wherein the same input parameter in a calculation expression of a pure function corresponds to the same output result; For a candidate calculation expression whose output result is determined by an input parameter, detecting whether the candidate calculation expression contains a global variable modification instruction, a static variable modification instruction, an I / O operation, or a non-pure function call instruction; Detecting whether the candidate calculation expression has a dependency relationship with an external variable state parameter; Detecting whether the candidate computation expression does not have referential transparency; If the detection results of the candidate calculation expression are all negative, then the candidate calculation expression is a pure function.
5. The method according to claim 4, characterized in that For a candidate calculation expression whose output result is determined by input parameters, if the candidate calculation expression is a combinatorial function, then Expand the nested calculation expressions of the composite function and perform all pure function detection operations on the nested calculation expressions one by one.
6. The method according to claim 2, characterized in that After marking the traversed node as the target computing unit, the method further includes: Classifying the target computing unit according to constant utilization; Establishing a dynamic buffer in the memory; the dynamic buffer is used to detect and update the target computing unit whose utilization rate is lower than a set threshold; The target computing unit is stored in a storage space corresponding to the dynamic cache area based on the constant utilization classification result.
7. The method according to claim 1, characterized in that After identifying whether the computing unit whose predicted occupancy is higher than the set threshold includes computing operations that meet the pre-computation conditions, if the computing unit includes computing operations that meet the pre-computation conditions, the method further includes: Dependency adjustment is performed on computation operations that are eligible for precomputation to isolate dependencies between computation operations that are eligible for precomputation and external mutable states.
8. The method according to claim 7, characterized in that After the dependency adjustment is performed on the computing operations that meet the pre-computation conditions, the method further includes: Verify whether the dependency adjustment process of the computing operations that meet the pre-computation conditions is successful; If it is verified that the dependency adjustment fails, the idle resources in the chip hardware simulation environment are called to perform computing operations that meet the pre-computation conditions to achieve computing acceleration of the target model.
9. The method according to claim 1, characterized in that: The step of predicting the computing resource occupancy of each computing unit in the chip hardware simulation environment to obtain the predicted occupancy of each computing unit includes: Calling a resource prediction model loaded in a chip hardware simulation environment; the resource prediction model is pre-trained based on the model structures of multiple models and historical model operation data; Based on the model structure of the target model, a virtual mapping model to be analyzed is constructed; the virtual mapping model and the target model have the same computing units and connection relationships between computing units; Collect current environmental parameters, combine with the historical model operation data, and construct a virtual input data set; Inputting the virtual input data set into the virtual mapping model for multi-modal operation monitoring, and obtaining the operation flow data of each computing unit in the virtual mapping model under different modes; The resource occupancy is predicted based on the operating traffic data of each computing unit in different modes to obtain the predicted occupancy of each computing unit.
10. A model optimization device, characterized in that: The device comprises at least the following units: An acquisition unit is configured to acquire a target model to be compiled from a chip hardware simulation environment; A screening unit is configured to traverse each computing unit in the target model and screen out a target computing unit that contains a target computing operation; The target computing operation does not depend on dynamic data at runtime, and the execution logic of the target computing operation is not associated with the runtime operation; a pre-calculation unit, configured to pre-calculate the target calculation unit to obtain a pre-calculation result of the target calculation unit; A replacement unit, configured to replace the constant calculation operation contained in the target calculation unit with the pre-calculation result to obtain an optimized target model; The prediction unit is configured to predict the computing resource occupancy of each computing unit in the chip hardware simulation environment after the acquisition unit acquires the target model to be compiled from the chip hardware simulation environment, and obtain the predicted occupancy of each computing unit; identify whether the computing unit with the predicted occupancy higher than the set threshold contains a computing operation that meets the pre-computation condition; the pre-computation condition is that the expression corresponding to the computing operation can be converted into a constant expression; if it contains a computing operation that meets the pre-computation condition, then perform expression conversion on the computing unit with the predicted occupancy higher than the set threshold; The converted constant expressions are pre-calculated, and the pre-calculated results are replaced into the computing units whose predicted occupancy is higher than the set threshold.
11. A chip, characterized in that: The chip includes a processor coupled to a transceiver, and is used to execute the model optimization method according to any one of claims 1 to 9.
12. An electronic device, characterized in that: The system comprises a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the model optimization method according to any one of claims 1 to 9.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program instructions, which, when executed, implement the model optimization method described in any one of claims 1 to 9.
Citation Information
Patent Citations
Compiling method and device of deep learning algorithm and related product
CN111667060A
Neural network optimization method and device and processor
CN112200297A