Model fusion method and related device

By identifying and optimizing the computing modules in the neural network model, operating fusion is carried out to optimize the connection relationship and data transmission method between modules, the problems of low model execution efficiency and waste of computing resources are solved, and hardware acceleration and performance improvement are achieved.

CN119150922BActive Publication Date: 2025-06-06SHANGHAI TAIZE SEMICONDUCTOR CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411269990.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-11
Publication Date
2025-06-06
Estimated Expiration
2044-09-11

AI Technical Summary

Technical Problem

The execution efficiency of the model and waste of computing resources due to excessive read and write operations in the neural network model.

Method used

By identifying the calculation module to be optimized, operating and fusion are carried out according to the matching module fusion method, optimizing the connection relationship and data transmission method between modules, thereby achieving hardware acceleration.

Benefits of technology

It reduces unnecessary memory read and write operations, improves the utilization rate of computing resources, shortens the time of the computing process, reduces energy consumption, and improves the overall performance and execution efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119150922B_ABST
    Figure CN119150922B_ABST
Patent Text Reader

Abstract

The embodiment of the present application provides a model fusion method and related devices. The method includes: identifying the first computing module and the second computing module to be optimized in the neural network model; the first computing module and the second computing module meet the preset optimization conditions; according to the module fusion method matching the model to be optimized, the first computing module and the second computing module are operated and fused to obtain the target neural network model; the module fusion method at least includes: the target connection relationship and data transmission method between the optimized first computing module and the second computing module; executing the target neural network model to achieve hardware acceleration of the target neural network model. This method helps to avoid the technical problems of reduced model execution efficiency and waste of computing resources due to too many read and write operations in the neural network model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of data processing. More specifically, the embodiments of the present application relate to a model fusion method and related devices. Background Art

[0002] In the field of modern artificial intelligence, especially in machine learning and deep learning applications, the efficiency and performance of algorithms are crucial. Neural network models, such as the Generative Pre-trained Transformation (GPT) model and the Bloom model in the Large Language Model (LLM), require a lot of computing resources for data processing and model training. These processes are mainly composed of many complex mathematical calculations, so how to efficiently use hardware resources to improve computing efficiency is crucial.

[0003] In the related art, the neural network model is composed of various functional modules, which are clearly divided among each other. However, data needs to be transferred between modules through memory space, which generates unnecessary read and write operations, making the model execution process longer and reducing the model execution efficiency. In addition, due to the division between modules, the subsequent modules need to wait for the previous modules to complete the calculation and obtain all the result data before continuing the calculation, which may cause the computing unit to be idle and waste hardware computing resources.

[0004] In summary, there is an urgent need to provide a new technical solution to solve at least one of the above-mentioned technical problems existing in the relevant technologies. Summary of the invention

[0005] In this context, the embodiments of the present application hope to provide a model fusion method and related devices to avoid the technical problems of decreased model execution efficiency and waste of computing resources due to excessive read and write operations in the neural network model.

[0006] In a first aspect of the embodiments of the present application, a model fusion method is provided, the method comprising:

[0007] Identify a first computing module and a second computing module to be optimized in the neural network model; the first computing module and the second computing module meet a preset optimization condition;

[0008] According to the module fusion method matching the model to be optimized, the first computing module and the second computing module are operated and fused to obtain a target neural network model; the module fusion method at least includes: a target connection relationship and a data transmission method between the optimized first computing module and the second computing module;

[0009] Execute the target neural network model to achieve hardware acceleration of the target neural network model.

[0010] In a second aspect of the embodiments of the present application, a model fusion device is provided, the device comprising at least:

[0011] An identification module, used to identify a first calculation module and a second calculation module to be optimized in the neural network model; the first calculation module and the second calculation module meet the preset optimization conditions;

[0012] A fusion module is used to operate and fuse the first computing module and the second computing module according to a module fusion method that matches the model to be optimized, so as to obtain a target neural network model; the module fusion method at least includes: a target connection relationship and a data transmission method between the optimized first computing module and the second computing module;

[0013] The execution module is used to execute the target neural network model and realize hardware acceleration of the target neural network model.

[0014] In a third aspect of the embodiments of the present application, a computing device is provided, the computing device comprising:

[0015] at least one processor, memory, and input-output unit;

[0016] The memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the model fusion method of the first aspect.

[0017] In a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, which includes instructions, and when the instructions are executed on a computer, the computer executes the model fusion method of the first aspect.

[0018] In a fifth aspect of the embodiments of the present application, a computer program product is provided, comprising a computer program, which implements the model fusion method of the first aspect when executed by a processor.

[0019] In a sixth aspect of the embodiments of the present application, a chip is provided, which includes a processor coupled to a transceiver, and is used to execute the model fusion method of the first aspect.

[0020] In the seventh aspect of the implementation of the present application, a chip system is provided, which includes: a communication interface for inputting and / or outputting information; a processor for executing a computer executable program so that a device equipped with the chip system executes the model fusion method of the first aspect.

[0021] In the implementation manner of the present application, a model fusion method and related devices are provided. In the implementation manner of the present application, first, the first computing module and the second computing module to be optimized in the neural network model are identified; the first computing module and the second computing module meet the pre-set optimization conditions. Secondly, according to the module fusion method that matches the model to be optimized, the first computing module and the second computing module are operationally fused to obtain a target neural network model; the module fusion method at least includes: the target connection relationship between the optimized first computing module and the second computing module and the data transmission method. Finally, the target neural network model is executed to achieve hardware acceleration of the target neural network model. In the implementation manner of the present application, by operationally fusion of the first computing module and the second computing module in the neural network model, the computing resource usage is reduced, the overall performance of the model is significantly improved, and the model execution efficiency is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] By reading the detailed description below with reference to the accompanying drawings, the above and other purposes, features and advantages of the exemplary embodiments of the present application will become readily understood. In the accompanying drawings, several embodiments of the present application are shown in an exemplary and non-limiting manner, wherein:

[0023] Figure 1 A schematic diagram of a flow chart of a model fusion method provided in one embodiment of the present application;

[0024] Figure 2 A schematic diagram of the implementation process of step S102 provided in an optional embodiment of the present application;

[0025] Figure 3 A schematic diagram of the principle of model fusion provided for an optional embodiment of the present application;

[0026] Figure 4 A schematic diagram of the structure of a multi-layer perceptron module of a large floodlight model provided in an optional embodiment of the present application;

[0027] Figure 5 A schematic diagram of the structure of a multi-layer perceptron module provided in a specific example of the present application;

[0028] Figure 6 A schematic diagram of the structure of a model fusion device provided in one embodiment of the present application;

[0029] Figure 7 A schematic diagram of the structure of a medium in an embodiment of the present application is schematically shown;

[0030] Figure 8 A schematic diagram of the structure of a computing device according to an embodiment of the present application is shown schematically.

[0031] In the drawings, the same or corresponding reference numerals represent the same or corresponding parts. DETAILED DESCRIPTION

[0032] The principles and spirit of the present application will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and implement the present application, and are not intended to limit the scope of the present application in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.

[0033] Those skilled in the art know that the embodiments of the present application can be implemented as a system, device, equipment, method or computer program product. Therefore, the present disclosure can be specifically implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0034] In the related art, the neural network model is composed of various functional modules, which are clearly divided among each other. However, data needs to be transferred between modules through memory space, which generates unnecessary read and write operations, making the model execution process longer and reducing the model execution efficiency. In addition, due to the division between modules, the subsequent modules need to wait for the previous modules to complete the calculation and obtain all the result data before continuing the calculation, which may cause the computing unit to be idle and waste hardware computing resources.

[0035] In summary, there is an urgent need to provide a new technical solution to solve at least one of the above-mentioned technical problems existing in the relevant technologies.

[0036] In order to overcome the above-mentioned technical problems existing in the existing solutions, according to the implementation mode of the present application, a model fusion method and related devices are proposed.

[0037] In the implementation mode of the present application, first, the first computing module and the second computing module to be optimized in the neural network model are identified; the first computing module and the second computing module meet the preset optimization conditions. Secondly, according to the module fusion method matching the model to be optimized, the first computing module and the second computing module are operated and fused to obtain the target neural network model; the module fusion method at least includes: the target connection relationship and the data transmission method between the optimized first computing module and the second computing module. Finally, the target neural network model is executed to achieve hardware acceleration of the target neural network model.

[0038] Specifically, first, operation fusion reduces unnecessary memory read and write operations, allowing data to be directly used for subsequent calculations before being stored in memory, thereby reducing the use of memory bandwidth and access delays, and improving the speed of reasoning. Secondly, the fusion technology allows the operations of the first and second modules to be executed synchronously, avoiding the idle state of computing resources, improving the utilization of computing resources, and shortening the time of the computing process. In addition, operation fusion optimizes data transmission between modules, reduces data transmission overhead and bandwidth consumption, especially in multi-layer neural networks, this optimization is particularly important for improving overall efficiency. At the same time, by reducing unnecessary memory access and transmission overhead, the implementation method of the present application also reduces energy consumption and improves the energy efficiency of hardware, which is particularly suitable for edge devices or low-power environments. Finally, operation fusion enhances the scalability of deep learning models, allowing models to become larger and more complex without sacrificing performance, effectively reducing resource usage and improving the efficiency of training and reasoning.

[0039] The technical solutions provided in the embodiments of the present application can be implemented by servers and / or terminal devices. It should be noted that the server involved in the embodiments of the present application can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, big data and artificial intelligence platforms.

[0040] The terminal device involved in the embodiments of the present application may be a device that provides voice and / or data connectivity to a user, a handheld device with a wireless connection function, or other processing devices connected to a wireless modem. For example, a mobile phone (or "cellular" phone) and a computer with a mobile terminal, for example, may be a portable, pocket-sized, handheld, computer-built-in or vehicle-mounted mobile device that exchanges voice and / or data with a wireless access network. For example, Personal Communication Service (PCS) phones, cordless phones, Session Initiation Protocol (SIP) phones, Wireless Local Loop (WLL) stations, Personal Digital Assistants (PDAs), and other devices.

[0041] It should be understood herein that any number of elements in the drawings is for illustration rather than limitation, and any naming is only for distinction rather than having any limiting meaning.

[0042] The principles and spirit of the present application are explained in detail below with reference to several representative implementations of the present application.

[0043] Reference below Figure 1 , Figure 1 A flow chart of a model fusion method provided in an embodiment of the present application. A model fusion method provided in an embodiment of the present application includes:

[0044] Step S101, identifying a first computing module and a second computing module to be optimized in a neural network model.

[0045] In an embodiment of the present application, the first computing module and the second computing module meet the pre-set optimization conditions. Specifically, the optimization conditions include at least one of the following: First, the computing module performs at least one access operation to the memory space during data processing. This condition refers to the memory read and write operations involved in the data processing of the computing module. When the module is processing data, if it is necessary to store the intermediate results in the memory and then read them from the memory to continue the subsequent calculations, this will increase the memory access overhead. The goal of this conditional trigger optimization is to reduce unnecessary memory reads and writes and optimize memory operations through module fusion, thereby improving execution efficiency. Exemplarily, when a computing module processes a part of the data, it writes the results to the memory, and the next module needs to read these results from the memory during calculation. After optimization, these data can be directly transferred in registers or caches without passing through the memory, reducing memory read and write operations.

[0046] Second, the data transmission process between computing modules includes at least one access operation to the memory space. This condition is mainly for the data transmission process between modules. If data needs to be transferred between two computing modules through memory, that is, the previous module writes data to the memory, and the latter module reads the data from the memory for processing, this will increase the transmission delay and the use of memory bandwidth. Through optimization, data can be directly transmitted between the two modules, avoiding the memory as an intermediary, reducing unnecessary transmission steps, and thus improving data transmission efficiency. For example, the data generated by the first computing module is written to the memory, and the second computing module retrieves the data from the memory for further calculation. After optimization, data can be directly transferred between modules, avoiding multiple accesses to the memory, saving time and resources.

[0047] Third, the computing modules each perform independent discrete operations. This condition indicates that the operations between the computing modules are independent discrete operations, and there is no strong dependence on each other. Since independent operations can be processed in parallel to a certain extent, the computing tasks of multiple modules can be fused and executed in parallel through optimization to avoid waiting states between modules. Through such fusion, computing resources can be better utilized, reducing idle waiting between modules and improving computing efficiency. Exemplarily, two independent computing modules perform their own discrete operations respectively, and originally need to wait for the previous module to be fully executed before performing the operation of the next module. After optimization, the parallel execution of some tasks can be achieved, reducing waiting time.

[0048] It should be noted that the purpose of these three optimization conditions is to maximize the computational efficiency of the neural network by reducing memory reads and writes, optimizing data transmission, and improving the parallel processing capabilities between modules. Each condition provides a specific optimization path for different operating scenarios of the system to ensure that performance improvements can be achieved for different module combinations in actual applications. Here, the above three optimization conditions are only examples. In actual applications, they are not limited to the above three optimization conditions, and other optimization conditions can also be used.

[0049] As an optional embodiment, in S101, first, a computational graph of the neural network is constructed to understand the input, output and dependencies of each module. The computational graph can help identify the data flow between modules and whether there are obvious memory read and write operations or data transmission bottlenecks between modules. For example, for dependencies, you can check which modules must rely on the results of other modules when executing, which helps determine the order of optimization and possible module fusion points.

[0050] Alternatively, you can monitor and analyze the frequency and type of memory accesses of each computing module during model execution. Focus on modules that contain frequent memory read and write operations during execution, which are usually the focus of optimization. In practical applications, you can use performance analysis tools, such as TensorBoard, NVIDIA Nsight, Intel VTune, etc., to monitor the memory usage and bandwidth consumption of the model during execution. Further, analyze whether the module performs independent discrete operations. If the calculation results of a module can be used directly within the module or through more efficient transmission methods without relying on the intermediate results of other modules, then this module may be a candidate for optimization. Identify modules that can be executed in parallel, whose calculations can be performed independently of other modules, so that performance can be improved through operation fusion.

[0051] Alternatively, you can use performance analysis tools to detect computational bottlenecks or performance hotspots in your model. Typically, modules that take longer to compute or use more resources are the focus of optimization. For example, generate heat maps or statistics of computational hotspots to help identify which modules have the greatest impact on overall performance. Experiment with different combinations of modules to simulate the effects of operational fusion. Experiment to verify which module combinations achieve better performance. Use benchmarks to evaluate the performance differences before and after module fusion to help determine the best module combination and optimization strategy.

[0052] Alternatively, you can also analyze the data transmission capabilities and computing resources of the hardware characteristics (such as GPU, TPU, CPU, etc.) that actually run the model. Select modules that can run efficiently on specific hardware for optimization. Identify bottlenecks on the hardware (such as bandwidth limitations, low utilization of computing units, etc.) and adjust the optimization strategy of the modules based on these bottlenecks.

[0053] In addition, it is also possible to determine which modules are suitable for operation fusion according to specific fusion rules (such as frequent data transmission, independent computing tasks, etc.). Thus, by executing module fusion experiments, the performance improvement effect after fusion can be tested.

[0054] In summary, through one or more of the above-mentioned analysis methods, the step of identifying the first computing module and the second computing module to be optimized in S101 can be achieved by comprehensively analyzing multiple factors such as computing graphs, memory operations, computing independence, performance bottlenecks, hardware characteristics, etc., by using performance analysis tools, experimental simulations and considering actual hardware characteristics, effectively determining the target module for optimization, thereby achieving performance improvement of the neural network model.

[0055] Step S102, according to the module fusion method matching the model to be optimized, the first computing module and the second computing module are operated and fused to obtain the target neural network model.

[0056] The module fusion method at least includes: a target connection relationship between the optimized first computing module and the second computing module and a data transmission method.

[0057] In step S102, the first computing module and the second computing module are operated and fused according to the module fusion mode that matches the model to be optimized. Obtaining the target neural network model means operating and fusion of specific computing modules (i.e., the first computing module and the second computing module) in the neural network to achieve a more efficient connection relationship and data transmission mode. The core of this step is to select an appropriate fusion strategy based on the structural characteristics and computing requirements of the model, which specifically includes two aspects: target connection relationship and data transmission mode.

[0058] It can be understood that the target connection relationship refers to the dependency and data flow mode between the optimized first computing module and the second computing module. Through operation fusion, the connection between the two modules can be redesigned to make them more efficient and compact. Specifically, the traditional module connection method may require that the second computing module will not start until the first computing module is fully executed. In the target connection relationship, the computing tasks of the two can be partially overlapped through operation fusion, allowing the second module to start computing when the first module is partially completed. This method reduces the waiting time between modules and improves the utilization of computing resources. Secondly, the optimized connection relationship can directly connect the data flow between modules, avoiding the additional overhead of main memory or intermediate buffering. The output of the first module can be directly used as the input of the second module without the need for storage and re-reading steps. In some cases, the first and second computing modules can be designed as a more tightly integrated module that shares the same computing resources or memory space. This method reduces the delay when data is transferred between the two modules and improves the overall computing efficiency.

[0059] The data transmission mode refers to how to transfer data between the first and second computing modules after the modules are merged. Through optimization, the transmission delay, bandwidth consumption and redundant operations can be reduced. Exemplarily, the following methods can be used to achieve this improvement and optimization. For example, in the optimized data transmission mode, the data between the two modules can be directly transferred through the register without passing through the main memory. This method can greatly reduce the time of data transmission because the access speed of the register is much faster than that of the memory. Alternatively, by transferring data in the cache instead of the main memory, the bandwidth bottleneck caused by frequent access to the memory can be avoided. When the data is transferred between modules, it is preferentially stored in the cache, reducing the frequency of accessing the main memory. Alternatively, through the pipeline design, the data flow of the first module can be quickly transferred to the second module, reducing pauses when processing data. This method is similar to the pipeline structure in hardware design, which can improve parallelism and make data transmission between modules smoother. In addition, in some cases, communication between modules can also be optimized by transferring data in batches. Unlike processing a single data item each time, batch transmission can reduce the total number of data transmissions and improve bandwidth utilization. In summary, by fusing the first computing module with the second computing module in a module fusion method that matches the model to be optimized, the target neural network model is obtained, which can reduce unnecessary memory reading and writing, that is, data can be directly transferred between modules, reducing storage and re-reading operations. It is also possible to improve module parallelism, that is, by optimizing the connection relationship, the first and second modules can be executed in parallel, avoiding the waste of computing resources. In addition, the efficiency of data transmission can be optimized. By using registers or caches for transmission, the delay of memory access and bandwidth consumption are reduced.

[0060] In step S102, the module fusion method determined mainly involves the target connection relationship between the first computing module and the second computing module and the data transmission method. By optimizing these two aspects, the delay and resource waste between computing modules can be effectively reduced, and the overall efficiency and performance of the neural network model can be improved.

[0061] In a specific embodiment, in S102, the first computing module and the second computing module are operated and integrated according to the module integration mode matching the model to be optimized to obtain the step of the target neural network model, such as Figure 2 As shown, including:

[0062] S201, obtaining the module structure, module parameters and connection relationship of the first computing module and the second computing module respectively;

[0063] S202, according to the module structure, module parameters and connection relationship, obtaining a unit to be optimized in the first computing module that matches the second computing module; the unit to be optimized includes at least a logic circuit unit;

[0064] S203: Reconstruct the unit to be optimized to achieve operational fusion between the first computing module and the second computing module.

[0065] In this specific embodiment, the module fusion in step S102 is implemented through a series of steps, including obtaining the structure and parameters of the module, matching the units to be optimized, and reconstructing these units to achieve operation fusion.

[0066] Specifically, in S201, the module structure, module parameters and connection relationship of each of the first computing module and the second computing module are obtained. In this step, the system first needs to understand in detail the internal structure, parameter settings and connection relationship between the first computing module and the second computing module. This information includes the computing logic, storage unit, input and output data flow in each module, and the data transfer method between modules. By analyzing this information, the system can identify the computing dependency, data transmission path and possible performance bottlenecks between the two modules. This is a key step to provide basic information for subsequent steps.

[0067] In this way, by fully understanding the structure and parameters of the module, the parts to be optimized can be identified more accurately, avoiding interference with unnecessary parts. This step provides detailed data support for the subsequent optimization process, ensuring that the operation fusion can proceed smoothly.

[0068] In S202, the units to be optimized in the first computing module that match the second computing module are obtained according to the module structure, module parameters, and connection relationship. After obtaining the detailed information of the modules, the system analyzes which units (such as logic circuit units) in the first computing module can match the units of the second computing module, and determines whether these units have overlapping or redundant computing tasks. Through this matching process, the system can identify specific logical units that can be merged or optimized. These units may involve similar computing operations or data processing processes. Through fusion, repeated operations can be reduced and overall efficiency can be improved.

[0069] Thus, the optimizable units can be accurately identified, avoiding unnecessary adjustments to the entire module and reducing system complexity. By identifying redundant computing units, the waste of computing resources can be significantly reduced and the operating efficiency of the model can be improved.

[0070] In S203, the unit to be optimized is reconstructed to achieve operational fusion between the first computing module and the second computing module. After identifying the units to be optimized, the system will reconstruct these units. This includes redesigning the logic circuit to merge similar or similar computing units of the first module and the second module into a unified structure. This reconstruction process not only includes physical merging (such as sharing registers or cache units), but may also include merging of computing processes to reduce data transmission and waiting time between modules. Through operational fusion, the calculation and data transmission between modules become more compact and efficient, thereby reducing delays and resource overhead.

[0071] In this way, by reducing redundant calculations and data transmission, the optimized modules can perform computing tasks more quickly, and the operating efficiency of the overall model is improved. By fusing logic circuit units and optimizing data flow paths, the waste of hardware resources is reduced, such as reducing frequent access to memory and reducing power consumption. The direct connection and tight integration between modules reduce waiting time and communication delays, especially in highly parallel computing environments, where the optimization effect is more significant. Through module fusion and reconstruction, some computing tasks can be executed in parallel, further improving the throughput of the system. This embodiment demonstrates an effective strategy for optimizing and fusing computing modules in a neural network through structured steps and analysis methods, thereby improving the overall performance and efficiency of the model.

[0072] As an optional implementation, in the above step S203, the operations in the first calculation module and the second calculation module that are logically associated are merged. Further, the logical association includes at least one of the following: having the same processing data, having similar processing logic.

[0073] The same processed data refers to the need to process the same data set or data type in two computing modules. In this case, the two modules may perform similar data preprocessing, format conversion, feature extraction and other operations. Operation fusion combines these repeated operations on the same data into a unified operation unit. For example, if both the first module and the second module need to normalize the input data, the normalization operation can be abstracted into a shared submodule, and the two modules jointly call this submodule to process the data.

[0074] In this way, data only needs to be processed once, and the results can be directly used by subsequent modules, reducing the number of repeated calculations. Thus, by merging repeated operations on the same data, the repeated execution of the same calculation is reduced, thereby reducing the use of computing resources. The steps of data processing are reduced, making the calculation process more concise and efficient. By reducing the redundant storage and transmission of data between different modules, the requirements for memory and bandwidth are reduced.

[0075] It is worth noting that the same processed data may not necessarily have the same role or function in the two modules. For example, the same processed data may be an intermediate calculation result obtained by the first calculation module or a model parameter value used in the second calculation module.

[0076] Similar processing logic means that two computing modules use similar algorithms or logical steps when processing data. Even if their input data is different, the processing algorithm logic is similar. For example, both modules may involve convolution operations, matrix multiplication, or activation function application. Operation fusion identifies and merges these similar logical operations and unifies them into a shared computing path. For example, if both the first module and the second module need to perform convolution operations, the two convolution operations can be merged into a single convolution layer, and the appropriate indexes and parameters can be used to distinguish the data flow.

[0077] The purpose of this is to reduce the repeated calculation nodes in the calculation graph and optimize the complexity and execution efficiency of the calculation graph. Therefore, by sharing similar calculation logic, the occupation of hardware resources (such as GPU, TPU, etc.) is reduced, making the utilization rate of hardware resources higher. Redundant operations in the calculation path are reduced, making the calculation graph more concise and the execution speed faster. By merging logical operations, the structure and calculation graph of the model are simplified, the complexity of the model is reduced, and the maintainability and scalability are improved.

[0078] It is worth noting that, due to the above-mentioned logical association characteristics, the present embodiment does not require the continuity between the first calculation module and the second calculation module. In the present application, the first calculation module and the second calculation module can be continuous or discontinuous.

[0079] The logical association feature refers to the fact that the first computing module and the second computing module improve the overall computing efficiency by fusing operations when processing the same data or using similar processing logic. This logical association does not depend on physical continuity or a fixed execution order. For example, if two modules both perform the same normalization operation on the input data, or both perform convolution operations, even if the two operations are not adjacent in the computational graph, they can still be logically associated and fused.

[0080] That is to say, since the key to operation fusion is to utilize the logical association between modules rather than the physical proximity, the position and order between the first computing module and the second computing module in the computational graph are allowed to be discontinuous. This flexibility means that the optimization between modules is not limited to their physical order, but is more focused on their functional similarities and commonalities in data processing. Therefore, logically related computing modules can share resources or reuse similar computing logic even if they are not directly adjacent in structure. For example, a deep neural network model may use convolution operations in the front layers, and then use similar convolution operations again in a later layer. In this case, although these operations are separated by multiple levels in the computational graph, they still have logical associations and can therefore be fused. In this way, even modules that are physically or temporally discontinuous can be merged and optimized.

[0081] In the related art, it is assumed that the first computing module is module 1 and the second computing module is module 2. It is assumed that module 2 is started to execute its own logic after module 1 is executed. Figure 3 As shown in the data flow, after the module reconstruction in the embodiment of the present application, part of the results of module 1 can be directly input into module 2, thereby reducing the hardware resource call of module 2 and accelerating the execution efficiency of module 2.

[0082] In summary, in the above step S102, by fusing operations with logical associations, the above implementation further improves the optimization effect of the model. Merging the same or similar data processing steps makes computing resources more concentrated and avoids invalid repeated calculations, especially when processing large-scale data. The optimized model achieves more efficient computing performance by reducing computing nodes and redundant operations. Through operation fusion, the number of data transmission and storage is reduced, thereby reducing system latency and bandwidth consumption, and optimizing system power consumption. By reorganizing and fusing computing operations, the parallel computing capabilities of the hardware can be better utilized, further improving the execution speed of the model. The model structure is simpler and more optimized, the code maintenance is easier, and it is more scalable when adding new functional modules in the future.

[0083] Specifically, in an example of the above steps, the operations with logical association in the first computing module and the second computing module are fused, which can be implemented as follows: if the neural network model is a floodlit model, the linear change operation of the multilayer perceptron module in the floodlit model and the addition operation in the residual connection module at the back end of the multilayer perceptron module are determined as the first unit to be optimized. Among them, the linear change operation and the addition operation both have addition operation processing, and the data processing objects involved in the linear change operation and the addition operation both include: linear transformation matrix.

[0084] The floodlight model involved in the above or following examples includes Figure 4 The neural network structure shown. Figure 4 In the figure, from left to right, we can see the input, decoding block, multi-head attention mechanism and fusion mechanism. In the input and embedding layer, the input tokens (Tokens), the input data is first represented as tokens (Token1, Token2, Token3). In the embedding layer (Embed), each input token is converted into a vector representation through the embedding layer (Embed), and these vectors capture the important semantic information of the token. Layer Normalization (LN), the embedded vectors are processed by layer normalization (Layer Normalization, LN) to help stabilize the training process. In the decoding block (Decoder Block), the decoding blocks are stacked, and the model structure includes multiple (for example 70) decoding blocks stacked. Each decoding block includes a multi-head attention mechanism and a multi-layer perceptron (MLP) module. Further, the specific structure of the multi-layer perceptron (MLP) module can be found in Figure 5As shown. Multi-Head Attention, Key-Query Product, calculates the dot product of the key (Key) and query (Query) of the input vector to capture the relationship between information at different positions. MASK (ALiBi Mask), adjusts the position encoding through ALBI Mask to ensure that each attention head can capture the importance of the mark at different positions. Softmax, calculates the weight distribution to determine the attention coefficient between different marks. Weighted Sum of Values, weighted sum of value vectors based on attention distribution. Multi-layer Perceptron (MLP) module, including layer normalization (LN) and fully connected layer (the embedded input is first processed by MLP). Residual connection and layer normalization, through residual connection (ResidualConnection) and layer normalization operations, improve the training stability of the model. In the Multi-Head Attention, the multi-head attention mechanism is multi-path parallel, and the multi-head attention mechanism captures different attention modes through parallel operations. Head Fusion, integrates the output results of multiple attention heads together. In the output layer, the final output is the weighted sum of the fused attention heads to obtain the decoding result. In general, this model uses deeply stacked decoding blocks and an efficient multi-head attention mechanism to capture complex semantic relationships and long-distance dependencies in the input data, thereby achieving high-quality natural language processing tasks.

[0085] Exemplarily, in a panoptic model (such as a variant of the Transformer model), computing resources can be optimized by fusing operations with logical associations. In this specific example, how to fuse the linear change operation in the multilayer perceptron (MLP) module and the addition operation in the residual connection module is discussed. Assume that the multilayer perceptron (MLP) module contains one or more linear change (linear transformation) operations, such as a fully connected layer. This operation can be expressed as: Y = X*W + b; where X is the input feature matrix, W is the weight matrix of the linear transformation, b is the bias vector, and Y is the linear transformation matrix output after the linear transformation. The residual connection module is implemented by directly adding the output of the MLP module to the input of the module, assuming that it is expressed as: Z = Y + X; where Y is the linear transformation matrix output by the MLP module, X is the input of the residual connection (also the input of the MLP), and Z is the output after the residual connection.

[0086] In this example, there is a logical association between the linear change operation (fully connected layer of the MLP module) and the addition operation of the residual connection module, because they both involve addition operations and the processed data objects contain linear transformation matrices. Therefore, we can perform operational fusion on these two operations. Here, the logical association can be understood as the linear change operation contains an addition operation (addition of the bias term in Y = X*W + b), and the residual connection operation contains an addition operation (Z = Y + X), both of which involve addition operations, and the data processing objects contain linear transformation matrices (such as the weight matrix W and the input data matrix X). By merging the linear change operation in the MLP module and the addition operation in the residual connection module into one operation unit, the calculation steps can be reduced. This merger allows two addition steps to be completed in one operation, and the calculation results can be applied to subsequent calculations in the register without storing the intermediate result Y in the memory space.

[0087] In this way, by fusing the linear change operation of the multi-layer perceptron module in the flood model with the addition operation of the residual connection module, the calculation steps and storage requirements can be effectively reduced, and the calculation efficiency can be optimized. This fusion strategy is particularly suitable for application on neural network accelerators, thereby significantly improving the overall performance of the model.

[0088] Exemplarily, in another example of the above steps, the operations with logical association in the first computing module and the second computing module are fused, which can also be implemented as follows: if the neural network model is a large floodlight model, the first linear transformation operation of the multilayer perceptron module in the large floodlight model and the activation function operation at the back end of the multilayer perceptron module are determined as the second unit to be optimized. Here, the first linear transformation operation at least includes part of the linear transformation operation logic before matrix addition. Furthermore, the data processing objects involved in the first linear transformation operation and the activation function operation both include: the linear transformation result calculated by the first linear transformation operation.

[0089] In this example, we will introduce how to perform fusion optimization of logically related operations in the Floodlight model, especially the first linear transformation operation and subsequent activation function operations in the Multilayer Perceptron (MLP) module. In the MLP module, the first linear transformation operation usually refers to the first fully connected layer. It linearly transforms the input features through matrix multiplication (and possible addition of bias terms) to generate a new feature representation. This linear transformation operation can be expressed as: Z = X*W1+b1, where X is the input feature, W1 is the weight matrix, and b1 is the bias. After the linear transformation, a nonlinear activation function (such as ReLU, GELU, etc.) is usually applied to introduce nonlinearity, enabling the model to learn more complex patterns. The input of the activation function is usually the output result of the first linear transformation operation, that is, `Z`.

[0090] In this example, there is a close logical connection between the first linear transformation operation and the activation function operation, because the activation function directly acts on the output of the linear transformation. Therefore, we can optimize the computation of these two operations by fusion. The output Z of the linear transformation operation is directly used as the input of the activation function operation. Therefore, there is a natural computational dependency between them. The objects processed by these operations are the same data - the result Z obtained by the linear transformation.

[0091] Thus, the linear transformation and activation function operations can be integrated into a composite operation. In some hardware implementations or optimization schemes, the activation function can be applied directly while calculating the linear transformation, thereby reducing the number of data transfers in memory and the delay of the calculation. In hardware optimization or software implementation, a fusion calculation unit can be designed to immediately apply the activation function to the result while performing matrix multiplication and addition of linear transformation. For example, A = ReLU (X*W1+b1), this can be completed in one calculation process instead of in two separate steps.

[0092] In this way, by completing the linear transformation and activation function in one step, the computing resources and time overhead of separate operations are reduced. The storage requirements of intermediate results are reduced because it is no longer necessary to store the results of the linear transformation separately before the activation function is applied. By reducing the separation steps between operations, the parallelization capability of the overall execution is improved, especially on hardware accelerators (such as GPUs and TPUs), the computing speed can be significantly improved. By fusing the first linear transformation operation with the activation function operation in the floodlight model, the computing resources can be effectively optimized, the intermediate steps can be reduced, and the execution efficiency of the model can be improved. This optimization strategy is particularly suitable for large neural network models and can significantly improve performance in high-performance computing environments.

[0093] As an optional implementation, in the above step S202, based on the module structure, module parameters and connection relationship, a first unit to be optimized that is associated with at least part of the computing logic in the second computing module is extracted from the first computing module; the first unit to be optimized is used to execute the computing logic to be optimized that contains at least one memory space interaction operation in the first computing module. Furthermore, in the above step S203, the intermediate calculation result obtained by the first unit to be optimized is integrated into the second computing module, and the computing logic to be optimized in the second computing module is reconstructed to reduce the memory space interaction operations in the first computing module and the second computing module.

[0094] In this optional implementation, it mainly involves optimizing the memory space interaction operations between computing modules, thereby improving the overall computing efficiency. First, based on the structure and parameters of the modules and the connection relationship between them, a first unit to be optimized associated with a part of the computing logic in the second computing module is extracted from the first computing module. The unit to be optimized includes at least one memory space interaction operation.

[0095] Here, memory space interaction refers to operations that require intermediate results to be stored in memory during the computation process and accessed again in subsequent steps. For example, intermediate results are output from one module and stored in memory, and then another module reads these intermediate results from memory.

[0096] In step S203, the intermediate calculation results generated by the first unit to be optimized are directly integrated into the second calculation module, without being transferred through memory interaction operations. This integration is achieved by redesigning the calculation logic in the second module, with the aim of directly utilizing the output results of the first module and reducing the number of memory accesses. At the same time as the integration, the calculation logic of the second calculation module will be rebuilt to adapt to the intermediate results from the first unit to be optimized. This step ensures that the dependency between the first module and the second module becomes closer, thereby eliminating unnecessary memory interactions.

[0097] It is understandable that the intermediate results between the first computing module and the second computing module need to be stored in the memory and accessed in the subsequent stage. By integrating this memory interaction, data can be directly transferred in the computing unit, reducing the dependence on memory bandwidth. Memory access is one of the bottlenecks of computing performance, especially in high-performance computing, reducing memory access can significantly speed up the calculation speed. By reducing the memory interaction between modules, the overall computing efficiency is greatly improved. Each memory interaction will introduce delays and increase energy consumption. By reducing this interaction, not only the computing delay is reduced, but also the consumption of computing resources can be effectively reduced, and the energy efficiency ratio of the system is improved. By directly integrating part of the logic of the first computing module into the second computing module, these modules can be better executed in parallel, which is especially suitable for hardware acceleration scenarios (such as GPU, TPU), further improving the parallel processing capability of the model.

[0098] This optimization can effectively improve the computing speed, reduce storage overhead, and optimize resource utilization by extracting relevant computing logic and reducing memory interactions between modules, thereby significantly improving the execution performance of the model. This strategy is very suitable for the optimization of large neural networks, especially in scenarios that require high performance and high energy efficiency.

[0099] Exemplarily, in the above step S202, the intermediate calculation result obtained by the first unit to be optimized is integrated into the second calculation module, and the calculation logic to be optimized in the second calculation module is reconstructed, which can be implemented as follows:

[0100] In the first computing module, a new port is set for outputting the intermediate calculation result; the new port is directly connected to the second computing module; through the new port, the intermediate calculation result is input into the second computing module, and the second computing module applies the intermediate calculation result to implement its own computing logic, so as to reduce the storage operation of the first computing module on the memory space and the reading operation of the second computing module on the content space.

[0101] The above steps add a port in the first computing module for outputting the intermediate calculation results, and the port is directly connected to the second computing module. In this way, the intermediate calculation results of the first computing module can be directly input to the second computing module without going through the memory storage operation, and the latter applies the intermediate results to its own calculation logic. This design reduces the need for the first computing module to store the results in the memory and the operation of the second computing module to read data from the memory, thereby optimizing memory usage and improving computing efficiency.

[0102] The above steps are described below using two different examples.

[0103] In one case, if the neural network model is a flood model, and the linear change operation of the multilayer perceptron module and the addition operation in the residual connection module are used as the first unit to be optimized, then, in the above steps, the first calculation module is provided with a new port for outputting the intermediate calculation result, including: configuring the first new port in the multilayer perceptron module, for directly outputting the linear change matrix obtained by the linear change operation of the multilayer perceptron module from the register. Furthermore, the intermediate calculation result is input into the second calculation module through the first new port, and the second calculation module applies the intermediate calculation result to realize its own calculation logic, including: inputting the linear change matrix into the residual connection module through the first new port, and applying the linear change matrix to the addition operation of the residual connection module, so as to realize the fusion of the linear change operation and the residual connection module.

[0104] In this case, it is assumed that the neural network model is a flood model, and the linear change operation of the multilayer perceptron (MLP) module and the addition operation in the residual connection module are fused as the units to be optimized. A first newly added port is configured in the multilayer perceptron module for directly outputting the matrix after the linear change operation (linear transformation) from the register. This means that after the addition operation after the linear transformation (adding the matrix to the bias matrix) is completed, the intermediate result is no longer stored in the memory, but the result is directly output through this port. Through the first newly added port, the intermediate calculation result after the linear change is directly transmitted to the second calculation module, that is, the residual connection module, to avoid the intermediate result being stored in the memory. In the second calculation module, the residual connection module receives the linear change matrix from the first newly added port and directly applies it to the addition operation of the residual connection. At this time, the linear change matrix and the input of the residual connection module are added to complete the residual connection operation.

[0105] This method ensures that after the linear transformation operation is completed, the result is still stored in the register and immediately enters the addition operation of the residual connection module. Through this operation fusion, the operation of writing and reading intermediate results from memory is successfully reduced, improving computing efficiency and resource utilization. Through the design and operation fusion of this new port, memory interaction operations can be effectively reduced, latency and energy consumption can be reduced, while improving computing parallelism and achieving more efficient model computing performance.

[0106] In another case, if the neural network model is a floodlight model, and the first linear transformation operation of the multilayer perceptron module and the activation function operation at the back end of the multilayer perceptron module are used as the second unit to be optimized, then in the first calculation module, a new port for outputting the intermediate calculation result is set, including: configuring a second new port in the multilayer perceptron module, for directly outputting from the register a plurality of linear transformation results obtained by the multilayer perceptron module after the first linear transformation operation; the first linear transformation operation at least includes part of the linear transformation operation logic before matrix addition. Further, through the first new port, the intermediate calculation result is input into the second calculation module, and the second calculation module applies the intermediate calculation result to realize its own calculation logic, including: through the second new port, the plurality of linear transformation results are respectively input into the activation function operation module, and the plurality of linear transformation results are respectively applied to the activation function operation to realize the activation function calculation, so as to realize the operation fusion of the first linear transformation operation and the activation function operation.

[0107] In another case, it is assumed that the neural network model is a flood model, and the first linear transformation operation and the subsequent activation function operation of the multilayer perceptron (MLP) module are used as the unit to be optimized. A second newly added port is configured in the multilayer perceptron module to directly output multiple linear transformation results obtained by the first linear transformation operation from the register. This operation includes some linear transformation logic, such as operations before matrix multiplication and bias matrix addition. Through the second newly added port, the intermediate results of these linear transformations are directly transferred from the register to the second calculation module, that is, the activation function operation module, to avoid the results being stored in the memory and then read. In the second calculation module, after receiving multiple linear transformation results, the activation function module immediately performs activation function (such as GeLU) calculation operations on these results. Since the activation function operates on a single value, the linear transformation results are directly transmitted from the register, and the activation function calculation can be started directly, reducing the read and write operations of data in the memory. By fusing the first linear transformation operation with the activation function operation, the intermediate results can be avoided from being written to the memory, and the results can be kept in the register. The results after the linear transformation do not need to be written or read from the memory, and directly participate in the activation function calculation. This greatly reduces memory interaction and improves computing efficiency and speed.

[0108] This operation fusion method performs activation function calculation immediately after the first linear transformation operation is completed, without the need to write and read intermediate results into and from memory. By directly using the data in the register, it can reduce latency in the calculation process, optimize memory utilization, and improve the overall performance of the model.

[0109] Further optionally, in step S102, assembly instructions can be used to perform operations on the hardware resource calling methods in the first computing module and the second computing module according to the module fusion method to obtain the target neural network model; the hardware resources at least include memory space resources.

[0110] Specifically, the core of module fusion is to integrate the operations of multiple computing modules to reduce redundant computing steps and resource occupation. Specifically, the intermediate results between the first computing module and the second computing module no longer need to be transferred through memory, but are directly transferred between computing units (such as registers). This method reduces the dependence on memory resources, thereby improving computing efficiency. Assembly instructions have the ability to directly manipulate underlying hardware resources, such as registers, memory, and CPU. Assembly instructions can be used to more accurately control the allocation and call of computing resources. When implementing operation fusion, assembly instructions can finely manage the use of registers, save the results of the first computing module directly in the register, and then immediately perform further operations in the second computing module without writing and reading data from the memory. Through assembly instructions, the calling method of hardware resources can be optimized so that multiple computing operations can be executed in parallel or in an efficient sequence. For example, after the first computing module completes the calculation, the assembly instruction can directly trigger the operation of the second computing module without waiting for the data storage and reading process. This method maximizes the computing power of registers and CPUs and reduces the load on memory.

[0111] In this way, operation fusion is performed through assembly instructions, which avoids redundant memory operations and makes data transfer between computing modules faster and more direct. This significantly improves computing efficiency, especially when processing large amounts of data and complex models, which can effectively shorten computing time. The application of assembly instructions reduces the occupancy of memory space, because intermediate results can be directly transferred in registers and further used. This not only reduces the frequency of memory reading and writing, but also reduces the pressure on memory bandwidth. By finely controlling the calling order and method of hardware resources, hardware resources such as CPU and registers can be used more efficiently. Multiple computing operations can be performed in parallel without conflict, thereby further improving the computing performance of the overall system. In general, the use of assembly instructions for operation fusion can maximize the use of underlying hardware resources and significantly improve the computing efficiency and performance of neural network models.

[0112] Step S103, execute the target neural network model to achieve hardware acceleration of the target neural network model.

[0113] Step S103 involves executing the target neural network model and optimizing its performance through hardware acceleration. Specifically, further optionally, before S103, a dedicated hardware accelerator, such as a GPU (graphics processing unit), a TPU (tensor processing unit) or an FPGA (field programmable gate array), is configured. The layout and allocation of hardware resources are determined to ensure that key components (such as memory, cache, and registers) can efficiently serve the needs of the neural network model. The neural network model is compiled into a low-level instruction set that can be executed by the hardware accelerator, such as CUDA (for NVIDIA GPU) or Vulkan. Hardware-specific optimizations are performed, such as memory alignment, compute intensity optimization, and reducing data transmission delays. Using operation fusion technology, several computing operations (such as matrix multiplication and activation functions) are combined into larger and more complex operations to reduce the storage and reading of intermediate data. Through the pipeline method, the model is decomposed into multiple independent computing stages so that each stage can run in parallel to improve overall efficiency. Registers are used to cache intermediate data to reduce frequent access to memory during the calculation process. Accelerate matrix operations and activation function calculations through hardware-specific instruction sets (such as SIMD instruction sets). For the multi-head self-attention mechanism, the calculation of query, key and value vectors is implemented in parallel processing hardware such as GPU, and the parallel computing capabilities of the hardware are used to increase the speed. Use special hardware features such as TensorCores to speed up matrix multiplication and other linear algebra operations. Implement data pipelining and caching strategies to ensure that data can be transmitted between computing units at high speed without forming bottlenecks. Store hot data and intermediate results in cache to increase access speed.

[0114] During the actual execution, the input data is converted into a format suitable for hardware processing, such as vectorizing image or text features and then loading them into GPU memory. The neural network model is initialized and the necessary GPU / TPU resources are allocated. Model inference is performed according to the compiled optimization instructions to perform feature extraction, attention mechanism calculation and final prediction. The parallel computing capability of the GPU is used to speed up the execution of matrix multiplication, convolution operations and activation functions. The hardware acceleration results are transferred from the GPU memory to the CPU for post-processing and output in a format suitable for application requirements (such as conversion into analysis reports, image outputs or text responses).

[0115] Through the above steps, based on the target neural network model after operation fusion, the hardware acceleration optimization of the complete execution process of the neural network model is realized, which significantly improves the speed and performance of model reasoning. Through parallel computing and specific hardware optimization, the operation speed of each layer of the neural network is greatly accelerated.

[0116] In an embodiment of the present application, first, the first computing module and the second computing module to be optimized in the neural network model are identified; the first computing module and the second computing module meet the pre-set optimization conditions. Secondly, according to the module fusion method that matches the model to be optimized, the first computing module and the second computing module are operationally fused to obtain a target neural network model; the module fusion method at least includes: the target connection relationship and the data transmission method between the optimized first computing module and the second computing module. Finally, the target neural network model is executed to achieve hardware acceleration of the target neural network model. In the implementation mode of the present application, by operationally fusion of the first computing module and the second computing module in the neural network model, the computing resource usage is reduced, the overall performance of the model is significantly improved, and the model execution efficiency is improved.

[0117] After introducing the method of the exemplary embodiment of the present application, next, refer to Figure 6 A model fusion device according to an exemplary embodiment of the present application is described, and the device includes the following modules:

[0118] An identification module, used to identify a first calculation module and a second calculation module to be optimized in the neural network model; the first calculation module and the second calculation module meet the preset optimization conditions;

[0119] A fusion module is used to operate and fuse the first computing module and the second computing module according to a module fusion method that matches the model to be optimized, so as to obtain a target neural network model; the module fusion method at least includes: a target connection relationship and a data transmission method between the optimized first computing module and the second computing module;

[0120] The execution module is used to execute the target neural network model and realize hardware acceleration of the target neural network model.

[0121] As an optional implementation, the optimization conditions include at least: the data processing process performed by the computing module includes at least one access operation to the memory space; and / or the data transmission process between the computing modules includes at least one access operation to the memory space; and / or the computing modules each perform independent discrete operations.

[0122] As an optional implementation, the fusion module is specifically used to: obtain the module structure, module parameters and connection relationship of each of the first calculation module and the second calculation module;

[0123] According to the module structure, module parameters and connection relationship, obtaining a unit to be optimized in the first computing module that matches the second computing module; the unit to be optimized includes at least a logic circuit unit;

[0124] The unit to be optimized is reconstructed to achieve operational fusion between the first computing module and the second computing module.

[0125] As an optional implementation, the fusion module reconstructs the unit to be optimized and is configured to: perform operational fusion on the operations with logical association in the first computing module and the second computing module; the logical association includes at least one of the following: having the same processing data and having similar processing logic.

[0126] As an optional implementation, the fusion module performs an operation fusion on the operations having logical association in the first computing module and the second computing module, specifically for:

[0127] If the neural network model is a floodlight model, determine a linear change operation of a multilayer perceptron module in the floodlight model and an addition operation in a residual connection module at the back end of the multilayer perceptron module as the first unit to be optimized;

[0128] Among them, the linear change operation and the addition operation both have addition operation processing, and the data processing objects involved in the linear change operation and the addition operation both include: a linear transformation matrix.

[0129] As an optional implementation, the fusion module performs an operation fusion on the first linear transformation layer and the activation function in the multi-layer perceptron module and is configured as follows:

[0130] If the neural network model is a floodlight model, determine the first linear transformation operation of the multilayer perceptron module in the floodlight model and the activation function operation at the back end of the multilayer perceptron module as the second unit to be optimized; the first linear transformation operation at least includes part of the linear transformation operation logic before matrix addition;

[0131] The data processing objects involved in the first linear transformation operation and the activation function operation both include: a linear transformation result calculated by the first linear transformation operation.

[0132] As an optional implementation, the fusion module obtains the unit to be optimized in the first computing module that matches the second computing module according to the module structure, module parameters and connection relationship, and is configured as follows:

[0133] Extracting a first unit to be optimized associated with at least part of the computing logic in the second computing module from the first computing module according to the module structure, module parameters and connection relationship; the first unit to be optimized is used to execute the computing logic to be optimized in the first computing module that includes at least one memory space interaction operation;

[0134] The fusion module reconstructs the unit to be optimized to achieve the operation fusion between the first computing module and the second computing module, and is configured as follows:

[0135] The intermediate calculation results obtained by the first unit to be optimized are integrated into the second calculation module, and the calculation logic to be optimized in the second calculation module is reconstructed to reduce the memory space interaction operations between the first calculation module and the second calculation module.

[0136] As an optional implementation, the fusion module, the intermediate calculation result obtained by the first unit to be optimized is merged into the second calculation module, and the calculation logic to be optimized in the second calculation module is reconstructed, and is configured as follows: in the first calculation module, a new port for outputting the intermediate calculation result is set; the new port is directly connected to the second calculation module;

[0137] The intermediate calculation result is input into the second calculation module through the newly added port, and the second calculation module applies the intermediate calculation result to implement its own calculation logic, so as to reduce the storage operation of the first calculation module on the memory space and the reading operation of the second calculation module on the content space.

[0138] As an optional implementation, if the neural network model is a flood model, and the linear change operation of the multilayer perceptron module and the addition operation in the residual connection module are used as the first unit to be optimized, then

[0139] The fusion module, in the first calculation module, is provided with a newly added port for outputting the intermediate calculation result, specifically for:

[0140] A first newly added port is configured in the multi-layer perceptron module, for directly outputting a linear change matrix obtained by the multi-layer perceptron module through a linear change operation from a register;

[0141] The fusion module inputs the intermediate calculation result into the second calculation module through the first newly added port, and the second calculation module applies the intermediate calculation result to implement its own calculation logic, specifically for:

[0142] The linear change matrix is ​​input into the residual connection module through the first newly added port, and the linear change matrix is ​​applied to the addition operation of the residual connection module to achieve the fusion of the linear change operation and the operation of the residual connection module.

[0143] As an optional implementation, if the neural network model is a flood model, and the first linear transformation operation of the multilayer perceptron module and the activation function operation at the back end of the multilayer perceptron module are used as the second unit to be optimized, then

[0144] The fusion module, in the first calculation module, is provided with a newly added port for outputting the intermediate calculation result, specifically for:

[0145] A second newly added port is configured in the multilayer perceptron module, for directly outputting from the register a plurality of linear transformation results obtained by the multilayer perceptron module after a first linear transformation operation; the first linear transformation operation at least includes a partial linear transformation operation logic before matrix addition;

[0146] The fusion module inputs the intermediate calculation result into the second calculation module through the first newly added port, and the second calculation module applies the intermediate calculation result to implement its own calculation logic, specifically for:

[0147] Through the second newly added port, multiple linear transformation results are respectively input into the activation function operation module, and the multiple linear transformation results are respectively applied to the activation function operation to realize the activation function calculation, so as to realize the operation fusion of the first linear transformation operation and the activation function operation.

[0148] As an optional implementation, the fusion module is specifically used for:

[0149] Assembly instructions are used to perform operations on the hardware resource calling methods in the first computing module and the second computing module according to the module fusion method to obtain the target neural network model; the hardware resources at least include memory space resources.

[0150] In an embodiment of the present application, through the model fusion device, the first computing module and the second computing module in the neural network model are operated and fused, thereby reducing the computing resource usage, significantly improving the overall performance of the model, and improving the model execution efficiency.

[0151] After introducing the devices, methods and apparatus of the exemplary embodiments of the present application, next, reference is made to Figure 7 The computer-readable storage medium of the exemplary embodiment of the present application is described. The computer-readable storage medium can be set in the model fusion device. Please refer to Figure 7 , the computer-readable storage medium shown is a CD 110, on which a computer program (i.e., a program product) is stored. When the computer program is executed by the processor, it will implement the steps described in the above method implementation, for example, identifying the first computing module and the second computing module to be optimized in the neural network model; the first computing module and the second computing module meet the preset optimization conditions; according to the module fusion method matching the model to be optimized, the first computing module and the second computing module are operated and fused to obtain the target neural network model; the module fusion method at least includes: the target connection relationship and data transmission method between the optimized first computing module and the second computing module; executing the target neural network model to achieve hardware acceleration of the target neural network model. The specific implementation method of each step will not be repeated here.

[0152] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical or magnetic storage media, which are not listed here one by one.

[0153] After introducing the devices, methods, media and apparatus of the exemplary embodiments of the present application, reference is now made to Figure 8 For the computing device for model fusion according to the exemplary embodiment of the present application, the computing device can be arranged in a model fusion device.

[0154] Figure 8 A block diagram of an exemplary computing device 120 suitable for implementing embodiments of the present application is shown. The computing device 120 may be a computer system or a server. Figure 8 The computing device 120 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0155] like Figure 8 As shown, the components of computing device 120 may include, but are not limited to: one or more processors or processing units 1201, a system memory 1202, and a bus 1203 connecting different system components (including system memory 1202 and processing unit 1201).

[0156] The computing device 120 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the computing device 120, including volatile and nonvolatile media, removable and non-removable media.

[0157] System memory 1202 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 12021 and / or cache memory 12022. Computing device 120 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, ROM 12023 may be used to read and write non-removable, non-volatile magnetic media ( Figure 8 is not shown in the Figure 8As shown in FIG. 1 , a disk drive for reading and writing a removable non-volatile disk (e.g., a "floppy disk") and an optical disk drive for reading and writing a removable non-volatile optical disk (e.g., a CD-ROM, a DVD-ROM, or other optical media) can be provided. In these cases, each drive can be connected to bus 1203 via one or more data medium interfaces. System memory 1202 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of each embodiment of the present application.

[0158] A program / utility 12025 having a set (at least one) of program modules 12024 may be stored, for example, in system memory 1202, and such program modules 12024 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment. Program modules 12024 generally perform the functions and / or methods of the embodiments described herein.

[0159] The computing device 120 may also communicate with one or more external devices 1204 (e.g., a keyboard, a pointing device, a display, etc.). Such communication may be performed via an input / output (I / O) interface 1205. In addition, the computing device 120 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 1206. Figure 8 As shown, the network adapter 1206 communicates with other modules (such as the processing unit 1201, etc.) of the computing device 120 via the bus 1203. It should be understood that although Figure 8 Not shown, other hardware and / or software modules may be used in conjunction with computing device 120 .

[0160] The processing unit 1201 executes various functional applications and data processing by running the program stored in the system memory 1202, for example, identifying the first computing module and the second computing module to be optimized in the neural network model; the first computing module and the second computing module meet the preset optimization conditions; according to the module fusion method matching the model to be optimized, the first computing module and the second computing module are operated and fused to obtain the target neural network model; the module fusion method at least includes: the target connection relationship and data transmission method between the optimized first computing module and the second computing module; execute the target neural network model to achieve hardware acceleration of the target neural network model. The specific implementation method of each step is not repeated here. It should be noted that although several units / modules or sub-units / sub-modules of the model fusion device are mentioned in the above detailed description, this division is only exemplary and not mandatory. In fact, according to the implementation method of the present application, the features and functions of two or more units / modules described above can be concretized in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided into multiple units / modules to be concretized.

[0161] In particular, according to an embodiment of the present disclosure, the process described with reference to the flowchart above may be implemented as a computer program product, which includes: a computer program, which implements the above model fusion method when executed by a processor.

[0162] In the description of the present application, it should be noted that the terms "first", "second" and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0163] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0164] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.

[0165] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0166] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0167] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium that is executable by a processor. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0168] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present application, which are used to illustrate the technical solutions of the present application, rather than to limit them. The protection scope of the present application is not limited thereto. Although the present application is described in detail with reference to the above-described embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-described embodiments within the technical scope disclosed in the present application, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

[0169] In addition, although the operations of the method of the present application are described in a specific order in the drawings, this does not require or imply that the operations must be performed in this specific order, or that all the operations shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.

Claims

1. A model fusion method, characterized in that: The method comprises: Identify a first computing module and a second computing module to be optimized in the neural network model; the first computing module and the second computing module meet a preset optimization condition; According to the module fusion method matching the model to be optimized, the first computing module and the second computing module are operated and fused to obtain a target neural network model; the module fusion method at least includes: a target connection relationship and a data transmission method between the optimized first computing module and the second computing module; Execute the target neural network model to achieve hardware acceleration of the target neural network model; The method of operating and fusing the first computing module and the second computing module according to the module fusion mode matching the model to be optimized to obtain the target neural network model includes: Obtain the module structure, module parameters and connection relationship of each of the first computing module and the second computing module; extract the first unit to be optimized associated with at least part of the computing logic in the second computing module from the first computing module according to the module structure, module parameters and connection relationship; the first unit to be optimized is used to execute the computing logic to be optimized in the first computing module that contains at least one memory space interaction operation; merge the intermediate calculation result obtained by the first unit to be optimized into the second computing module, and reconstruct the computing logic to be optimized in the second computing module to reduce the memory space interaction operations in the first computing module and the second computing module; Among them, if the neural network model is a large flood model, and the first linear transformation operation of the multilayer perceptron module and the activation function operation at the back end of the multilayer perceptron module are used as the second unit to be optimized, a second newly added port is configured in the multilayer perceptron module to directly output from the register a plurality of linear transformation results obtained by the multilayer perceptron module after the first linear transformation operation; the first linear transformation operation at least includes part of the linear transformation operation logic before matrix addition; through the second newly added port, the plurality of linear transformation results are respectively input into the activation function operation module, and the plurality of linear transformation results are respectively applied to the activation function operation to realize the activation function calculation, so as to realize the operational fusion of the first linear transformation operation and the activation function operation.

2. The model fusion method according to claim 1, characterized in that: The optimization conditions include at least: the data processing process performed by the computing module includes at least one access operation to the memory space; and / or the data transmission process between the computing modules includes at least one access operation to the memory space; and / or the computing modules each perform independent discrete operations.

3. The model fusion method according to claim 1, characterized in that: The method of operating and fusing the first computing module and the second computing module according to the module fusion mode matching the model to be optimized to obtain the target neural network model includes: Obtaining the module structure, module parameters and connection relationship of the first computing module and the second computing module respectively; According to the module structure, module parameters and connection relationship, obtaining a unit to be optimized in the first computing module that matches the second computing module; the unit to be optimized includes at least a logic circuit unit; The unit to be optimized is reconstructed to achieve operational fusion between the first computing module and the second computing module.

4. The model fusion method according to claim 3, characterized in that: The reconstructing the unit to be optimized includes: The operations having logical association in the first computing module and the second computing module are merged; the logical association includes at least one of the following: having the same processing data and having similar processing logic.

5. The model fusion method according to claim 4, characterized in that: The step of fusing the operations having logical association in the first computing module and the second computing module comprises: If the neural network model is a floodlight model, determine a linear change operation of a multilayer perceptron module in the floodlight model and an addition operation in a residual connection module at the back end of the multilayer perceptron module as the first unit to be optimized; Among them, the linear change operation and the addition operation both have addition operation processing, and the data processing objects involved in the linear change operation and the addition operation both include: a linear transformation matrix.

6. The model fusion method according to claim 4, characterized in that: The step of fusing the operations having logical association in the first computing module and the second computing module comprises: If the neural network model is a floodlight model, determine the first linear transformation operation of the multilayer perceptron module in the floodlight model and the activation function operation at the back end of the multilayer perceptron module as the second unit to be optimized; the first linear transformation operation at least includes part of the linear transformation operation logic before matrix addition; The data processing objects involved in the first linear transformation operation and the activation function operation both include: a linear transformation result calculated by the first linear transformation operation.

7. The model fusion method according to claim 1, characterized in that: The step of fusing the intermediate calculation result obtained by the first unit to be optimized into the second calculation module and reconstructing the calculation logic to be optimized in the second calculation module includes: In the first calculation module, a new port for outputting the intermediate calculation result is provided; the new port is directly connected to the second calculation module; The intermediate calculation result is input into the second calculation module through the newly added port, and the second calculation module uses the intermediate calculation result to implement its own calculation logic, so as to reduce the storage operation of the first calculation module on the memory space and the reading operation of the second calculation module on the content space.

8. The model fusion method according to claim 7, characterized in that: If the neural network model is a flood model, and the linear change operation of the multilayer perceptron module and the addition operation in the residual connection module are used as the first unit to be optimized, then in the first calculation module, a new port for outputting the intermediate calculation result is set, including: A first newly added port is configured in the multi-layer perceptron module, for directly outputting a linear change matrix obtained by the multi-layer perceptron module through a linear change operation from a register; The step of inputting the intermediate calculation result into the second calculation module through the first newly added port, and the second calculation module using the intermediate calculation result to implement its own calculation logic includes: The linear change matrix is ​​input into the residual connection module through the first newly added port, and the linear change matrix is ​​applied to the addition operation of the residual connection module to achieve the fusion of the linear change operation and the operation of the residual connection module.

9. The model fusion method according to any one of claims 1 to 8, characterized in that: The method of operating and fusing the first computing module and the second computing module according to the module fusion mode matching the model to be optimized to obtain the target neural network model includes: Assembly instructions are used to perform operations on the hardware resource calling methods in the first computing module and the second computing module according to the module fusion method to obtain the target neural network model; the hardware resources at least include memory space resources.

10. A model fusion device, characterized in that: The device at least comprises: An identification module, used to identify a first calculation module and a second calculation module to be optimized in the neural network model; the first calculation module and the second calculation module meet the preset optimization conditions; A fusion module is used to operate and fuse the first computing module and the second computing module according to a module fusion method that matches the model to be optimized, so as to obtain a target neural network model; the module fusion method at least includes: a target connection relationship and a data transmission method between the optimized first computing module and the second computing module; An execution module, used for executing the target neural network model to realize hardware acceleration of the target neural network model; The fusion module, according to the module fusion mode matching the model to be optimized, operates and fuses the first computing module and the second computing module to obtain the target neural network model, is specifically used to: Obtain the module structure, module parameters and connection relationship of each of the first computing module and the second computing module; extract the first unit to be optimized associated with at least part of the computing logic in the second computing module from the first computing module according to the module structure, module parameters and connection relationship; the first unit to be optimized is used to execute the computing logic to be optimized in the first computing module that contains at least one memory space interaction operation; merge the intermediate calculation result obtained by the first unit to be optimized into the second computing module, and reconstruct the computing logic to be optimized in the second computing module to reduce the memory space interaction operations in the first computing module and the second computing module; Among them, the fusion module, when integrating the intermediate calculation result obtained by the first unit to be optimized into the second calculation module and reconstructing the calculation logic to be optimized in the second calculation module, is specifically used for: if the neural network model is a large floodlight model, and the first linear transformation operation of the multilayer perceptron module and the activation function operation at the back end of the multilayer perceptron module are used as the second unit to be optimized, a second newly added port is configured in the multilayer perceptron module for directly outputting from the register a plurality of linear transformation results obtained by the multilayer perceptron module after the first linear transformation operation; the first linear transformation operation at least includes part of the linear transformation operation logic before matrix addition; through the second newly added port, the plurality of linear transformation results are respectively input into the activation function operation module, and the plurality of linear transformation results are respectively applied to the activation function operation to realize the activation function calculation, so as to realize the fusion of the first linear transformation operation and the activation function operation.

11. A computing device, characterized in that: The computing device comprises: at least one processor, memory, and input-output unit; The memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the model fusion method as described in any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that: The method comprises instructions which, when executed on a computer, cause the computer to execute the model fusion method as described in any one of claims 1 to 9.

13. A computer program product, characterized in that It comprises a computer program, which, when executed by a processor, implements the model fusion method as described in any one of claims 1 to 9.

14. A chip, characterized in that: The chip includes a processor coupled to a transceiver, and is used to execute the model fusion method as described in any one of claims 1 to 9.

15. A chip system, characterized in that: The chip system includes: A communication interface for inputting and / or outputting information; A processor is used to execute a computer executable program so that a device equipped with the chip system executes the model fusion method as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Operator fusion method for neural network and related device

    CN118171683A