Method, device, equipment and storage medium for deploying machine learning model
Patent Information
- Application Number
- CN202311281626.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-28
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-09-28
Smart Images

Figure CN117291280B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of this disclosure relate to a pipelined awareness deployment method, apparatus, device, and non-transitory computer-readable medium for in-memory computing systems. Background Technology
[0002] With the remarkable achievements of deep learning in many real-world tasks, such as natural language processing and image generation, it is gradually changing people's daily work and lives. However, highly effective deep learning models require powerful computing capabilities. Traditional computing devices based on the von Neumann architecture, such as CPUs and GPUs, face the problem of separating storage and computing units, making it impossible to reduce energy consumption and thus increasing the cost of inference for deep learning models. In-memory computing systems can achieve higher energy efficiency in many deep learning inference tasks compared to traditional computing systems and have been widely researched both domestically and internationally in recent years, showing promise for application in many terminal devices. However, as the structure of deep learning models becomes increasingly diverse, manually deploying algorithms to in-memory computing systems will incur significant costs. Summary of the Invention
[0003] This disclosure provides methods, apparatus, devices, and non-transitory computer-readable media for deploying machine learning models in a computer system, which can fully utilize hardware features to optimize the deployment of machine learning models and thereby improve computing performance.
[0004] For example, at least one embodiment of this disclosure provides a method for deploying a machine learning model in a computer system. The method includes: extracting multiple computational tasks of the machine learning model; dividing multiple computing units of the computer system into multiple sets of computing unit groups based on the hardware characteristics of the computer system, wherein the multiple sets of computing unit groups have different computational performance; sorting the multiple computational tasks of the machine learning model based on the amount of computation; for each of the sorted multiple computational tasks, selecting one of one or more computing unit groups with the required computational performance from one or more available computing unit groups; and mapping the computational task to a computing unit in the selected computing unit group.
[0005] For example, in a method for deploying a machine learning model in a computer system according to at least one embodiment of the present disclosure, the process of grouping multiple computing units of the computer system into multiple sets of computing unit groups based on the hardware characteristics of the computer system includes: dividing multiple computing unit groups of the computer system into multiple sets of computing unit groups with different computing performance based on the number of one or more computing units included in the computing unit group and the access relationship between one or more computing units included in the computing unit group and the storage unit of the computer system.
[0006] For example, in a method for deploying a machine learning model in a computer system according to at least one embodiment of the present disclosure, the plurality of computing unit groups include a first computing unit group, a second computing unit group, and a third computing unit group, wherein the computing unit group in the first computing unit group includes a plurality of first computing units and the plurality of first computing units are configured to access different first storage units simultaneously, the computing unit group in the second computing unit group includes a plurality of second computing units and the plurality of second computing units are configured to access the same second storage unit in a time-sharing manner, and the computing unit group in the third computing unit group includes only independent third computing units and the third computing units are configured to access third storage units.
[0007] For example, in a method for deploying a machine learning model in a computer system according to at least one embodiment of the present disclosure, the computer system is a memory computing system and the computing unit is a memory computing unit, wherein the memory computing unit includes at least one memory computing array and peripheral driving circuitry for the at least one memory computing array.
[0008] For example, in a method for deploying a machine learning model in a computer system according to at least one embodiment of the present disclosure, the machine learning model is a deep learning model.
[0009] For example, in a method for deploying a machine learning model in a computer system according to at least one embodiment of the present disclosure, each of a plurality of computational tasks includes computation between adjacent network layers of a deep learning model, and the plurality of computational tasks of the machine learning model are ordered based on computational cost, including: ordering the plurality of computational tasks of the deep learning model according to the computational cost between adjacent network layers of the deep learning model, wherein the computational cost is the number of matrix multiplications required to complete one feature map computation of the deep learning model.
[0010] For example, in a method for deploying a machine learning model in a computer system according to at least one embodiment of the present disclosure, each of a plurality of computing unit groups includes a first computing unit and a second computing unit, and wherein mapping a computing task to a computing unit in a selected computing unit group includes: mapping the computation of a first network layer included in the computing task of a deep learning model to a first computing unit in the selected computing unit group, and mapping the computation of a second network layer included in the computing task of a deep learning model to a second computing unit in the selected computing unit group, wherein the second network layer performs computation based on the output of the first network layer, and the second computing unit performs computation based on the output of the first computing unit.
[0011] For example, in a method for deploying a machine learning model in a computer system according to at least one embodiment of the present disclosure, extracting multiple computational tasks of the machine learning model includes: extracting multiple computational layers of a deep learning model for computation in a storage unit and parameters corresponding to the multiple computational layers based on parameter analysis.
[0012] For example, in a method for deploying a machine learning model in a computer system according to at least one embodiment of the present disclosure, the method further includes: expanding the weights of the extracted multiple computing layers into weights for computation by the in-memory unit, and splitting the weights based on the size of the in-memory unit.
[0013] For example, at least one embodiment of this disclosure provides an apparatus for deploying a machine learning model in a computer system. The apparatus includes: a parsing module for extracting multiple computational tasks of the machine learning model; a grouping module for dividing multiple computational units of the computer system into multiple sets of computational unit groups based on the hardware characteristics of the computer system, wherein the multiple sets of computational unit groups have different computational performance; a sorting module for sorting the multiple computational tasks of the machine learning model based on computational load; and a mapping module for selecting one of one or more computational unit groups with the highest computational performance from one or more available computational unit groups for each of the sorted computational tasks, and mapping the computational task to a computational unit in the selected computational unit group.
[0014] For example, at least one embodiment of this disclosure provides an apparatus for deploying a machine learning model in a computer system, the apparatus comprising: a memory storing executable instructions; and at least one processor coupled to the memory, wherein the at least one processor is configured to implement the methods described in any of the various embodiments of this disclosure when executing the executable instructions.
[0015] For example, at least one embodiment of the present disclosure provides a non-transitory computer-readable medium storing executable instructions, wherein the executable instructions, when executed by a processor, cause the processor to perform the methods described in any of the various embodiments of the present disclosure. Attached Figure Description
[0016] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings of the embodiments will be briefly described below. Obviously, the drawings described below only relate to some embodiments of this disclosure and are not intended to limit this disclosure.
[0017] Figure 1 A method for deploying a machine learning model in a computer system according to at least one embodiment of the present disclosure is shown;
[0018] Figure 2 An apparatus for deploying a machine learning model in a computer system according to at least one embodiment of the present disclosure is shown;
[0019] Figure 3 An apparatus for deploying a machine learning model in a computer system according to at least one embodiment of the present disclosure is shown;
[0020] Figure 4A An example of an in-memory computing hardware system is shown;
[0021] Figure 4B Another example of an in-memory computing hardware system is shown;
[0022] Figure 5 An example of a memory computing unit in a memory computing hardware system is shown;
[0023] Figure 6 Different types of computing resources are illustrated in a memory computing hardware system according to at least one embodiment of the present disclosure;
[0024] Figure 7 A pipeline-aware deployment method for in-memory computing systems according to at least one embodiment of the present disclosure is illustrated;
[0025] Figure 8 A pipelined sensing deployment apparatus for in-memory computing systems according to at least one embodiment of the present disclosure is shown;
[0026] Figure 9 At least one embodiment of the present disclosure is shown. Figure 8 The grouping process of the pipeline sensing deployment device is shown;
[0027] Figure 10 At least one embodiment of the present disclosure is shown. Figure 8The sequencing process of the pipeline sensing deployment device is shown.
[0028] Figure 11 At least one embodiment of the present disclosure is shown. Figure 8 The mapping process of the pipeline sensing deployment device is shown;
[0029] Figure 12A The architecture of an existing in-memory computing compiler is shown;
[0030] Figure 12B The architecture of a compiler for a fusion pipeline-aware deployment method according to at least one embodiment of the present disclosure is shown;
[0031] Figure 13 Performance improvements of a pipeline-aware deployment method according to at least one embodiment of the present disclosure are illustrated. Detailed Implementation
[0032] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.
[0033] Unless otherwise defined, the technical or scientific terms used in this disclosure shall have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms “first,” “second,” and similar terms used in this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, the terms “an,” “a,” or “the,” and similar terms do not indicate a quantity limitation, but rather indicate the presence of at least one. The terms “including,” “comprising,” or “containing,” and similar terms mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. The terms “connected,” “linked,” or similar terms are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. The terms “upper,” “lower,” “left,” and “right,” etc., are used only to indicate relative positional relationships, and these relative positional relationships may change accordingly when the absolute position of the described objects changes.
[0034] In deploying machine learning models in computer systems, researchers have introduced software tools such as compilers to reduce the overhead of manual labor, effectively completing the deployment of machine learning models to computer systems. However, previous compiler designs did not consider the dataflow computation characteristics of computer system hardware, and could only guarantee the correctness of deployment, not the computational performance of the algorithm model after deployment, resulting in significant room for optimization in the deployment scheme.
[0035] At least one embodiment of this disclosure proposes a pipeline-aware deployment device for in-memory computing systems, based on the hardware characteristics of such systems, which can be integrated into a compiler. This allows for the full utilization of hardware characteristics to optimize deployment schemes and improve the performance of in-memory computing systems when inferring different deep learning tasks, while automatically completing task deployment.
[0036] The following is combined with Figure 1 This disclosure describes a method for deploying machine learning models in a computer system.
[0037] Figure 1 A method for deploying a machine learning model in a computer system according to at least one embodiment of the present disclosure is shown. The method includes steps S110-S140.
[0038] In step S110, multiple computational tasks of the machine learning model are extracted.
[0039] The machine learning model according to this disclosure can include, but is not limited to, deep learning models such as neural network models. For example, neural network models can include, but are not limited to, convolutional neural networks, recurrent neural networks, etc. For deep learning models, multiple computational tasks can be, for example, computations between adjacent neural network layers, but the embodiments of this disclosure are not limited thereto.
[0040] In step S120, based on the hardware characteristics of the computer system, the multiple computing units of the computer system are grouped into multiple computing unit group sets, wherein the multiple computing unit group sets have different computing performance.
[0041] The computing system according to this embodiment of the present disclosure may include, but is not limited to, a memory computing system, and the computing unit may be a memory computing unit in a memory computing system that includes at least one memory computing array and peripheral driving circuitry for at least one memory computing array, but the embodiments of the present disclosure are not limited thereto.
[0042] In step S130, the multiple computational tasks of the machine learning model are sorted based on computational cost.
[0043] For example, when the machine learning model is a deep learning model, multiple computational tasks can be sorted according to the computational cost between adjacent neural networks of the deep learning model. In the deployment of the deep learning model, the computational cost can refer to the number of matrix multiplications required to complete one feature map computation of the neural network.
[0044] In step S140, for each of the sorted multiple computing tasks, one computing unit group with the required computing performance is selected from one or more available computing unit groups, and the computing task is mapped to a computing unit in the selected computing unit group.
[0045] For example, computing tasks can be sorted from high to low based on computational load, starting with the computing task with the highest computational load. For each computing task, the computing unit group with the highest computing performance can be selected from one or more available computing unit groups, and the computing task can be mapped to the computing unit in the selected computing unit group. However, the embodiments of this disclosure are not limited to this.
[0046] based on Figure 1 The method shown for deploying machine learning models in a computer system can fully utilize the hardware characteristics of the computer system to optimize the deployment of machine learning models, thereby improving computing performance.
[0047] Figure 2 An apparatus for deploying a machine learning model in a computer system according to at least one embodiment of the present disclosure is shown.
[0048] like Figure 2 As shown, an apparatus 200 for deploying a machine learning model in a computer system according to at least one embodiment of the present disclosure is illustrated. The apparatus 200 may include a parsing module 210, a grouping module 220, a sorting module 230, and a mapping module 240. It will be understood that for... Figure 2 The device shown can also be equipped with other additional modules. For example, it can employ... Figure 2 The apparatus 200 is used to implement the methods according to various embodiments of the present disclosure. These modules can be implemented, for example, by hardware, software, firmware, or any combination thereof.
[0049] like Figure 2The parsing module 210 shown can be used, for example, to extract multiple computational tasks from a machine learning model. The machine learning model according to embodiments of this disclosure may include, but is not limited to, deep learning models such as neural network models, for example, convolutional neural networks, recurrent neural networks, etc. For deep learning models, multiple computational tasks may be, for example, computations between adjacent neural network layers, but embodiments of this disclosure are not limited thereto.
[0050] The grouping module 220 can be used, for example, to group multiple computing units of a computer system into multiple sets of computing unit groups based on the hardware characteristics of the computer system, wherein the multiple sets of computing unit groups have different computing performance. The computing system according to at least one embodiment of this disclosure may include, but is not limited to, an in-memory computing system, and the computing unit may be an in-memory computing unit that includes at least one in-memory computing array and peripheral driving circuitry for the at least one in-memory computing array; however, embodiments of this disclosure are not limited thereto.
[0051] The sorting module 230 can, for example, be used to sort multiple computational tasks of a machine learning model based on computational cost. For example, when the machine learning model is a deep learning model, multiple computational tasks can be sorted according to the computational cost between adjacent neural networks of the deep learning model, where computational cost in the deployment of the deep learning model can refer to the number of matrix multiplications required to complete one feature map computation of the neural network.
[0052] The mapping module 240 can be used to select, for each of a sorted plurality of computing tasks, one of one or more computing unit groups with the highest computing performance from one or more available computing unit groups, and map the computing task to a computing unit in the selected computing unit group. For example, the computing tasks can be sorted from high to low based on the amount of computing power, starting with the computing task with the highest amount of computing power, and for each computing task, the computing unit group with the highest computing performance can be selected from one or more available computing unit groups, and the computing task can be mapped to a computing unit in the selected computing unit group. However, the embodiments of this disclosure are not limited thereto.
[0053] Figure 3 An apparatus 300 for deploying a machine learning model in a computer system according to at least one embodiment of the present disclosure is shown.
[0054] Figure 3 The device 300 shown may include a memory 310 and at least one processor 320, wherein the at least one processor 320 is coupled to the memory 310.
[0055] The memory 310 may store executable instructions that, when executed by the processor 320, at least partially implement the method for deploying the machine learning model described above. For example, the memory 310 may include, but is not limited to, read-only memory (ROM), random access memory (RAM), hard disk drive, optical disc (CD), digital video disc (DVD), or any other type of memory.
[0056] like Figure 3 The at least one processor 320 shown may include, but is not limited to, a central processing unit (CPU), an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field-programmable gate array (FPGA), a system-on-a-chip (SoC), a programmable or non-programmable logic device or array, digital circuits, a microprocessor, an application-specific integrated circuit (ASIC), etc. The at least one processor 320 may be configured to, for example, perform... Figure 1 The method shown is for deploying a machine learning model in a computer system according to at least one embodiment of the present disclosure.
[0057] It should be understood that, for Figure 3 The device shown can also have other additional components added. For example, it can employ... Figure 3 The device 300 in the computer system deploys machine learning models to implement methods according to various embodiments of the present disclosure.
[0058] The method for deploying machine learning models in a computer system according to at least one embodiment of this disclosure can be applied, for example, to in-memory computing hardware systems. The in-memory computing system mentioned in this disclosure can refer to a computing system composed of devices capable of implementing an in-memory computing architecture, such as, but not limited to, novel memories such as memristors, resistive random access memory (RRAM), phase-change memory (PCM), and magnetic random access memory (MRAM), as well as memories such as static random access memory (SRAM) or flash memory. Correspondingly, the in-memory computing array may include a memristor array, and the peripheral driving circuitry includes, but is not limited to, word line drivers, source line drivers, and bit line drivers.
[0059] The in-memory computing hardware system involved in the embodiments of this disclosure can be, for example, a hardware system having in-memory computing units and their interconnection structures. Furthermore, the embodiments of this disclosure do not impose any special restrictions on the interconnection topology of the in-memory computing system, and may include, but are not limited to, for example, network-on-chip (NOC), bus interconnection, star interconnection, ring interconnection, etc.
[0060] The following is combined with Figures 4A-4B Examples of in-memory computing hardware systems according to at least one embodiment of the present disclosure are described.
[0061] Figure 4A An example of an in-memory computing hardware system is shown. Figure 4A The in-memory computing hardware system shown is a bus-interconnected in-memory computing hardware system.
[0062] like Figure 4A As shown, an in-memory computing hardware system may include a global control unit, multiple in-memory computing unit blocks, a bus, and a shared storage unit. In such a case... Figure 4A In the in-memory computing hardware system shown, each of the multiple in-memory computing unit blocks may include one or more in-memory computing units and local storage units, and the global control unit, multiple in-memory computing unit blocks and shared storage units may be connected via a bus. Figure 4A The shared and local storage units shown may include, but are not limited to, read-only memory (ROM), random access memory (RAM), or any other type of memory. Figure 4A The in-memory computing unit may include, but is not limited to, at least one in-memory computing array and peripheral driving circuitry for at least one in-memory computing array.
[0063] For example, Figure 4A The in-memory computing hardware system described herein can be used to implement methods according to various embodiments of this disclosure.
[0064] Figure 4B Another example of an in-memory computing hardware system is shown. Figure 4B The in-memory computing hardware system shown is an NOC interconnect in-memory computing hardware system.
[0065] like Figure 4B As shown, an in-memory computing hardware system can include multiple in-memory computing unit blocks and shared storage units. In, for example... Figure 4B In the in-memory computing hardware system shown, each of the multiple in-memory computing unit blocks may include a control unit, one or more in-memory computing units, and a local storage unit, and the multiple in-memory computing unit blocks and the shared storage unit are connected in a NOC manner. Figure 4BThe shared and local storage units shown may include, but are not limited to, read-only memory (ROM), random access memory (RAM), or any other type of memory. Figure 4B The in-memory computing unit may include, but is not limited to, at least one in-memory computing array and peripheral driving circuitry for at least one in-memory computing array.
[0066] For example, Figure 4B The in-memory computing hardware system described herein can be used to implement methods according to various embodiments of this disclosure.
[0067] Figure 5 An example of a memory computing unit in a memory computing hardware system is shown. Figure 5 The storage unit shown can be used for, for example Figure 4A , Figure 4B The in-memory computing hardware system shown.
[0068] Figure 5 The in-memory computing unit shown may include, but is not limited to, at least one in-memory computing array and peripheral driving circuitry for at least one in-memory computing array. For example, the peripheral driving circuitry for the in-memory computing array may include input registers, digital-to-analog converters, analog-to-digital converters, shift-accumulator units, output registers, special function units, control units, etc., but it should be understood that these may be omitted. Figure 5 One or more of the peripheral circuits shown may be included, or additional peripheral circuits may be added. For example, the in-memory computing array may include a memristor array; the memristor array includes multiple memristor cells, each memristor cell may include at least one transistor and at least one memristor, and may be composed of 1T1R (i.e., one transistor + one memristor), or 2T2R (i.e., two transistors + two memristors), etc., which will not be elaborated further, and the embodiments of this disclosure do not limit this. The peripheral driving circuits may be implemented by digital and / or analog circuits, for example, and the embodiments of this disclosure also do not limit this.
[0069] Figure 6 Different types of computing resources are shown in a memory computing hardware system according to at least one embodiment of the present disclosure.
[0070] like Figure 6 As shown, according to at least one embodiment of this disclosure, the hardware computing resources of a memory computing hardware system can be categorized. For example, different types of computing resources can have different computing performance. The categorization can be based on the hardware characteristics of the memory computing hardware system; for example, the types of computing resources can be categorized based on the number of one or more computing units and the access relationships between one or more computing units and the memory units of the memory computing hardware system. According to at least one embodiment of this disclosure, Figure 6The computing unit shown may include, but is not limited to, for example: Figure 4A and Figure 4B The memory storage units and / or memory storage unit blocks, etc., and Figure 6 The storage unit shown may include, but is not limited to, for example: Figure 4A and Figure 4B Local storage units and / or shared storage units, etc.
[0071] For example, such as Figure 6 As shown, since in-memory computing systems operate on a data-flow-driven computation model, the primary dependency between computing units is data. Once the data meets the requirements of the computation task, the computing unit can begin computation. This data-flow-driven model enables in-memory computing units to exhibit higher computing performance, such as lower latency, compared to traditional computing devices like CPUs. Based on the characteristics of computing units accessing storage units to obtain input data, computing resources in in-memory computing systems can be categorized into strongly pipelining, weakly pipelining, and non-pipelined systems.
[0072] According to at least one embodiment of this disclosure, strong pipelining refers to the allocation of different computing tasks to different computing units, and the read, write, and storage operations between different computing units are independent of each other. For example, as Figure 6 As shown, computation task 1 is assigned to computation unit 1, computation task 2 is assigned to computation unit 2, and computation unit 1 and computation unit 2 access storage unit 1 and storage unit 2 independently, respectively. According to... Figure 6 In high-flow computing, computational units 1 and 2 can execute calculations independently and continuously. As long as the data requirements of the computation task are met, all computational units can be fully utilized, achieving the highest hardware utilization rate and thus maximizing performance. Figure 6 The high-flow type shown illustrates an example of assigning two computing tasks to two computing units, but the embodiments of this disclosure are not limited thereto; for example, a greater number of computing tasks can be assigned to a greater number of different computing units.
[0073] According to at least one embodiment of this disclosure, weak pipelining refers to the allocation of different computing tasks to different computing units, but the reading and writing of storage between computing units is shared, and at any given time, only one computing unit can access the storage. For example, as Figure 6 As shown, computation task 1 is assigned to computation unit 1, and computation task 2 is assigned to computation unit 2, but computation unit 1 and computation unit 2 share storage unit 1. According to... Figure 6 In weak pipelined scenarios, in addition to meeting the data requirements of the computation task, the read / write requirements of the storage unit also need to be met. This can lead to incomplete utilization of hardware resources. For example, ... Figure 6As shown, in the weak pipelining type, there is a time gap between the computation and memory access of Task 1, therefore the computational performance is lower compared to the strong pipelining. Figure 6 The weak pipeline type shown illustrates an example of assigning two computing tasks to two computing units, but the embodiments of this disclosure are not limited to this. For example, a greater number of computing tasks can be assigned to a greater number of different computing units.
[0074] According to at least one embodiment of this disclosure, pipelining refers to different computational tasks being assigned to the same computational unit. For example, such as Figure 6 As shown, computational task 1 and computational task 2 are assigned to computational unit 1. Figure 6 As shown, in weak pipelining, only one computational task can be executed at any given time. The computational tasks are executed completely serially, resulting in the lowest computational performance for the in-memory computing system. For example, non-pipelined computing resources can be used when hardware resources are scarce and each computational task cannot be allocated a separate computing unit.
[0075] According to embodiments of this disclosure, by dividing computing resources using the above method, the computing performance of different computing resources can be distinguished earlier. Moreover, this process is only related to hardware design and is independent of deep learning models. Thus, different types of computing resources can be selected for the computing tasks of deep learning models to efficiently guide the efficient deployment of different deep learning models.
[0076] It should be understood that Figure 6 This is merely an example of classifying different types of computing resources in an in-memory computing hardware system according to at least one embodiment of the present disclosure, and embodiments of the present disclosure are not limited thereto. Computing resources can be classified in other ways according to embodiments of the present disclosure. For example, the computing resources of an in-memory computing hardware system can be classified into fewer or more types. For example, the types of computing resources can be classified based on the size of the computing unit and / or the characteristics of the computing unit device, etc.
[0077] The method for deploying a machine learning model in a computer system according to at least one embodiment of the present disclosure can, for example, be used to deploy a deep learning model in a memory computing hardware system. For instance, it can be used to deploy a neural network model in a memory computing hardware system, wherein the neural network may include, but is not limited to, convolutional neural networks, recurrent neural networks, etc.
[0078] The following is combined with Figure 7 A pipeline-aware deployment method for deploying deep learning models in a memory computing hardware system is described according to at least one embodiment of the present disclosure.
[0079] Figure 7A pipeline-aware deployment method for in-memory computing systems according to at least one embodiment of the present disclosure is shown. The pipeline-aware deployment method includes steps S710 to S740.
[0080] like Figure 7 As shown, in step S710, multiple computational tasks of the neural network model are extracted.
[0081] For example, this step may include, but is not limited to, parsing the neural network model and extracting neural network layers suitable for in-memory computing systems; and / or transforming the extracted neural network layers and expanding the weights of the extracted neural network layers into weights suitable for in-memory computing systems, and splitting the weights of the neural network layers into appropriate sizes according to the hardware scale of the in-memory computing system; and / or parsing the parameters required to calculate the extracted neural network layers. For example, when the extracted neural network layer is a convolutional layer, the parameters required to calculate the convolutional layer may include, for example, the kernel size, the stride of the convolution, the padding size (zero padding operation on the input image), the size of the input convolution group, etc.
[0082] In step S720, based on the hardware characteristics of the in-memory computing system, the multiple computing units of the in-memory computing system are divided into multiple computing unit group sets, wherein the multiple computing unit group sets have different computing performance.
[0083] For example, an in-memory computing system can be like this Figure 4A and 4B The in-memory computing hardware system shown. The computing unit can be, for example, such as... Figure 4A and 4B The storage units and / or storage unit blocks shown are as follows.
[0084] For example, according to at least one embodiment of this disclosure, the hardware computing resources of an in-memory computing hardware system can be categorized. Different types of computing resources may have different computing performance. The categorization can be based on the hardware characteristics of the in-memory computing hardware system. For example, the types of computing resources can be categorized based on the number of one or more computing units and the access relationships between one or more computing units and the storage units of the in-memory computing hardware system. Storage units may include, but are not limited to, for example… Figure 4A and Figure 4B Local storage units and / or shared storage units, etc.
[0085] In step S730, multiple computational tasks of the deep learning model are sorted based on computational cost. For example, multiple computational tasks can be sorted according to the computational cost between adjacent neural networks of the deep learning model. In the deployment of the deep learning model, computational cost can refer to the number of matrix multiplications required to complete one feature map computation of the neural network.
[0086] In step S740, for each of the sorted multiple computing tasks, one computing unit group with the required computing performance is selected from one or more available computing unit groups, and the computing task is mapped to a computing unit in the selected computing unit group.
[0087] For example, computing tasks can be sorted from high to low based on computational load, starting with the computing task with the highest computational load. For each computing task, the computing unit group with the highest computing performance can be selected from one or more available computing unit groups, and the computing task can be mapped to the computing unit in the selected computing unit group. However, the embodiments of this disclosure are not limited to this.
[0088] based on Figure 7 The pipeline-aware deployment method for in-memory computing systems shown can fully utilize the hardware characteristics of in-memory computing hardware systems to optimize the deployment of deep learning algorithms during the deployment of deep learning models, thereby effectively completing the deployment task of deep learning models to in-memory computing systems and improving computing performance.
[0089] Figure 8 A pipelined awareness deployment system for in-memory computing systems according to at least one embodiment of the present disclosure is shown.
[0090] like Figure 8 As shown, a pipeline-aware deployment system 800 for in-memory computing systems according to at least one embodiment of this disclosure may include a parsing module 810, a grouping module 820, a sorting module 830, a mapping module 840, etc. It should be understood that... Figure 8 The pipeline-aware deployment system 810 for in-memory computing systems shown is merely an example; additional components, modules, and functions can be added as needed. For example, Figure 8 The pipeline-aware deployment system 810 for in-memory computing systems can be used to implement methods according to various embodiments of this disclosure. These modules can be implemented, for example, by software, hardware, firmware, or any combination thereof.
[0091] According to at least one embodiment of this disclosure, such as Figure 8The parsing module 810 shown can be used to extract multiple computational tasks from a neural network model. For example, the parsing module can receive the algorithm model 850 of the neural network model (including but not limited to Open Neural Network Exchange (ONNX) and scripts of neural network frameworks such as PyTorch and TensorFlow), and parse the algorithm description of the neural network model. For example, this parsing operation may include, but is not limited to, extracting neural network layers suitable for in-memory computing systems; and / or transforming the extracted neural network layers and expanding the weights of the extracted neural network layers into weights suitable for in-memory computing systems, and splitting the weights of the neural network layers into appropriate sizes according to the hardware scale of the in-memory computing system; and / or parsing the parameters required to compute the extracted neural network layers. For example, when the extracted neural network layer is a convolutional layer, the parameters required to compute the convolutional layer may include, for example, the kernel size, the stride of the convolution, the padding (zero-padding operation on the input image), and the size of the input convolution group.
[0092] According to at least one embodiment of this disclosure, such as Figure 8 The grouping module 820 shown can be used to classify different computing units (CUs) into different categories based on the hardware characteristics of the in-memory computing system. For example, the grouping module 820 can convert the hardware description 860 of the in-memory computing system into a hardware grouping description. The hardware grouping description can be represented in text format or serialized language, including but not limited to text files (TXT), JavaScript object notation (JSON), YAML, etc. For example, the grouping module 820 can obtain the relationships between different computing units (CUs) based on the current hardware description 860 of the in-memory computing system, group different computing resources (e.g., grouping them in pairs, but the embodiments of this disclosure are not limited to this), and classify the computing unit groups into different types based on the number of computing units within each computing unit group and the relationship between the computing units and the storage units. For example, according to the embodiments of this disclosure, it can be used to... Figure 6 The diagram illustrates how computing resources are categorized into three types: highly pipelining, weakly pipelining, and non-pipelined. Highly pipelining resources offer the highest performance and can operate entirely in a pipelining manner. Weakly pipelining resources offer moderate performance and can operate partially in a pipelining manner. Non-pipelined resources offer the lowest performance and can only operate in a serial manner.
[0093] According to embodiments of this disclosure, the grouping module 820 can pre-distinguish the computational performance of different computing resources. This process is only related to hardware design and is independent of the deep learning model. Therefore, different types of computing resources can be selected for the computational tasks of the deep learning model to efficiently guide the deployment of different deep learning models. For example, based on the classification of strong pipelining, weak pipelining, and no pipelining, all hardware computing resources can be assigned to these three types respectively. Then, different computing layers can be allocated to these three types of hardware resources to complete the deployment. It should be understood that the method of classifying computing resource types described above is merely an example, and embodiments of this disclosure are not limited thereto.
[0094] According to embodiments of this disclosure, the sorting module 830 can be used to sort multiple computational tasks of a deep learning model based on computational complexity. For example, computational tasks may include computations between adjacent layers of a neural network model, and the sorting module 830 can sort the computations between adjacent neural network layers generated by the parsing module 810 according to the computational complexity of adjacent neural network layers. For example, computational complexity may be the number of matrix multiplications required to complete one feature map computation of the neural network.
[0095] According to embodiments of this disclosure, the mapping module 840 can be used to map the hardware resources generated by the grouping module 820 to the computing tasks generated by the sorting module 830 until all computing tasks are mapped to appropriate hardware resources, thereby outputting a complete model deployment scheme 870. The mapping module 840 can integrate different mapping strategies according to different objectives. For example, for each computing task among multiple sorted computing tasks, it can select one of one or more computing unit groups with the required computing performance from one or more available computing unit groups and map the computing task to a computing unit in the selected computing unit group. Alternatively, for example, the computing tasks can be sorted from high to low based on the amount of computing power, starting with the computing task with the highest computing power. For each computing task, it can select the computing unit group with the highest computing performance from one or more available computing unit groups and map the computing task to a computing unit in the selected computing unit group. However, embodiments of this disclosure are not limited to these methods.
[0096] The following is combined with Figures 9-11 Detailed description of at least one embodiment based on the present disclosure Figure 8 The diagram illustrates the grouping, sorting, and mapping processes of the pipeline-aware deployment system.
[0097] Figure 9 At least one embodiment of the present disclosure is shown. Figure 8 The diagram shows the grouping process of the pipeline-aware deployment system.
[0098] according to Figure 9 In the illustrated embodiment, the in-memory computing system 910 includes four in-memory computing units: CU1, CU2, CU3, and CU4. CU1 and CU2 share a local storage unit, and CU3 and CU4 share a local storage unit. It should be understood that... Figure 9 The in-memory computing system 910 shown is merely an example, and the embodiments disclosed herein are not limited thereto.
[0099] like Figure 9 As shown, the in-memory computing units in the in-memory computing system 910 are grouped in pairs, and each group of in-memory computing units is classified into one of three types: strong flow, weak flow, and no flow. Figure 9 Since there are no shared memory units between CU1, CU2 and CU3, CU4, the relationships between CU1 and CU3, CU4, as well as between CU2 and CU3, CU4, satisfy strong pipelining. The relationships between CU1 and CU2, and between CU3 and CU4, satisfy weak pipelining. The relationships within CU1, CU2, CU3, CU4 are non-pipelining.
[0100] It should be understood that Figure 9 The grouping process shown is merely an example. According to various embodiments of this disclosure, the computing resources of the in-memory computing system can be divided in different grouping ways. For example, the in-memory computing units in the in-memory computing system 910 can be grouped in different numbers, or the in-memory computing units can be divided into more types, or the in-memory computing blocks including one or more in-memory computing units can be divided, etc.
[0101] For ease of explanation, based on Figure 9 The grouping process shown below Figure 10 and Figure 11 Describe it.
[0102] Figure 10 At least one embodiment of the present disclosure is shown. Figure 8 The sequencing process of the pipeline-aware deployment system is shown.
[0103] like Figure 10 The sorting process shown prioritizes the computational tasks of deep learning models based on the computational complexity of adjacent neural network layers. For example, by considering the computational complexity of adjacent neural network layers, the priority of computational tasks can be determined: layers with higher computational complexity have higher priority, and vice versa. Priority then determines which computational resources should be used to process the task during mapping; higher priority means using hardware resources with higher computational performance.
[0104] For example, such as Figure 10As shown, the extracted neural network layers are convolutional layers and fully connected layers, where Conv1 and Conv2 represent convolutional computations, and FC represents fully connected computations. The sorted results show that the computational cost of Conv1 and Conv2 is greater than that of Conv2 and FC. Therefore, in the subsequent mapping process, the deployment of Conv1 and Conv2 should be prioritized. Figure 10 The sorting process shown is merely an example, but the embodiments disclosed herein are not limited thereto.
[0105] Figure 11 At least one embodiment of the present disclosure is shown. Figure 8 The mapping process of the pipeline-aware deployment system is shown.
[0106] Figure 11 The mapping process shown can Figure 10 The sorted computational tasks shown are Figure 9 The categorized computing resources shown are mapped one-to-one to determine that each computing task can be completed by a single computing resource. For example, in order to... Figure 10 The computing tasks shown are deployed to Figure 9 The mapping principle for the categorized computing resources shown is to first consider finding suitable computing resources from the highly pipelining types, and only after finding no suitable resources in the highly pipelining types should the search proceed to the less pipelining types, and finally consider the non-pipelined types. Based on the priority of computational load, the mapping prioritizes the deployment of Conv1 and Conv2. In the highly pipelining computing resources, in-memory unit groups CU1 and CU3 are found, and CU1 and CU3 within these groups are assigned to Conv1 and Conv2 respectively. Then, the computational tasks of Conv2 and FC are considered. Since Conv2 has already been assigned to CU3, the search should first look for in-memory unit groups starting with CU3 in the highly pipelining computing resources. Finding no suitable computing resources, the search then proceeds to the less pipelining computing resources, where suitable in-memory unit groups—CU3 and CU4—are found. The final deployment scheme is that Conv1 is computed in CU1, Conv2 in CU3, and FC in CU4. The above mapping process is merely an example, and the embodiments disclosed herein are not limited thereto. For example, for the deployment of Conv1 and Conv2, a group of CU2 and CU4 can be selected in the high-pipeline type computing resources, and CU2 and CU4 can be assigned to Conv1 and Conv2 respectively. In the low-pipeline type computing resources, a group of CU4 and CU3 can be selected to allocate the computing tasks of Conv2 and FC. The final deployment scheme is that Conv1 is computed in CU2, Conv2 is computed in CU4, and FC is computed in CU3.
[0107] The pipeline-aware deployment method for in-memory computing systems according to at least one embodiment of this disclosure can be integrated into an in-memory computing compiler. Figure 12A The architecture of an existing in-memory compiler is shown, while Figure 12B The architecture of a compiler for a fusion pipeline-aware deployment method according to at least one embodiment of the present disclosure is shown.
[0108] like Figure 12A As shown, existing compiler implementations are mainly divided into two parts: a front-end and a back-end. The front-end is primarily responsible for handling model-related transformations, such as model parsing and hardware-independent operator optimizations, like operator fusion. For the processed model, the front-end can output a new representation format, generally called an intermediate representation (IR). This format facilitates processing by the back-end to generate a hardware-executable file. The back-end then maps the algorithm's intermediate representation to hardware resources based on hardware information, including the allocation of computational units and storage units. Finally, the code generation module generates the executable code for the algorithm model on the hardware based on the allocation results.
[0109] like Figure 12B As shown, the pipeline-aware deployment method according to at least one embodiment of this disclosure can be integrated into the compiler implementation by adding an operator sorting module in the compiler front-end and a hardware grouping module in the compiler back-end. For example, the operator sorting module can sort different computational tasks in the optimized algorithm according to their computational load. The hardware grouping module can divide hardware resources into different types according to their computational performance, thereby providing guidance for the mapping module. Finally, the model mapping module only needs to map the sorted computational tasks one by one to the categorized computational resources. According to embodiments of this disclosure, Figure 12B The operator sorting module, hardware grouping module, and model mapping module shown can respectively correspond to, for example, the following modules: Figure 8 The diagram shows the sorting module, grouping module, and mapping module in the pipeline-aware deployment system. Compared to existing implementations, the compiler implementation that integrates the pipeline-aware deployment method can utilize hardware information more effectively, resulting in a higher-performance deployment solution.
[0110] Figure 13 A performance comparison chart is shown between a pipeline-aware deployment scheme according to at least one embodiment of the present disclosure and an existing deployment scheme.
[0111] like Figure 13As shown, this paper mainly compares the normalized inference latency of three different deep learning networks—LeNet, VGG11, and ResNet-18—in existing deployment schemes and pipelined aware deployment schemes. Lower latency indicates better performance. Normalized inference latency is obtained by normalizing the average time required to infer a single image in existing deployment schemes. Existing deployment schemes refer to algorithms that sort inputs according to their order of appearance, and hardware resources that sort them according to storage unit dependencies (first sorting computation units with shared storage units, then sorting computation units with non-shared storage units), then mapping computational tasks to hardware resources one-to-one according to their order of appearance. Figures 9-11 Taking the in-memory computing system and corresponding neural network model as an example, in the existing deployment scheme, the order of computing tasks is Conv1, Conv2, FC, and the order of hardware resources is computing units CU1, CU2, CU3, CU4. Therefore, the final mapping scheme is that computing task Conv1 is calculated in computing unit CU1, computing task Conv2 is calculated in computing unit CU2, and computing task FC is calculated in computing unit CU3.
[0112] exist Figure 13 In the example shown, the performance simulator of the in-memory computing system was used to perform model inference for different deployment schemes. A total of 5 images were inferred, and the average of the 5 inference times was taken as the statistical result. Figure 13 The results show that the pipeline-aware deployment scheme achieves lower inference latency on all three network tasks shown. The deeper the neural network layers of the deep learning model, the more significant the latency reduction effect.
[0113] When deploying machine learning models in traditional computer systems, only the correctness of the deployment can be guaranteed, not the computational performance of the algorithm model after deployment. To address this issue, at least one embodiment of this disclosure proposes a method for deploying machine learning models based on the hardware characteristics of a computer system. This method can fully utilize the hardware characteristics of the computer system during the deployment of machine learning models, optimize the deployment of machine learning models, and thus improve computational performance.
[0114] At least one embodiment of this disclosure proposes a compiler architecture that can be used for a pipeline-aware deployment scheme and a fused pipeline-aware deployment scheme for deploying deep learning models in in-memory computing systems. Based on the pipeline-aware deployment scheme for in-memory computing systems, the deployment of deep learning models can be optimized based on the hardware characteristics of the in-memory computing system, thereby improving computational performance.
[0115] As used in this article, “element,” “unit,” “module,” “component,” “device,” or “system” can refer to software, hardware, or a combination of software and hardware.
[0116] In some exemplary embodiments, any element and / or unit and / or functional block and / or module and / or device and / or system described with reference to the accompanying drawings may include or be implemented in a processing circuit, which is, for example, hardware including logic circuitry; a hardware / software combination, such as a processor executing software; or a combination thereof. Specifically, the processing circuit may include, but is not limited to, a central processing unit (CPU), an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field-programmable gate array (FPGA), a system-on-a-chip (SoC), a programmable or non-programmable logic device or array, digital circuitry, a microprocessor, an application-specific integrated circuit (ASIC), etc. The processing circuit may include electronic components such as transistors, resistors, capacitors, etc. The processing circuit may include electronic components such as logic gates (including at least one of AND gates, OR gates, NAND gates, NOT gates, etc.).
[0117] Multiple processors, multiple controllers, and / or processing circuitry can be specially programmed to perform actions or steps (such as using an FPGA or ASIC), or can be configured to perform actions or steps by executing instructions received from memory, or a combination thereof.
[0118] The above description is merely a specific embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any person skilled in the art can make various changes or substitutions within the technical scope disclosed in this disclosure, and such changes or substitutions should all be covered within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.
Claims
1. A method for deploying a machine learning model in a computer system, the method comprising: Extract multiple computational tasks from the machine learning model. Based on the hardware access characteristics between the computing units and storage units of the computer system, the multiple computing units of the computer system are grouped into multiple computing unit group sets, wherein the multiple computing unit group sets have different computing performance. The multiple computational tasks of the machine learning model are ranked based on computational cost. For each of the sorted computing tasks, select one of one or more computing unit groups with the required computing performance from one or more available computing unit groups, and map the computing task to a computing unit in the selected computing unit group. Specifically, based on the hardware access characteristics between the computing units and storage units of the computer system, the plurality of computing units of the computer system are grouped into a plurality of computing unit group sets, including: Based on the number of one or more computing units included in a computing unit group and the access relationship between the one or more computing units included in the computing unit group and the storage unit of the computer system, the multiple computing unit groups of the computer system are divided into multiple computing unit group sets with different computing performance.
2. The method according to claim 1, wherein, The plurality of computing unit group sets include a first computing unit group set, a second computing unit group set, and a third computing unit group set, wherein, The computing unit group in the first computing unit group set includes multiple first computing units, and the multiple first computing units are configured to simultaneously access different first storage units. The computing unit grouping in the second computing unit grouping set includes multiple second computing units, and the multiple second computing units are configured to access the same second storage unit in a time-sharing manner. The computing units in the third computing unit group set include only independent third computing units, and the third computing units are configured to access third storage units.
3. The method according to claim 1, wherein, The computer system is an in-memory computing system, and The computing unit is a memory computing unit, wherein the memory computing unit includes at least one memory computing array and peripheral driving circuitry for the at least one memory computing array.
4. The method according to claim 3, wherein, The machine learning model mentioned is a deep learning model.
5. The method according to claim 4, wherein, Each of the plurality of computational tasks includes computation between adjacent network layers of the deep learning model, and The process of ranking the multiple computational tasks of the machine learning model based on the computational load includes: The multiple computational tasks of the deep learning model are ordered according to the computational cost between adjacent network layers, wherein... The computational complexity refers to the number of matrix multiplications required to complete one feature map calculation of the deep learning model.
6. The method according to claim 5, wherein, Each of the plurality of computing unit groups includes a first computing unit and a second computing unit, and Mapping the computational tasks to the computational units in the selected group of computational units includes: Mapping the computation of the first network layer included in the computational task of the deep learning model to the first computational unit in the selected group of computational units, and The computation of the second network layer included in the computational task of the deep learning model is mapped to the second computational unit in the selected group of computational units. The second network layer performs calculations based on the output of the first network layer, and the second calculation unit performs calculations based on the output of the first calculation unit.
7. The method according to claim 4, wherein, Extracting the multiple computational tasks of the machine learning model, including: Based on parameter analysis, multiple computational layers of the deep learning model used for computation in the in-memory unit and the parameters corresponding to the multiple computational layers are extracted.
8. The method according to claim 7, further comprising: The extracted weights of the multiple computing layers are expanded into weights for computation in the in-memory unit, and The weights are split based on the size of the storage unit.
9. An apparatus for deploying a machine learning model in a computer system, the apparatus comprising: The parsing module is used to extract multiple computational tasks from the machine learning model. The grouping module is used to divide multiple computing units of the computer system into multiple computing unit group sets based on the hardware access characteristics between the computing units and storage units of the computer system. These multiple computing unit group sets have different computing performance. The sorting module is used to sort the multiple computational tasks of the machine learning model based on computational cost. The mapping module is used to select, for each of a sorted plurality of computational tasks, one of one or more computational unit groups with the highest computational performance from one or more available computational unit groups, and map the computational task to a computational unit in the selected computational unit group. Specifically, based on the hardware access characteristics between the computing units and storage units of the computer system, the plurality of computing units of the computer system are grouped into a plurality of computing unit group sets, including: Based on the number of one or more computing units included in a computing unit group and the access relationship between the one or more computing units included in the computing unit group and the storage unit of the computer system, the multiple computing unit groups of the computer system are divided into multiple computing unit group sets with different computing performance.
10. An apparatus for deploying a machine learning model in a computer system, the apparatus comprising: Memory, which stores executable instructions; At least one processor coupled to the memory, wherein the at least one processor is configured to perform the method as described in any one of claims 1-8 when executing the executable instructions.
11. A non-transitory computer-readable medium storing executable instructions, wherein, The executable instructions, when executed by a processor, cause the processor to perform the method as described in any one of claims 1-8.
Citation Information
Patent Citations
Method, apparatus for processing computing task and computer program product
CN110389824A
Neural network compiling method for storage and calculation integrated platform
CN112465108A
Method for deploying deep learning model to acceleration unit
CN113743567A
Topology aware grouping and provisioning of GPU resources in GPU-as-a-Service platform
US10325343B1