In-memory computing chip architecture using sparsity and in-memory computing chip
Through the sparse in-memory computing chip architecture, the edge-end AI model training process is optimized, which solves the problem of low hardware utilization and achieves more efficient training efficiency and hardware utilization.
Patent Information
- Application Number
- CN202510365811.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-11
AI Technical Summary
The existing edge-end AI model training technology has challenges in hardware utilization, training accuracy and efficiency, especially due to the imbalance of forward propagation and backpropagation computational volume and large storage requirements, resulting in low hardware utilization.
The sparse in-memory computing chip architecture is adopted to accelerate the model training process through the first and second CIM units, and the load scheduler is used to split and combine the forward propagation and backpropagation tasks, generate a preliminary scheduling plan, and adjust the task allocation when the no-load time exceeds the preset upper limit to improve hardware utilization.
By optimizing task scheduling, the no-load time is shortened, the hardware utilization and training efficiency during model training is improved, and the training time is reduced.
Smart Images

Figure CN120295730A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of chip architectures, and in particular, to a memory - in - compute chip architecture and a memory - in - compute chip that utilize sparsity. Background Art
[0002] With the development of Artificial Intelligence (AI), AI models such as Transformer and Convolutional Neural Network (CNN) have achieved excellent performance in many fields. Currently, most neural network trainings are deployed in the cloud, consuming a large amount of computing resources, electricity, and network transmission bandwidth. Therefore, people have gradually started to study the deployment of these neural network trainings at the edge. Edge training enables devices to update parameters without an Internet connection, preventing user data from being uploaded and protecting personal privacy. In addition, edge training can also perform personalized model optimization according to users' usage habits, making the model more adaptable to users' preferences. Deploying the model directly at the edge enables the device to transmit data without a network, featuring low latency.
[0003] Compute - in - Memory (CIM) is an emerging computing paradigm that tightly couples computing units with storage units through customized circuits, eliminating the overhead of data movement and giving great advantages in computing energy efficiency and density compared to traditional von Neumann architectures. Currently, many edge - side inference chips utilize CIM as the computing unit. However, current edge - side AI model deployment technologies focus on model inference and rarely study edge - side training deployment.
[0004] The training of AI models mainly includes: forward propagation, backward propagation, and gradient update. Compared with inference, training requires additional backward propagation, and during forward propagation, the activation values of the intermediate layers of the model also need to be stored for subsequent backward propagation calculations. In addition, the requirements for data accuracy in model training are more stringent than those in model inference. Mainstream edge - side inference of AI models uses integer data precision, but training requires floating - point numbers to ensure that the training results can converge. Therefore, training has the characteristics of large computational volume, large storage requirements, and diverse computational modes compared to inference, posing challenges to current edge - side hardware in terms of hardware utilization, training accuracy, and training efficiency. Summary of the Invention
[0005] Multiple aspects of this application provide a memory - in - compute chip architecture and a memory - in - compute chip that utilize sparsity to improve hardware utilization during model training.
[0006] An embodiment of this application provides a memory - in - compute chip architecture that utilizes sparsity, including:
[0007] A central processing unit, which includes a first CIM unit and a load scheduler;
[0008] A neural network processor, where the neural network processor includes: a second CIM unit;
[0009] A memory, which is used to store data related to model training;
[0010] The load scheduler includes a configuration storage unit, a first scheduling unit, and a second scheduling unit;
[0011] The configuration storage unit is used to store the structure information of the model, the sparsity configured for the model, and a preset upper limit of the no-load duration; the model includes multiple network layers;
[0012] The configuration storage unit is further used to: determine the forward propagation calculation amount and forward propagation storage amount required for each network layer during forward propagation and the backward propagation calculation amount and backward propagation storage amount required for each network layer during backward propagation according to the structure information and the sparsity;
[0013] The first scheduling unit is used to: split and combine the i-th forward propagation task and the (i - 1)-th backward propagation task of the model according to the forward propagation storage amount and the backward propagation storage amount to obtain multiple task groups; each task group includes a first task packet and a second task packet that can be processed in parallel; where i is an integer greater than 1;
[0014] The second scheduling unit is used to determine the first total calculation amount of the first task packet and the second total calculation amount of the second task packet according to the forward propagation calculation amount and the backward propagation calculation amount; determine a preliminary scheduling plan and the no-load duration corresponding to the preliminary scheduling plan according to the first total calculation amount, the second total calculation amount, the configuration information of the first CIM unit, and the configuration information of the second CIM unit, where the preliminary scheduling plan is used to schedule one of the first task packet and the second task packet to the first CIM unit and the other to the second CIM unit; if the no-load duration corresponding to the preliminary scheduling plan is greater than or equal to the upper limit of the no-load duration, then schedule a part of the task packet responsible for the CIM unit without no-load situation in the preliminary scheduling plan to the CIM unit with no-load situation to obtain a target scheduling plan; perform task scheduling according to the target scheduling plan.
[0015] The embodiment of the present application further provides an in-memory computing chip, including: the in-memory computing chip architecture using sparsity described above.
[0016] In the technical solution provided by the embodiment of the present application, the first CIM unit and the second CIM unit in the in-memory computing chip architecture are used to accelerate the model training process. The first scheduling unit splits and combines the forward propagation task and the backward propagation task to obtain multiple task groups. The second scheduling unit first generates a preliminary scheduling plan for the first task packet and the second task packet in each task group, and then, when the idle time corresponding to the preliminary scheduling plan is greater than or equal to the preset upper limit of the idle time, part of the tasks in the task packets responsible for the CIM units with idle conditions in the preliminary scheduling plan are scheduled to the CIM units without idle conditions to obtain a target scheduling plan. Through this processing, the second scheduling unit can shorten the idle time, thereby improving the hardware utilization rate during model training. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0018] Figure 1 is the first schematic structural diagram of the chip architecture provided by an exemplary embodiment of the present application;
[0019] Figure 2 is the second schematic structural diagram of the chip architecture provided by an exemplary embodiment of the present application;
[0020] Figure 3 is the schematic structural diagram of the microarchitecture of the pipeline tightly coupled expansion module provided by an exemplary embodiment of the present application;
[0021] Figure 4 is the schematic structural diagram of the microarchitecture of the CIM workload scheduler provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Apparently, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0023] Before introducing the technical solution provided by the embodiment of the present application, a brief introduction to the technical terms involved in the technical solution provided by the embodiment of the present application is given:
[0024] Forward Propagation: It refers to the process in which input data is passed forward through the various network layers of a neural network to generate an output. In this process, the neurons in each layer calculate their activation values based on the current parameters and pass these values to the next layer until a final output prediction is produced.
[0025] Backward Propagation: It refers to the process that starts from the output layer and calculates the contribution of the parameters (including weights and biases) of each network layer to the loss (i.e., the parameter gradients) layer by layer backward, and updates the network parameters based on these parameter gradients.
[0026] Sparsity: It refers to the proportion of non-zero elements in data or calculations. High sparsity means that most elements are zero and only a small number of elements are non-zero. In deep learning, there are sparsity of parameters, sparsity of activation values, and sparsity of parameter gradients.
[0027] Sparse pattern: It is represented by a 0-1 sparse mask.
[0028] In practical applications, in order to reduce the computational amount of the edge segment AI model, the prior art performs sparsification processing (sparsity) on the model. For training, the sparsity degrees of forward propagation and backward propagation are often set differently, resulting in more imbalance in the originally mismatched computational amounts after introducing sparsity. When deploying these two imbalanced computations to the same hardware, it will lead to low hardware utilization.
[0029] The storage space of edge-side hardware is very limited, and at the same time, the training has a very large demand for storage space. The prior art reduces the storage demand of the model through sparsity and quantization. However, the higher the degree of sparse quantization, the worse the final effect of the model. Therefore, it is necessary to balance the two metrics of the degree of sparse quantization and model performance.
[0030] To solve or partially solve the above technical problems, the embodiments of the present application provide an in-memory computing chip architecture that utilizes sparsity. In this architecture, the first CIM unit and the second CIM unit in the in-memory computing chip architecture are used to accelerate the model training process. The first scheduling unit splits and combines the forward propagation tasks and backward propagation tasks to obtain multiple task groups. The second scheduling unit first generates a preliminary scheduling plan for the first task packet and the second task packet in each task group. Then, when the idle time corresponding to the preliminary scheduling plan is greater than or equal to the preset upper limit of the idle time, it schedules some tasks in the task packets responsible for the CIM units with idle situations in the preliminary scheduling plan to the CIM units without idle situations to obtain a target scheduling plan. Through this processing, the second scheduling unit can shorten the idle time, thereby improving the hardware utilization during model training.
[0031] The following will introduce in detail the in-memory computing chip architecture that utilizes sparsity provided by the embodiments of the present application in conjunction with the accompanying drawings.
[0032] Figure 1 The schematic diagram of the in-memory computing chip architecture that utilizes sparsity provided by an embodiment of the present application is shown. As Figure 1 shown, the chip architecture includes:
[0033] A central processing unit 1, where the central processing unit 1 includes a first CIM unit 11 and a load scheduler 12;
[0034] A neural network processor 2, where the neural network processor 2 includes: a second CIM unit 21;
[0035] A memory 3 for storing data related to model training;
[0036] The load scheduler 12 includes a configuration storage unit 121, a first scheduling unit 122, and a second scheduling unit 123;
[0037] The configuration storage unit 121 is used to store the structural information of the model, the sparsity configured for the model, and the preset upper limit of the idle time duration; the model includes multiple network layers;
[0038] The configuration storage unit 121 is further used to perform the following steps:
[0039] 101. According to the structural information and the sparsity, determine the forward propagation calculation amount and forward propagation storage amount required for each network layer during forward propagation, and the backward propagation calculation amount and backward propagation storage amount required for each network layer during backward propagation.
[0040] The first scheduling unit 122 is used to perform the following steps:
[0041] 102. According to the forward propagation storage amount and the backward propagation storage amount, split and combine the i-th forward propagation task and the (i - 1)-th backward propagation task of the model to obtain multiple task groups.
[0042] The second scheduling unit 123 is used to perform the following steps:
[0043] 103. According to the forward propagation calculation amount and the backward propagation calculation amount, determine the first total calculation amount of the first task packet and the second total calculation amount of the second task packet.
[0044] 104. According to the first total calculation amount, the second total calculation amount, the configuration information of the first CIM unit, and the configuration information of the second CIM unit, determine the preliminary scheduling plan and the idle time duration corresponding to the preliminary scheduling plan.
[0045] 105. If the no-load duration corresponding to the preliminary scheduling plan is greater than or equal to the preset upper limit of the no-load duration, then part of the tasks in the task packages responsible for the CIM units without no-load conditions in the preliminary scheduling plan are scheduled to the CIM units with no-load conditions, and a target scheduling plan is obtained.
[0046] 106. Execute task scheduling according to the target scheduling plan.
[0047] In some embodiments, the structural information of the model may include, but is not limited to, the model type, the number of network layers, the number of parameters of each network layer, and the weight size. The sparsity configured for the model may include: the sparsity configured for the parameters of each network layer, the sparsity configured for the activation values of each network layer, and the sparsity configured for the parameter gradients of each network layer. The storage upper limit of the memory is determined by the hardware resources of the memory. The size of the preset upper limit of the no-load duration can be set according to actual needs, and the embodiments of the present application do not make specific limitations on this.
[0048] In the above step 101, the forward propagation calculation amount and forward propagation storage amount required for each of the network layers during forward propagation refer to the forward propagation calculation amount and forward propagation storage amount required when each of the network layers is executed during one forward propagation process. The backpropagation calculation amount and backpropagation storage amount required for each of the network layers during backpropagation refer to the backpropagation calculation amount and backpropagation storage amount required when each of the network layers is executed during one backpropagation process.
[0049] In the above step 102, each task group includes a first task package and a second task package that can be processed in parallel; the total storage amount required when the first task package and the second task package are executed is less than the storage upper limit; where i is an integer greater than 1.
[0050] The i-th forward propagation task of the model refers to the task required to complete the i-th forward propagation process of the model, which includes the forward propagation subtasks of each network layer in multiple network layers during the i-th forward propagation process. The (i - 1)-th backpropagation task of the model refers to the task required to complete the (i - 1)-th backpropagation process of the model, which includes the backpropagation subtasks of each network layer in multiple network layers during the (i - 1)-th backpropagation process.
[0051] In some embodiments, the model includes N network layers, where N is an integer greater than 1. The forward propagation subtask of the nth network layer in the i-th forward propagation process can be combined with the backward propagation subtask of the (N + 1 - n)th network layer in the (i - 1)-th backward propagation process to form a task group, where n is an integer and n ranges from 1 to N. For example, when N is 9, the forward propagation subtask of the first network layer can be combined with the backward propagation subtask of the ninth network layer to form a task group, the forward propagation subtask of the second network layer can be combined with the backward propagation subtask of the eighth network layer to form a task group, the forward propagation subtask of the third network layer can be combined with the backward propagation subtask of the seventh network layer to form a task group, and the forward propagation subtask of the fourth network layer can be combined with the backward propagation subtask of the fifth network layer to form a task group.
[0052] In the embodiments of the present application, the number of task groups is N.
[0053] To reduce the resource consumption occupied by scheduling and the latency brought by scheduling, it can be achieved by reducing the number of scheduling times. In some other embodiments, the above-mentioned configuration storage unit is further configured to store the storage upper limit of the memory. The first scheduling unit, when executing step 102 above, is specifically configured to perform the following steps:
[0054] 1021. According to the forward propagation storage amount, the backward propagation storage amount, and the storage upper limit of the memory, split and combine the i-th forward propagation task and the (i - 1)-th backward propagation task of the model to obtain multiple task groups.
[0055] Among them, the total storage amount required for executing the first task packet and the second task packet in each task group is less than the storage upper limit; the multiple task groups include a first task group and / or a second task group; among them, the first task packet in the first task group includes the forward propagation subtasks of at least two consecutive first network layers respectively, and the second task packet in the first task group includes the backward propagation subtasks of at least one consecutive second network layer respectively; the first task packet in the second task group includes the forward propagation subtasks of at least one consecutive third network layer respectively, and the second task packet in the second task group includes the backward propagation subtasks of at least two consecutive fourth network layers respectively.
[0056] In practical applications, the above-mentioned multiple task groups may further include a third task group or a fourth task group. Among them, the first task package in the third task group includes a part of the forward propagation subtasks of each of at least one consecutive fifth network layer, and the first task package in the third task group includes another part of the forward propagation subtasks of each of the at least one consecutive fifth network layer. Among them, the first task package in the fourth task group includes a part of the backward propagation subtasks of each of at least one consecutive sixth network layer, and the first task package in the third task group includes another part of the backward propagation subtasks of each of the at least one consecutive sixth network layer.
[0057] It should be noted that the rules for splitting and combining can be set according to actual needs. The embodiments of the present application do not make specific limitations on this, as long as it is ensured that the first task package and the second task package in each task group meet the conditions for parallel processing, and the total storage required when the first task package and the second task package are executed in parallel is less than the storage upper limit, and the number of the multiple task groups is less than the number of the multiple network layers.
[0058] The sum of the first total storage required by the first task package and the second total storage required by the second task package is the total storage. Among them, the first total storage required by the first task package and the second total storage required by the second task package can be calculated according to the forward propagation storage and the backward propagation storage calculated in the above 101. The specific calculation method can be set according to actual needs, and the embodiments of the present application do not make specific limitations on this.
[0059] In some alternative embodiments, for the first task group (the first task package of which includes the forward propagation subtasks of each of at least two first network layers, and the second task package in the first task group includes the backward propagation subtasks of each of at least one second network layer), whether adding the forward propagation subtasks of the seventh network layer to the first task package or adding the backward propagation subtasks of the eighth network layer to the second task package will cause the total storage required when the first task package and the second task package in the first task group are executed to exceed the storage upper limit. Among them, the seventh network layer is the first network layer behind at least two first network layers among the multiple network layers, and in the forward propagation process, the seventh network layer is executed later than at least two first network layers. Among them, the eighth network layer is the first network layer in front of at least one second network layer among the multiple network layers, and in the backward propagation process, the eighth network layer is executed later than at least one second network layer.
[0060] In some alternative embodiments, for the second task group (where the first task packet includes forward propagation subtasks of at least one third network layer each, and the second task packet in the second task group includes backward propagation subtasks of at least two fourth network layers each), whether adding the forward propagation subtask of the ninth network layer to the first task packet or adding the backward propagation subtask of the tenth network layer to the second task packet will cause the total storage amount required when the first task packet and the second task packet in the second task group are executed to exceed the storage upper limit. Among them, the ninth network layer is the first network layer among the multiple network layers that is behind at least one third network layer, and during forward propagation, the ninth network layer is executed later than at least one third network layer. Among them, the tenth network layer is the first network layer among the multiple network layers that is in front of at least two fourth network layers, and during backward propagation, the tenth network layer is executed later than at least two fourth network layers.
[0061] Each of the first network layer to the tenth network layer mentioned above is any one of the multiple network layers.
[0062] In the embodiments of the present application, by making full use of the storage resources of the memory, the number of task groups is reduced, thereby reducing the number of task scheduling times. With the reduction of the number of task scheduling times, the computing resources occupied by task scheduling and the time consumed by task scheduling can be reduced, which helps to improve the training efficiency and reduce the training time.
[0063] In the above 103, according to the forward propagation computation amount and the backward propagation computation amount, the first total computation amount required when the first task packet is executed and the second total computation amount required when the second task packet is executed are determined. The specific calculation method can be set according to actual needs, and the embodiments of the present application do not make specific limitations on this.
[0064] In the above 104, among them, the preliminary scheduling plan is used to schedule one of the first task packet and the second task packet to the first CIM unit and the other to the second CIM unit.
[0065] The configuration information of the first CIM unit is used to represent the processing performance of the first CIM unit, and the configuration information of the second CIM unit is used to represent the processing performance of the second CIM unit.
[0066] In an alternative embodiment, when the second scheduling unit determines the preliminary scheduling plan and the no-load duration corresponding to the preliminary scheduling plan according to the first total computation amount, the second total computation amount, the configuration information of the first CIM unit, and the configuration information of the second CIM unit, it is specifically used to perform the following steps:
[0067] 1041. Determine the first processing duration required for the first CIM unit to process the first task package, the second processing duration required for the second CIM unit to process the second task package, the third processing duration required for the first CIM unit to process the second task package, and the fourth processing duration required for the second CIM unit to process the first task package according to the first total computing amount, the second total computing amount, the configuration information of the first CIM unit, and the configuration information of the second CIM unit.
[0068] 1042. Determine the delay and idle time corresponding to the first alternative scheduling plan and the delay and idle time corresponding to the second alternative scheduling plan according to the first processing duration, the second processing duration, the third processing duration, and the fourth processing duration.
[0069] Among them, the first alternative scheduling plan is used to allocate the first task package to the first CIM unit and the second task package to the second CIM unit; the second alternative scheduling plan is used to allocate the first task package to the second CIM unit and the second task package to the first CIM unit.
[0070] 1043. Determine the alternative scheduling plan with the minimum delay among the first alternative scheduling plan and the second alternative scheduling plan as the preliminary scheduling plan.
[0071] In the above 1041, the calculation of the processing duration can be calculated according to a preset calculation formula. The input of the preset calculation formula is the total computing amount of the task package and the configuration information of the CIM unit, and the output is the processing duration.
[0072] The above preset calculation formula can be designed according to actual needs, and the embodiments of the present application do not make specific limitations on this.
[0073] In the above 1042, among them, the delay corresponding to the first alternative scheduling plan is the maximum value of the above first processing duration and the second processing duration, and the idle time corresponding to the first alternative scheduling plan is the difference between the first processing duration and the second processing duration. The delay corresponding to the second alternative scheduling plan is the maximum value of the above third processing duration and the fourth processing duration, and the idle time corresponding to the second alternative scheduling plan is the difference between the above third processing duration and the fourth processing duration.
[0074] In the above 1043, determine the alternative scheduling plan with the minimum delay among the first alternative scheduling plan and the second alternative scheduling plan as the preliminary scheduling plan.
[0075] In the embodiments of the present application, the alternative scheduling plan with the minimum delay is determined as the preliminary scheduling plan, and then the preliminary scheduling plan is optimized later to shorten the no-load duration, which can not only improve the efficiency of model training but also improve the hardware utilization rate.
[0076] Of course, as an option, the alternative scheduling plan with the minimum no-load duration among the first alternative scheduling plan and the second alternative scheduling plan can be determined as the preliminary scheduling plan.
[0077] In the above 105, the no-load duration corresponding to the preliminary scheduling plan is compared with the preset upper limit of the no-load duration to determine whether the no-load duration corresponding to the preliminary scheduling plan is greater than or equal to the preset upper limit of the no-load duration.
[0078] If the no-load duration corresponding to the preliminary scheduling plan is greater than or equal to the preset upper limit of the no-load duration, then some tasks in the task packets responsible for the CIM units without no-load situations in the preliminary scheduling plan are scheduled to the CIM units with no-load situations to obtain the target scheduling plan.
[0079] In some embodiments, the task packets responsible for the CIM units with no-load situations in the preliminary scheduling plan include the propagation subtasks (forward propagation subtasks or backward propagation subtasks) of each network layer in X network layers. The propagation subtasks of Y network layers that are executed later in the task packets responsible for the CIM units with no-load situations can be split into a first part and a second part. Among them, the first part includes the partial propagation subtasks of each network layer in these Y network layers, and the second part includes the partial propagation subtasks of each network layer in these Y network layers. Wherein, X and Y are positive integers, and Y is less than X.
[0080] If the no-load duration corresponding to the preliminary scheduling plan is less than the preset upper limit of the no-load duration, then the preliminary scheduling plan is determined as the target scheduling plan, that is, it is not necessary to optimize or adjust the preliminary scheduling plan.
[0081] 106. Execute task scheduling according to the target scheduling plan.
[0082] Tasks are sent to the first CIM unit and the second CIM unit according to the target scheduling plan.
[0083] In the technical solution provided by the embodiments of the present application, the first CIM unit and the second CIM unit in the in-memory computing chip architecture are used to accelerate the model training process. The first scheduling unit splits and combines the forward propagation tasks and the backward propagation tasks to obtain multiple task groups. The second scheduling unit first generates a preliminary scheduling plan for the first task packet and the second task packet in each task group, and then, when the no-load duration corresponding to the preliminary scheduling plan is greater than or equal to the preset upper limit of the no-load duration, part of the tasks in the task packets responsible for the CIM units with no-load situations in the preliminary scheduling plan are scheduled to the CIM units without no-load situations to obtain a target scheduling plan. Through this processing, the second scheduling unit can shorten the no-load duration, thereby improving the hardware utilization rate during model training.
[0084] In some embodiments, as Figure 1 shown, the above architecture may further include: an expansion module 11 integrated in the pipeline of the central processing unit. The expansion module 11 includes the first CIM unit 11 and a sparse module 12.
[0085] The sparse module 12 is used to update the parameter sparse patterns of each network layer according to the parameter gradient distribution generated during the model training process, following the strategies of parameter discarding and parameter regeneration, and the sparse ratios corresponding to the parameter sparse patterns of each network layer remain unchanged before and after the update.
[0086] By the strategies of parameter discarding and parameter regeneration, it can be ensured that the sparse ratios corresponding to the parameter sparse patterns of each network layer remain unchanged before and after the update.
[0087] Among them, the parameter gradients may include weight gradients and bias gradients.
[0088] In an alternative embodiment, for each network layer, the multiple parameter gradients of the network layer are sorted, k parameters with the smallest parameter gradients are selected from the multiple non-zero parameters of the network layer, and these k parameters are set to zero, that is, parameter discarding. k parameters with the largest parameter gradients are selected from the multiple zero parameters of the network layer, and these k parameters are set to random non-zeros, that is, parameter regeneration.
[0089] In this embodiment, through the parameter loss and parameter regeneration strategies, the model accuracy or model performance can be improved under a certain sparse ratio.
[0090] In some embodiments, as Figure 1 shown, the expansion module 10 further includes a quantization module 13.
[0091] The quantization module 13 is configured to: determine a target floating-point format from multiple preset floating-point formats according to the data distribution of multiple parameters of each network layer; quantize the multiple parameters of the network layer according to the target floating-point format; wherein, the bit width of the floating-point format of the multiple parameters of each network layer before quantization is greater than the bit width of the preset floating-point format, and the multiple preset floating-point formats have the same bit width, different precisions, and different numerical representation ranges.
[0092] Exemplarily, the floating-point format of the multiple parameters of each network layer before quantization is BF16; the multiple preset floating-point formats include: E4M3 and E5M2. Among them, both E4M3 and E5M2 belong to FP8, where the precision of E4M3 is higher than that of E5M2, and the numerical representation range of E4M3 is smaller than that of E5M2.
[0093] In some optional embodiments, when the quantization module 13 determines the target data precision from two preset data precisions according to the data distribution of the multiple parameters of each network layer, it is specifically configured to: if the maximum value among the multiple parameters of each network layer exceeds the numerical representation range of E4M3, then determine E5M2 as the target floating-point format; otherwise, determine E4M3 as the target floating-point format.
[0094] In the embodiments of the present application, before quantizing the multiple parameters of the network layer, the data distribution of the multiple parameters will be considered, so as to select a more suitable quantization method. In this way, not only can the two indicators of the quantization degree and the model accuracy be well balanced.
[0095] Optionally, the quantization module 13 is further configured to: determine a target floating-point format from multiple preset floating-point formats according to the data distribution of the parameter gradients of each network layer; quantize the parameter gradients of the network layer according to the target data precision.
[0096] Optionally, the quantization module 13 is further configured to: determine a target floating-point format from multiple preset floating-point formats according to the data distribution of the activation values of each network layer; quantize the activation values of the network layer according to the target data precision.
[0097] It should be noted that the quantization process of the quantization module 13 for the parameter gradients and the activation values can refer to the quantization process of the quantization module 13 for the parameters in the above embodiments, and will not be elaborated here.
[0098] In some embodiments, the central processing unit 1 may include: a central processing unit based on the RISC-V instruction set architecture.
[0099] Optionally, there may be multiple first CIM units 11, and the first CIM units may be integrated inside the pipeline of the central processing unit 1.
[0100] In some embodiments, the above-mentioned memory 3 is used to store data such as parameters, activation values, and parameter gradients of the model.
[0101] In some embodiments, the memory 3 may be an embedded Dynamic Random Access Memory (eDRAM). Among them, the eDRAM can be implemented by 4 transistors. Among them, the eDRAM can be used as a cache to store intermediate calculation results during model training, such as: parameters, activation values, and parameter gradients of the model, thereby improving the calculation efficiency.
[0102] In some embodiments, the data stored in the memory 3 can be transmitted to the first CIM unit 11 or the second CIM unit 21 through a dedicated data link to increase the data bandwidth.
[0103] In some embodiments, the storage capacity of the memory 3 can be set according to actual needs, and the embodiments of the present application do not make specific limitations in this regard. Exemplarily, the storage capacity of the memory 3 can be 5 Mb.
[0104] Optionally, as Figure 1 shown, the above chip architecture may further include: a peripheral circuit 4, which includes a Static Random-Access Memory (SRAM). The SRAM is used to store information such as the running program of the CPU. The storage capacity of the SRAM can be set according to actual needs, and the embodiments of the present application do not make specific limitations in this regard. Exemplarily, the storage capacity of the SRAM is 128 KB.
[0105] Figure 2Shows a sparse training accelerator architecture using hybrid CIM (i.e., the chip architecture in the above text). The accelerator architecture consists of an open-source RISC-V CPU (i.e., the central processor 1 in the above text), a CIM-based NPU (i.e., the neural network model 2 in the above text), 5Mb eDRAM (i.e., the memory 3 in the above text), and 128KB SRAM (i.e., the peripheral circuit 4 in the above text). The pipelined tightly coupled CIM extension module (i.e., the extension module 10 in the above text) is located inside the CPU's pipeline as a coprocessor. This module contains two 128Kb SRAM CIM-C macrocells (i.e., the first CIM cell 11 in the above text) and an online fine-tuning cluster, which includes the above sparse module 12 and quantization module 13 for performing hierarchical sparsity and mixed-precision adjustment based on the runtime gradient distribution. A CIM workload scheduler (i.e., the load scheduler in the above text) is also designed in the CPU to parse custom RISC-V extension instructions and determine the workload (i.e., task) dispatched to hybrid CIM, that is, forward propagation or backward propagation on CIM-C or CIM-N. In the design of the CIM-based NPU, four 64Kb CIM-N macrocells (i.e., the second CIM cell 11 in the above text) are dedicated to AI acceleration. A CIM-aware floating-point quantization unit is also designed inside the NPU to reduce CIM computation cycles by dynamically pruning mantissas. Data is stored in eDRAM implemented by 4 transistors with a density 21.47% higher than SRAM and transmitted to hybrid CIM through a dedicated data link to increase data bandwidth. During model training, hybrid CIM can maintain a hardware utilization rate of >90%.
[0106] Figure 3 Describes the microarchitecture details of the pipelined tightly coupled CIM extension module. Each 128Kb CIM-C macrocell consists of a 512×256 SRAM array. An efficient sparse training process is achieved through the following four steps: forward propagation, backward propagation, weight update, and sparse adjustment. Sparse adjustment includes weight dropout and weight regeneration.
[0107] Considering sparsity and accuracy, a quantization module is also designed for layer-by-layer data precision adjustment (between BF16 and FP8). In addition, FP8 also provides two formats, E4M3 (higher precision) and E5M2 (wider range), and two data precisions are selected according to the data distribution. Through quantization adjustment, the memory footprint can be reduced, and the accuracy loss can be ignored.
[0108] Figure 4Describes the microarchitecture details of the CIM workload scheduler. The CIM workload scheduler consists of three parts: a storage configuration register (i.e., the configuration storage unit 121 mentioned above) for configuring the memory, an intra-layer allocator (i.e., the first scheduling unit 122 mentioned above), and an inter-layer scheduler (i.e., the second scheduling unit 123 mentioned above). The storage configuration register stores information about the structure of the model and sparsity. The CPU configures the model size, the forward pass sparsity lookup table, and the backward pass sparsity lookup table in this storage via the bus. Based on the lookup tables writable by the CPU, the storage configuration register infers the activation dimensions, the forward pass computational load (i.e., the forward pass computational amount mentioned above), and the backward pass computational load (i.e., the backward pass computational amount mentioned above). In addition, the CPU also needs to configure the available storage upper limit, the upper limit of the allowed CIM idle duration, and the type of the model. The data in the storage configuration register guides the scheduling of the forward and backward pass computational workloads. The inter-layer scheduler determines the workload blocks according to the data in the configuration register, including the propagation type (forward propagation or backward propagation) and the layer range (e.g., from the first layer to a certain intermediate layer), and generates a possible scheduling plan (i.e., Figure 3 the scheduling strategy shown) for the hybrid CIM and stores it in the inter-layer cache. The intra-layer scheduler selects a scheduling plan with less latency and further optimizes it through two workload fusion schemes to reduce the idle time. Workload fusion can dispatch some workloads to the idle CIM. Through such workload scheduling, the hardware utilization rate can be improved.
[0109] In summary, the hybrid CIM scheduling can effectively improve the CIM utilization rate, and the online sparsity and quantization adjustment can save memory with almost no loss of accuracy. Solutions such as the hybrid CIM architecture, online sparsity and quantization adjustment, and triple fixed data flow can effectively reduce the training time, thereby improving the system energy efficiency.
[0110] The embodiment of the present application also provides a memory computing chip, including: the memory computing chip architecture using sparsity described in the above embodiments.
[0111] In some embodiments, the chip is an intelligent edge chip.
[0112] It should also be noted that the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, commodity or device. Without further limitation, the element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, commodity or device including the element.
[0113] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. An in-memory computing chip architecture utilizing sparsity, characterized in that, Including: A central processing unit, which includes a first CIM unit and a load scheduler; A neural network processor, wherein the neural network processor includes: a second CIM unit; A memory for storing data related to model training; The load scheduler includes a configuration storage unit, a first scheduling unit, and a second scheduling unit; The configuration storage unit is used to store the structural information of the model, the sparsity configured for the model, and a preset upper limit of the no-load duration; the model includes multiple network layers; The configuration storage unit is further used to: determine the forward propagation calculation amount and forward propagation storage amount required for each network layer during forward propagation and the backward propagation calculation amount and backward propagation storage amount required for each network layer during backward propagation according to the structural information and the sparsity; The first scheduling unit is used to: split and combine the i-th forward propagation task and the (i - 1)-th backward propagation task of the model according to the forward propagation storage amount and the backward propagation storage amount to obtain multiple task groups; each task group includes a first task packet and a second task packet that can be processed in parallel; where i is an integer greater than 1; The second scheduling unit is used to determine the first total calculation amount of the first task packet and the second total calculation amount of the second task packet according to the forward propagation calculation amount and the backward propagation calculation amount; determine a preliminary scheduling plan and the no-load duration corresponding to the preliminary scheduling plan according to the first total calculation amount, the second total calculation amount, the configuration information of the first CIM unit, and the configuration information of the second CIM unit, where the preliminary scheduling plan is used to schedule one of the first task packet and the second task packet to the first CIM unit and the other to the second CIM unit; if the no-load duration corresponding to the preliminary scheduling plan is greater than or equal to the upper limit of the no-load duration, then schedule some tasks in the task packet responsible for the CIM unit without no-load situation in the preliminary scheduling plan to the CIM unit with no-load situation to obtain a target scheduling plan; perform task scheduling according to the target scheduling plan.
2. The architecture according to claim 1, characterized in that, When the second scheduling unit determines the preliminary scheduling plan and the no-load duration corresponding to the preliminary scheduling plan according to the first total calculation amount, the second total calculation amount, the configuration information of the first CIM unit, and the configuration information of the second CIM unit, it is specifically used for: Determine the first processing duration required for the first CIM unit to process the first task packet, the second processing duration required for the second CIM unit to process the second task packet, the third processing duration required for the first CIM unit to process the second task packet, and the fourth processing duration required for the second CIM unit to process the first task packet according to the first total calculation amount, the second total calculation amount, the configuration information of the first CIM unit, and the configuration information of the second CIM unit; Based on the first processing duration, the second processing duration, the third processing duration, and the fourth processing duration, determine the delay and idle duration corresponding to the first alternative scheduling plan and the delay and idle duration corresponding to the second alternative scheduling plan. The first alternative scheduling plan is used to allocate the first task package to the first CIM unit and the second task package to the second CIM unit; The second alternative scheduling plan is used to allocate the first task package to the second CIM unit and the second task package to the first CIM unit; Determine the alternative scheduling plan with the minimum delay among the first alternative scheduling plan and the second alternative scheduling plan as the preliminary scheduling plan.
3. The architecture according to claim 1 or 2, characterized in that, Further includes: An expansion module integrated in the pipeline of the central processing unit. The expansion module includes the first CIM unit and a sparse module; The sparse module is used to update the parameter sparse pattern of each network layer according to the parameter gradient distribution generated during the model training process and in accordance with the strategies of parameter discarding and parameter regeneration. The sparse ratio corresponding to the parameter sparse pattern of each network layer remains unchanged before and after the update.
4. The architecture according to claim 3, characterized in that, The expansion module further includes a quantization module; The quantization module is used to: determine a target floating-point format from multiple preset floating-point formats according to the data distribution of multiple parameters of each network layer; quantize the multiple parameters of this network layer according to the target floating-point format. Among them, the bit width of the floating-point format of the multiple parameters of each network layer before quantization is greater than the bit width of the preset floating-point format, and the bit widths of the multiple preset floating-point formats are the same, the precisions are different, and the numerical representation ranges are different.
5. The architecture according to claim 4, characterized in that, The multiple preset floating-point formats include E4M3 and E5M2; When the quantization module determines a target data precision from two preset data precisions according to the data distribution of multiple parameters of each network layer, it is specifically used to: If the maximum value among the multiple parameters of each network layer exceeds the numerical representation range of E4M3, then determine E5M2 as the target floating-point format; Otherwise, determine E4M3 as the target floating-point format.
6. The architecture according to claim 4, wherein The quantization module is further used to: determine a target floating-point format from multiple preset floating-point formats according to the data distribution of the parameter gradients of each network layer; quantize the parameter gradients of this network layer according to the target data precision.
7. The architecture according to claim 4, wherein The quantization module is further used to: determine a target floating-point format from multiple preset floating-point formats according to the data distribution of the activation values of each network layer; quantize the activation values of this network layer according to the target data precision.
8. The architecture according to claim 1 or 2, characterized in that, The configuration storage unit is further used to store the storage upper limit of the memory; When splitting and combining the i-th forward propagation task and the (i - 1)-th backward propagation task of the model according to the forward transmission storage amount and the backward transmission storage amount to obtain a plurality of task groups, the first scheduling unit specifically is configured to: split and combine the i-th forward propagation task and the (i - 1)-th backward propagation task of the model according to the forward transmission storage amount, the backward transmission storage amount, and the storage upper limit of the memory to obtain a plurality of task groups; the total storage amount required when the first task packet and the second task packet in each task group are executed is less than the storage upper limit; the number of the plurality of task groups is less than the number of the plurality of network layers; The plurality of task groups include a first task group and / or a second task group; Wherein, the first task packet in the first task group includes forward propagation subtasks of at least two first network layers respectively, and the second task packet in the first task group includes backward propagation subtasks of at least one second network layer respectively; The first task packet in the second task group includes forward propagation subtasks of at least one third network layer respectively, and the second task packet in the second task group includes backward propagation subtasks of at least two fourth network layers respectively.
9. An in-memory computing chip, characterized in that, Comprising: The in-memory computing chip architecture using sparsity according to any one of claims 1 to 8 above.
10. The chip according to claim 9, characterized in that, The chip is an intelligent edge chip.