Large model acceleration system

By combining the large model acceleration chip and near-memory computing chip in the large model acceleration system, unified scheduling of computing tasks and data in the large model decoding stage is achieved, solving the problem of low resource utilization in the existing technology, and improving the computing efficiency and hardware resource utilization rate.

CN120104559APending Publication Date: 2025-06-06SHANGHAI INFINIGENCE AI INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510165707.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing technology has not developed a large-model acceleration system for near-existing computing scenarios, resulting in a lack of unified scheduling of computing tasks and computing data, making it difficult to adapt to large models and computing models of different scales, limiting the utilization rate of hardware resources.

Method used

A large-model acceleration system is provided, including a large-model acceleration chip and multiple near-memory computing chips. The large-model acceleration chip is responsible for the preprocessing of vectors and the distribution and scheduling of computing instructions. The near-memory computing chip performs matrix vector multiplication calculation tasks and sends the results back to the large-model acceleration chip.

Benefits of technology

By uniformly scheduling computing tasks and data during the decoding stage, the delay and power loss caused by data transmission are reduced, the computing efficiency of the large model is improved, and hardware resources are fully utilized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104559A_ABST
    Figure CN120104559A_ABST
Patent Text Reader

Abstract

The invention provides a large model acceleration system, comprising: a large model acceleration chip configured to preprocess an input vector in a decoding stage of a large model, distribute the preprocessed vector to a near memory computing chip, and send a computing instruction to the near memory computing chip, the near memory computing chip executes a matrix vector multiplication computing task; and a plurality of near memory computing chips, each near memory computing chip is configured to execute a matrix vector multiplication computing task based on the distributed vector and the stored weight matrix of the large model according to the computing instruction, and send a computing result to the large model acceleration chip.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of artificial intelligence technology, and specifically relates to a large model acceleration system. Background Art

[0002] Since the Transformer neural network algorithm was proposed, neural network algorithms based on encoder and decoder structures have shown strong performance in natural language processing, intelligent content generation and other fields. In the large model based on the Transformer algorithm, the decoding stage mainly performs matrix-vector multiplication (MVM) calculations of vector and matrix multiplication. The vector of the calculation object is obtained dynamically by processing the input data, and the matrix is ​​usually a parameter stored in the memory, such as a weight matrix. Since the calculation time of the decoding stage is usually much higher than that of the encoding stage, it is possible to consider accelerating the calculation in the decoding stage. Summary of the invention

[0003] In the acceleration system for the decoding stage of large models known to the inventors of this application, the computing performance of matrix-vector multiplication is improved mainly by optimizing the computing efficiency on the computing chip or by increasing the memory access bandwidth of the computing chip. Such an acceleration system usually still adopts an architecture in which the computing unit is separated from the memory. In the large model scenario, the weight matrix is ​​huge, and frequent large-scale data handling exacerbates the problems of high latency and high energy consumption. In addition, the nonlinear processing before and after the matrix-vector multiplication (such as activation function, normalization, etc.) and the inefficiency of computing task scheduling are inefficient, further limiting the system performance.

[0004] In the related art, near-memory computing can reduce data handling overhead, delays and power consumption caused by data transmission by deploying computing units close to the memory. However, in the related art, no corresponding large model acceleration system has been developed for near-memory computing scenarios, resulting in a lack of unified scheduling of computing tasks and computing data, making it difficult to adapt to large models and computing modes of different sizes, and limiting the utilization of hardware resources.

[0005] In response to the problems existing in the prior art, the present application provides a large model acceleration system to solve the problem that the relevant technology has not developed a corresponding large model acceleration system for the near-memory computing scenario, so that while trying to avoid the delay and power loss caused by data transmission, the computing tasks and computing data of the decoding stage of the large model are uniformly scheduled, the computing efficiency is improved, and hardware resources are fully utilized.

[0006] In a first aspect, an embodiment of the present application provides a large model acceleration system, including:

[0007] A large model acceleration chip is configured to preprocess an input vector in a decoding stage of the large model, distribute the preprocessed vector to a near-memory computing chip, and send a computing instruction to the near-memory computing chip so that the near-memory computing chip performs a matrix-vector multiplication computing task; and

[0008] Multiple near-memory computing chips, each near-memory computing chip is configured to perform matrix-vector multiplication calculation tasks according to the calculation instructions, based on the distributed vectors and the stored weight matrix of the large model, and send the calculation results to the large model acceleration chip.

[0009] Furthermore, the near memory computing chip includes:

[0010] A memory configured to store a weight matrix of a large model; and

[0011] The calculation part includes:

[0012] a plurality of variable accumulation channels, each variable accumulation channel comprising a plurality of fixed accumulation groups and a first variable accumulator, and each variable accumulation channel corresponds to a memory read channel of the weight matrix, wherein the first variable accumulator selectively accumulates a plurality of primary accumulation results generated by performing matrix-vector multiplication calculations on the plurality of fixed accumulation groups according to a predetermined accumulator configuration to obtain a first accumulation result; and

[0013] The second variable accumulator selectively accumulates the multiple first accumulation results generated by the multiple variable accumulation channels according to a predetermined accumulator configuration to obtain a second accumulation result, and sends the second accumulation result as a calculation result of the near-memory computing chip to the large model acceleration chip.

[0014] Further, each of the plurality of fixed accumulation groups includes:

[0015] a plurality of multipliers, each multiplier being configured to perform a multiplication calculation on an element of the vector and an element of the weight matrix; and

[0016] The fixed accumulator accumulates the multiplication calculation results of the plurality of multipliers to obtain the primary accumulation result.

[0017] In addition, the large model acceleration chip may include:

[0018] a nonlinear computing unit configured to perform the pre-processing before the matrix-vector multiplication calculation and the post-processing after the matrix-vector multiplication calculation on the vector in the decoding stage;

[0019] An operation controller is configured to control the scheduling of computing tasks and computing data using computing instructions, wherein the computing tasks include computing in a nonlinear computing unit and computing in a near-memory computing chip, and the computing data includes the vector and the weight matrix;

[0020] A near-memory computing configuration memory configured to store configuration information included in a computing instruction from an operation controller, the configuration information including configuration information indicating a predetermined accumulator configuration and configuration information indicating how to distribute the computing data to near-memory computing chips;

[0021] A near memory chip access controller is configured to read the configuration information from the near memory computing configuration memory and control the sending of the computing data to the near memory computing chip according to the configuration information;

[0022] A matrix write cache is configured to send the weight matrix to the near memory computing chip under the control of the near memory chip access controller before the calculation of the decoding stage starts, so that the weight matrix is ​​written into the memory of the near memory computing chip; and

[0023] The vector cache is configured to cache the vector before the vector is sent to the near memory computing chip, and send the vector to the near memory computing chip under the control of the near memory chip access controller.

[0024] Furthermore, the large model acceleration chip also includes:

[0025] The vector part and the accumulation unit are configured to accumulate the calculation results of the near-memory calculation chip to obtain a third accumulation result when the calculation results of the near-memory calculation chip still need to be accumulated, and send the third accumulation result to the nonlinear calculation unit for post-processing; and

[0026] The vector part and cache are configured to cache the third accumulation result generated by the vector part and the accumulation unit.

[0027] Furthermore, the large model acceleration chip also includes:

[0028] The matrix loading flow control unit is configured to receive a signal from the near memory computing chip, wherein the signal indicates whether the near memory computing chip can receive a new computing instruction to perform a new matrix-vector multiplication computing task.

[0029] When the signaling indicates that the near memory computing chip can receive new computing instructions, the matrix loading flow control unit notifies the near memory chip access controller so that the near memory chip access controller controls the matrix write cache to send the weight matrix to the near memory computing chip.

[0030] In addition, during the decoding stage, the large model acceleration chip and the near-memory computing chip can perform calculations of the nonlinear computing unit and matrix-vector multiplication calculations alternately in the first thread and the second thread, respectively.

[0031] In a second aspect, an embodiment of the present application provides a large model acceleration chip, which is configured to preprocess an input vector in a decoding stage of a large model, distribute the preprocessed vector to a near-memory computing chip, and send a calculation instruction to the near-memory computing chip so that the near-memory computing chip performs a matrix-vector multiplication calculation task, wherein the large model acceleration chip includes:

[0032] a nonlinear computing unit configured to perform the pre-processing before the matrix-vector multiplication calculation and the post-processing after the matrix-vector multiplication calculation on the vector in the decoding stage;

[0033] An operation controller is configured to control the scheduling of computing tasks and computing data using computing instructions, wherein the computing tasks include computing in the nonlinear computing unit and computing in the near-memory computing chip, and the computing data includes the vector and a weight matrix of the large model;

[0034] A near-memory computing configuration memory configured to store configuration information included in a computing instruction from an operation controller, the configuration information including configuration information indicating a predetermined accumulator configuration and configuration information indicating how to distribute the computing data to near-memory computing chips;

[0035] A near memory chip access controller is configured to read the configuration information from the near memory computing configuration memory and control the sending of the computing data to the near memory computing chip according to the configuration information;

[0036] A matrix write cache is configured to send the weight matrix to the near memory computing chip under the control of the near memory chip access controller before the calculation of the decoding stage starts, so that the weight matrix is ​​written into the memory of the near memory computing chip; and

[0037] The vector cache is configured to cache the vector before the vector is sent to the near memory computing chip, and send the vector to the near memory computing chip under the control of the near memory chip access controller.

[0038] In a third aspect, an embodiment of the present application provides a near-memory computing chip, which is configured to perform a matrix-vector multiplication calculation task based on the distributed vectors and the stored weight matrix of the large model in accordance with a calculation instruction from the large model acceleration chip during the decoding stage of the large model, and send the calculation result to the large model acceleration chip, wherein the near-memory computing chip includes:

[0039] A memory configured to store a weight matrix of a large model; and

[0040] The calculation part includes:

[0041] a plurality of variable accumulation channels, each variable accumulation channel comprising a plurality of fixed accumulation groups and a first variable accumulator, and each variable accumulation channel corresponds to a memory read channel of the weight matrix, wherein the first variable accumulator selectively accumulates a plurality of primary accumulation results generated by performing matrix-vector multiplication calculations on the plurality of fixed accumulation groups according to a predetermined accumulator configuration to obtain a first accumulation result; and

[0042] The second variable accumulator selectively accumulates the multiple first accumulation results generated by the multiple variable accumulation channels according to a predetermined accumulator configuration to obtain a second accumulation result, and sends the second accumulation result as a calculation result of the near-memory computing chip to the large model acceleration chip.

[0043] Further, each of the plurality of fixed accumulation groups includes:

[0044] a plurality of multipliers, each multiplier being configured to perform a multiplication calculation on an element of the vector and an element of the weight matrix; and

[0045] The fixed accumulator accumulates the multiplication calculation results of the plurality of multipliers to obtain the primary accumulation result.

[0046] Other optional features and technical effects of the embodiments of the present application are partially described below, and partially can be understood by reading this document.

[0047] Compared with the related art, this application has the following beneficial technical effects:

[0048] A large model acceleration system with a large model acceleration chip and multiple near-memory computing chips is provided. The large model acceleration chip is used for the distribution and scheduling of computing instructions and computing data, and the multiple near-memory computing chips are used for matrix-vector multiplication calculations. An acceleration system for the decoding stage of large models is developed for near-memory computing scenarios, which tries to avoid delays and power losses caused by data transmission, improves the computing efficiency of the decoding stage of large models, and makes full use of computing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with specific implementation methods and drawings. Here, the illustrative implementation methods and descriptions of the present application are used to explain the present application, but are not intended to limit the present application.

[0050] As used herein, the term "including" and its variations mean open inclusion, i.e., "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "based at least in part on". The terms "an example embodiment" and "an embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0051] Figure 1 A schematic diagram of the structure of a large model acceleration system according to an embodiment of the present application is shown;

[0052] Figure 2 A schematic diagram of the structure of a near-memory computing chip and a large model acceleration chip according to an embodiment of the present application is shown;

[0053] Figure 3 A schematic diagram of the structure of a large model acceleration unit of a large model acceleration chip according to an embodiment of the present application is shown;

[0054] Figure 4 A schematic diagram of multi-thread scheduling according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0055] The embodiments of the present application are described in detail below with reference to the accompanying drawings.

[0056] The following describes the implementation methods of the present application through specific specific examples, and those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The present application can also be implemented or applied through other different specific implementation methods, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that, in the absence of conflict, the following embodiments and the features in the embodiments can be combined with each other. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work belong to the scope of protection of the present application.

[0057] It should be noted that various aspects of the embodiments within the scope of the appended claims are described below. It should be apparent that the aspects described herein may be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on the present application, it should be understood by those skilled in the art that an aspect described herein may be implemented independently of any other aspect, and two or more of these aspects may be combined in various ways. For example, any number of aspects described herein may be used to implement the device and / or practice the method. In addition, other structures and / or functionalities other than one or more of the aspects described herein may be used to implement this device and / or practice this method.

[0058] It should also be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present application. The drawings only show components related to the present application rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.

[0059] Additionally, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, it will be understood by those skilled in the art that the aspects described may be practiced without these specific details.

[0060] Figure 1 A schematic diagram of the structure of a large model acceleration system according to an embodiment of the present application. Figure 1 As shown, the large model acceleration system according to an embodiment of the present application includes a large model acceleration chip and multiple near-memory computing chips.

[0061] The large model acceleration chip is configured to perform preprocessing such as layer normalization (LayerNorm) on the input vector (e.g., the vector obtained after predetermined processing of the text data) during the decoding stage of the large model, distribute the preprocessed vector to the near-memory computing chip, and send a calculation instruction to the near-memory computing chip so that the near-memory computing chip performs the matrix-vector multiplication calculation task. It should be understood that in this application, the decoding stage of the large model refers to the calculation stage performed using the decoder structure during the training or reasoning process of the large model, and the large model is not limited to the large language model based on Transformer, but can be any large model for any purpose including the decoder structure.

[0062] Each of the multiple near-memory computing chips is configured to perform matrix-vector multiplication calculation tasks in accordance with calculation instructions from the large model acceleration chip, based on the distributed vectors and the stored weight matrix of the large model (written by the large model acceleration chip before the calculation starts), and send the calculation results to the large model acceleration chip.

[0063] In some embodiments, the near-memory computing chip may include a memory configured to store a weight matrix of a large model. Figure 2 As shown in the structural schematic diagram of the near-memory computing chip and the large model acceleration chip according to the embodiment of the present application, the near-memory computing chip may also include a computing part, and the computing part includes multiple variable accumulation channels, such as Figure 2 In the variable accumulation channel 0 to variable accumulation channel M. In addition, each variable accumulation channel can correspond to a memory read channel of the weight matrix, for example Figure 2 The variable accumulation channels 0 to 1 in the embodiment may correspond to the memory read channels 0 to 1 of the weight matrix. As shown in the variable accumulation channel 0, each variable accumulation channel may include, for example Figure 2 Multiple fixed accumulation groups of fixed accumulation group 0 to fixed accumulation group N and a first-level variable accumulator (first variable accumulator) in the fixed accumulation group 0. As shown in fixed accumulation group 0, each fixed accumulation group may include, for example, multiple multipliers of multiplier 0 to multiplier K (K can be set to, for example, 4, 8, 16, etc. as needed) and a fixed accumulator of level 0 (fixed accumulator). Each multiplier can be configured to perform multiplication calculations on an element of a vector and an element of a weight matrix. Specifically, after receiving the vector to be processed, the near-memory computing chip will place the vector element in an input port register of the corresponding multiplier according to the configuration information (described later). After receiving the calculation instruction, the near-memory computing chip continuously reads the stored weight matrix from its memory according to the calculation instruction, and sends the matrix element to another port of the multiplier, so that the multiplier performs multiplication calculations on the matrix element and the vector element stored in the input port register. After the multiplication calculation is completed, the multiplier sends the multiplication calculation result to the fixed accumulator of level 0. The fixed accumulator at level 0 accumulates the multiplication calculation results of all the multipliers in the fixed accumulation group, i.e., multipliers 0 to K, to obtain the primary accumulation result. Since the vector scale is often large in the matrix-vector multiplication calculation of the large model, by using the fixed accumulator to accumulate the multiplication results in the fixed accumulation group and then output them, it is unnecessary to configure any additional accumulator for the fixed accumulation group, thus simplifying the complexity of the accumulator configuration.

[0064] In each variable accumulation channel, each of the multiple fixed accumulation groups, fixed accumulation group 0 to fixed accumulation group N, sends the primary accumulation result obtained by performing matrix-vector multiplication calculation to the first-stage variable accumulator. The first-stage variable accumulator selectively accumulates multiple primary accumulation results of multiple fixed accumulation groups according to a predetermined accumulator configuration to obtain a first accumulation result. Figure 2As shown, the calculation part of the near memory computing chip also includes a second-level variable accumulator (second variable accumulator). A total of M+1 first accumulation results of multiple variable accumulation channels, namely variable accumulation channel 0 to variable accumulation channel M, are sent to the second-level variable accumulator. The second-level variable accumulator selectively accumulates multiple first accumulation results generated by multiple variable accumulation channels according to a predetermined accumulator configuration to obtain a second accumulation result, and sends the second accumulation result as the calculation result of the near memory computing chip to the large model acceleration chip.

[0065] The core processing module of the large model acceleration chip is the large model acceleration unit, such as Figure 3 As shown, the large model acceleration unit may include a nonlinear unit (nonlinear computing unit), an operation controller, a near memory computing configuration memory, a near memory chip access controller, a matrix write cache, a vector cache and other components.

[0066] The nonlinear unit is configured to perform preprocessing on the input vector (e.g., a vector obtained after predetermined processing of text data) before matrix-vector multiplication calculation and post-processing after matrix-vector multiplication calculation in the decoding stage of the large model, such as layer normalization (LayerNorm), exponential normalization (Softmax), etc.

[0067] The operation controller is configured to use calculation instructions to control the scheduling of calculation tasks and calculation data. The calculation tasks include calculations in nonlinear units (the aforementioned preprocessing and postprocessing) and calculations in near-memory computing chips (matrix-vector multiplication), and the calculation data include vectors and weight matrices. The calculation data may, for example, also include vector parts and intermediate results as described later. In other words, the operation controller generates calculation instructions and sends them to the corresponding components of the large model acceleration chip and the near-memory computing chip, so that each component accesses the calculation data and performs the corresponding calculation tasks. By setting the operation controller to uniformly schedule the calculation tasks and calculation data of the large model acceleration chip and the near-memory computing chip, the large model acceleration system provided in the embodiment of the present application can adapt to large models and calculation modes of different sizes, and support any number of near-memory computing chips.

[0068] The near memory computing configuration memory is configured to store configuration information included in the computing instructions from the operation controller. The configuration information may include, for example, configuration information indicating the predetermined accumulator configuration as described above and configuration information indicating how to distribute the computing data to the near memory computing chips (which computing data is distributed to which near memory computing chips).

[0069] The near memory chip access controller is configured to read configuration information from the near memory computing configuration memory, control the sending of computing data to the near memory computing chip according to the configuration information, and send the computing instructions from the operation controller to the corresponding near memory computing chip. In other words, the near memory chip access controller is used to schedule the sending of computing data and computing instructions to the near memory computing chip.

[0070] The matrix write cache is configured to cache the weight matrix before it is sent to the near memory computing chip, and to send the weight matrix to the near memory computing chip under the control of the near memory chip access controller before the calculation of the decoding stage starts, so that the weight matrix is ​​written into the memory of the near memory computing chip.

[0071] The vector cache is configured to cache the vector before the vector is sent to the near memory computing chip, and send the vector to the near memory computing chip under the control of the near memory chip access controller at appropriate timing.

[0072] As mentioned above, the second variable accumulator ( Figure 2 After obtaining the second accumulated result in the matrix-vector multiplication calculation, the second accumulated result is sent to the large model acceleration chip as the calculation result of the near-memory computing chip. According to the predetermined accumulator configuration, such a calculation result may be the final calculation result of the matrix-vector multiplication calculation, or it may be an intermediate calculation result (vector partial sum) that needs further accumulation. Figure 3 As shown, the large model acceleration unit of the large model acceleration chip may further include a vector part and an accumulation unit, which may be configured to accumulate the calculation results of the near-memory calculation chip when the calculation results of the near-memory calculation chip still need to be accumulated, to obtain a third accumulation result, and to send the third accumulation result to the nonlinear calculation unit ( Figure 3 The vector part and the accumulation unit may correspond to Figure 2 The third-level on-chip accumulator in the large model acceleration chip in the large model acceleration chip. Corresponding to the vector part and the accumulation unit, a vector part and a cache can be set in the large model acceleration unit of the large model acceleration chip to cache the third accumulation result generated by the vector part and the accumulation unit. When the third accumulation result still needs to be further accumulated with the second accumulation result from the near-memory computing chip, the vector part and the accumulation unit can read the cached third accumulation result from the vector part and the cache, and further accumulate it with the newly received second accumulation result. Alternatively, when the calculation is completed and no further accumulation is required, the vector part and the cache can send the final third accumulation result as the final calculation result of the matrix-vector multiplication to the nonlinear computing unit for post-processing.

[0073] Since the large model acceleration chip and the near memory computing chip are connected through the inter-chip interconnection, in order to ensure that the instructions received by the near memory computing chip will not overflow, the near memory chip access controller can use a credit-based flow control mechanism when sending computing instructions to the near memory computing chip. Figure 3 As shown, the large model acceleration unit of the large model acceleration chip may also include a matrix loading flow control unit. The matrix loading flow control unit may be configured to receive a signal from a near-memory computing chip, which indicates whether the near-memory computing chip can receive a new calculation instruction to perform a new matrix-vector multiplication calculation task. That is, when the near-memory computing chip has sufficient storage space to receive a new weight matrix and has sufficient computing resources to start a new calculation, it can send a signal to the matrix loading flow control unit of the large model acceleration chip indicating that the near-memory computing chip can receive a new calculation instruction to perform a new matrix-vector multiplication calculation task. At this time, that is, when the signal indicates that the near-memory computing chip can receive a new calculation instruction, the matrix loading flow control unit notifies the near-memory chip access controller so that the near-memory chip access controller controls the matrix write cache to send the weight matrix to the near-memory computing chip. Subsequently, after the vector to be processed is input into the large model acceleration chip and the nonlinear computing unit of the large model acceleration chip preprocesses the vector, the near-memory chip access controller sends the calculation instruction from the operation controller to the corresponding near-memory computing chip so that the near-memory computing chip starts to perform a new matrix-vector multiplication calculation task. When the near memory computing chip has started to execute the matrix-vector multiplication calculation task, a signal can be sent to the matrix loading flow control unit to indicate that the near memory computing chip cannot receive new calculation instructions to execute new matrix-vector multiplication calculation tasks, so as to ensure that the instructions in the near memory computing chip will not overflow.

[0074] In some embodiments, during the decoding stage of the large model, the large model acceleration chip and the near-memory computing chip may perform calculations of the nonlinear computing unit (the aforementioned preprocessing and postprocessing) and matrix-vector multiplication calculations alternately in the first thread and the second thread, respectively.

[0075] Specifically, in the related art, in the calculation of the decoding stage of the large model, nonlinear calculations such as preprocessing and postprocessing as described above need to be performed before and after a matrix-vector multiplication calculation. When performing nonlinear calculations, the calculation unit used for matrix-vector multiplication calculations has no calculation tasks and will be idle for a period of time. Figure 4 As shown in the upper figure, there is a time gap between the two matrix-vector multiplication calculations, which wastes hardware resources. In order to further improve the computing efficiency of the large model acceleration system, the embodiment of the present application provides a multi-threaded scheduling method to improve the computing performance of the system. Figure 4As shown in the figure below, there are two computing threads in the large model acceleration system. When the near memory computing chip is performing matrix-vector multiplication calculations in the first thread, thread 0, the large model acceleration chip can perform calculations of nonlinear computing units in the second thread, thread 1. Next, when the near memory computing chip completes the matrix-vector multiplication calculations in the first thread, thread 0, it can immediately perform the matrix-vector multiplication calculations in the second thread, thread 1. At this time, the large model acceleration chip that has completed the calculation of the nonlinear computing unit can also immediately perform the calculation of the nonlinear computing unit in the first thread, thread 0. Such multi-threaded scheduling can also be executed by the operation control unit of the large model acceleration chip, for example. From Figure 4 It can be seen that the running time of completing 4 matrix-vector multiplication calculations in the case of multi-thread scheduling in the lower figure is shorter than the running time of single-thread scheduling in the upper figure. Therefore, by making the large model acceleration chip and the near-memory computing chip alternately perform the calculation of the nonlinear computing unit and the matrix-vector multiplication calculation in the first thread and the second thread respectively, the hardware computing resources of the large model acceleration chip and the near-memory computing chip can be fully utilized to further improve the computing efficiency.

[0076] To facilitate understanding of the collaboration between the components of the large model acceleration chip and the near memory computing chip, a specific example of a matrix-vector multiplication calculation is given below. First, before the calculation starts, the near memory chip access controller of the large model acceleration chip receives signaling from the near memory computing chip. When the signaling indicates that the near memory computing chip can receive new calculation instructions to perform new matrix-vector multiplication calculation tasks, the near memory chip access controller controls the matrix write cache to write the weight matrix to the memory of the corresponding near memory computing chip according to the configuration information included in the calculation instructions from the operation controller stored in the near memory computing configuration memory. When the vector to be processed is input to the large model acceleration chip, the operation controller controls the nonlinear calculation unit to preprocess the vector and sends the preprocessed vector to the vector cache. The operation controller controls the near memory chip access controller to control the vector cache to distribute the vector to the corresponding near memory computing chip according to the dimension of the vector and the distribution of the weight matrix in the memory of the near memory computing chip, and then sends the calculation instruction to the near memory computing chip through the near memory chip access controller.

[0077] After receiving the vector to be processed, the near-memory computing chip will place the vector elements in an input port register of the multiplier of the fixed accumulation group of the corresponding variable accumulation channel according to the configuration information. After receiving the calculation instruction, the near-memory computing chip continuously reads the stored weight matrix from its memory according to the calculation instruction, and sends the matrix elements to another port of the multiplier for multiplication calculation. After the multiplication calculation is completed, the result data will pass through the three-level accumulator of the near-memory computing chip, namely the 0th level fixed accumulator, the 1st level variable accumulator, and the 2nd level variable accumulator, to obtain the matrix-vector multiplication calculation result of the near-memory computing chip.

[0078] Next, the near-memory computing chip sends the calculation result to the large model acceleration chip. Assuming that the calculation result still needs to be further accumulated, the calculation result is sent to the vector part and cache of the large model acceleration chip to wait for other calculation results from the near-memory computing chip. When the large model acceleration chip receives other calculation results to be accumulated from the near-memory computing chip, the vector part and the accumulation unit take out the calculation results in the vector part and the cache and accumulate them with the newly received other calculation results. The vector part and the accumulation unit send the accumulated results obtained to the nonlinear computing unit for post-processing. After the nonlinear computing unit performs post-processing on the vector, it determines whether the current decoding calculation task is completed. If the calculation is judged to be completed, the result data is sent back to the CPU; if the calculation is judged not to be completed, the result data is sent to the vector cache to wait for the next matrix-vector multiplication calculation. After receiving the vector, the vector cache will send the vector to the corresponding near-memory computing chip under the control of the near-memory chip access controller when the signaling of the near-memory computing chip indicates that the near-memory computing chip can perform a new matrix-vector multiplication calculation, and perform the next matrix-vector multiplication calculation.

[0079] It should be noted that in Figure 2 and Figure 3 In the structural diagram in FIG. 1 , not all connections between components are shown. Components in the large model acceleration chip and the near memory computing chip should be considered to be connected when there is data / instruction interaction.

[0080] The systems, devices, modules or units described in the above embodiments may be implemented by a computer or its associated components. The computer may be, for example, a mobile terminal, a smart phone, a personal computer, a laptop computer, a vehicle-mounted human-computer interaction device, a personal digital assistant, a media player, a navigation device, a game console, a tablet computer, a wearable device, a smart TV, an Internet of Things system, a smart home, an industrial computer, a server or a combination thereof.

[0081] Although not shown, in an embodiment of the present application, a computer-readable storage medium is provided, on which a computer program / instruction is stored. When the computer program / instruction is executed by a processor, the method described in the embodiment is implemented.

[0082] The storage media in the embodiments of the present application include permanent and non-permanent, removable and non-removable items that can be used to store information by any method or technology. Examples of storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission medium that can be used to store information that can be accessed by a computing device.

[0083] Although not shown, the embodiments of the present application also provide a computer program product, including: a computer program / instruction, which implements the method described in the embodiments when the computer program / instruction is executed by a processor.

[0084] The methods, programs, systems, devices, etc. of the embodiments of the present application may be executed or implemented in a single or multiple networked computers, or may be practiced in a distributed computing environment. In the embodiments of the present specification, in these distributed computing environments, tasks may be performed by remote processing devices connected via a communication network.

[0085] Those skilled in the art should understand that the embodiments of the present specification can be provided as methods, systems or computer program products. Therefore, those skilled in the art can imagine that the implementation of the functional modules / units or controllers and related method steps described in the above embodiments can be implemented in software, hardware or a combination of software / hardware.

[0086] Unless explicitly stated, the actions or steps of the methods, programs, and embodiments of the present application do not have to be performed in a specific order and can still achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0087] In this article, multiple embodiments of the present application are described, but for the sake of simplicity, the description of each embodiment is not exhaustive, and the same or similar features or parts between the various embodiments may be omitted. In this article, "one embodiment", "some embodiments", "example", "specific example", or "some examples" are intended to be applicable to at least one embodiment or example according to the present application, rather than all embodiments. The above terms do not necessarily mean to refer to the same embodiment or example. In the absence of mutual contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples.

[0088] The exemplary systems and methods of the present application have been specifically shown and described with reference to the above embodiments, which are merely examples of the best modes for implementing the present systems and methods. It will be appreciated by those skilled in the art that various changes may be made to the embodiments of the systems and methods described herein when implementing the present systems and / or methods without departing from the spirit and scope of the present application as defined in the appended claims.

Claims

1. A large model acceleration system, characterized in that: include: The large model acceleration chip is configured to preprocess the input vector in the decoding stage of the large model, distribute the preprocessed vector to the near-memory computing chip, and send a computing instruction to the near-memory computing chip so that the near-memory computing chip performs a matrix-vector multiplication computing task; as well as Multiple near-memory computing chips, each near-memory computing chip is configured to perform matrix-vector multiplication calculation tasks according to the calculation instructions, based on the distributed vectors and the stored weight matrix of the large model, and send the calculation results to the large model acceleration chip.

2. The large model acceleration system according to claim 1, characterized in that: The near-memory computing chip comprises: A memory configured to store a weight matrix of a large model; and The calculation part includes: a plurality of variable accumulation channels, each variable accumulation channel comprising a plurality of fixed accumulation groups and a first variable accumulator, and each variable accumulation channel corresponds to a memory read channel of the weight matrix, wherein the first variable accumulator selectively accumulates a plurality of primary accumulation results generated by performing matrix-vector multiplication calculations on the plurality of fixed accumulation groups according to a predetermined accumulator configuration to obtain a first accumulation result; and The second variable accumulator selectively accumulates the multiple first accumulation results generated by the multiple variable accumulation channels according to a predetermined accumulator configuration to obtain a second accumulation result, and sends the second accumulation result as a calculation result of the near-memory computing chip to the large model acceleration chip.

3. The large model acceleration system according to claim 2, characterized in that: Each of the plurality of fixed accumulation groups comprises: a plurality of multipliers, each multiplier being configured to perform a multiplication calculation on an element of the vector and an element of the weight matrix; and The fixed accumulator accumulates the multiplication calculation results of the plurality of multipliers to obtain the primary accumulation result.

4. The large model acceleration system according to any one of claims 1 to 3, characterized in that: The large model acceleration chip includes: a nonlinear computing unit configured to perform the pre-processing before the matrix-vector multiplication calculation and the post-processing after the matrix-vector multiplication calculation on the vector in the decoding stage; An operation controller is configured to control the scheduling of computing tasks and computing data using computing instructions, wherein the computing tasks include computing in a nonlinear computing unit and computing in a near-memory computing chip, and the computing data includes the vector and the weight matrix; A near-memory computing configuration memory configured to store configuration information included in a computing instruction from an operation controller, the configuration information including configuration information indicating a predetermined accumulator configuration and configuration information indicating how to distribute the computing data to near-memory computing chips; a near memory chip access controller configured to read the configuration information from the near memory computing configuration memory, control the sending of the computing data to the near memory computing chip according to the configuration information, and send the computing instruction from the operation controller to the corresponding near memory computing chip; A matrix write cache is configured to send the weight matrix to the near memory computing chip under the control of the near memory chip access controller before the calculation of the decoding stage starts, so that the weight matrix is ​​written into the memory of the near memory computing chip; and The vector cache is configured to cache the vector before the vector is sent to the near memory computing chip, and send the vector to the near memory computing chip under the control of the near memory chip access controller.

5. The large model acceleration system according to claim 4, characterized in that: The large model acceleration chip also includes: The vector part and the accumulation unit are configured to accumulate the calculation results of the near-memory calculation chip to obtain a third accumulation result when the calculation results of the near-memory calculation chip still need to be accumulated, and send the third accumulation result to the nonlinear calculation unit for post-processing; and The vector part and cache are configured to cache the third accumulation result generated by the vector part and the accumulation unit.

6. The large model acceleration system according to claim 5, characterized in that: The large model acceleration chip also includes: The matrix loading flow control unit is configured to receive a signal from the near memory computing chip, wherein the signal indicates whether the near memory computing chip can receive a new computing instruction to perform a new matrix-vector multiplication computing task. When the signaling indicates that the near memory computing chip can receive new computing instructions, the matrix loading flow control unit notifies the near memory chip access controller so that the near memory chip access controller controls the matrix write cache to send the weight matrix to the near memory computing chip.

7. The large model acceleration system according to claim 4, characterized in that: In the decoding stage, the large model acceleration chip and the near-memory computing chip perform calculations of the nonlinear computing unit and matrix-vector multiplication calculations alternately in the first thread and the second thread respectively.

8. A large model acceleration chip is configured to preprocess an input vector in a decoding stage of a large model, distribute the preprocessed vector to a near-memory computing chip, and send a calculation instruction to the near-memory computing chip so that the near-memory computing chip performs a matrix-vector multiplication calculation task, the large model acceleration chip comprising: a nonlinear computing unit configured to perform the pre-processing before the matrix-vector multiplication calculation and the post-processing after the matrix-vector multiplication calculation on the vector in the decoding stage; An operation controller is configured to control the scheduling of computing tasks and computing data using computing instructions, wherein the computing tasks include computing in the nonlinear computing unit and computing in the near-memory computing chip, and the computing data includes the vector and a weight matrix of the large model; A near-memory computing configuration memory configured to store configuration information included in a computing instruction from an operation controller, the configuration information including configuration information indicating a predetermined accumulator configuration and configuration information indicating how to distribute the computing data to near-memory computing chips; A near memory chip access controller is configured to read the configuration information from the near memory computing configuration memory and control the sending of the computing data to the near memory computing chip according to the configuration information; A matrix write cache is configured to send the weight matrix to the near memory computing chip under the control of the near memory chip access controller before the calculation of the decoding stage starts, so that the weight matrix is ​​written into the memory of the near memory computing chip; as well as The vector cache is configured to cache the vector before the vector is sent to the near memory computing chip, and send the vector to the near memory computing chip under the control of the near memory chip access controller.

9. A near-memory computing chip is configured to perform a matrix-vector multiplication calculation task based on the distributed vectors and the stored weight matrix of the large model in accordance with the calculation instructions from the large model acceleration chip during the decoding stage of the large model, and send the calculation results to the large model acceleration chip, the near-memory computing chip comprising: A memory configured to store a weight matrix of a large model; as well as The calculation part includes: a plurality of variable accumulation channels, each variable accumulation channel comprising a plurality of fixed accumulation groups and a first variable accumulator, and each variable accumulation channel corresponds to a memory read channel of the weight matrix, wherein the first variable accumulator selectively accumulates a plurality of primary accumulation results generated by performing matrix-vector multiplication calculations on the plurality of fixed accumulation groups according to a predetermined accumulator configuration to obtain a first accumulation result; as well as The second variable accumulator selectively accumulates the multiple first accumulation results generated by the multiple variable accumulation channels according to a predetermined accumulator configuration to obtain a second accumulation result, and sends the second accumulation result as a calculation result of the near-memory computing chip to the large model acceleration chip.

10. The near memory computing chip according to claim 9, characterized in that: Each of the plurality of fixed accumulation groups comprises: a plurality of multipliers, each multiplier being configured to perform a multiplication calculation on an element of the vector and an element of the weight matrix; and The fixed accumulator accumulates the multiplication calculation results of the plurality of multipliers to obtain the primary accumulation result.

Citation Information

Cited By

  • Large language model accelerator architecture based on three-dimensional NAND flash memory

    CN121501740A

  • Large language model accelerator system based on three-dimensional NAND flash memory

    CN121501740B