Deep learning processor architecture system containing attention mechanism and application method

Through the integrated design of attention mechanism control section and computing unit control section, the problems of excessive load and inefficiency in the traditional deep learning processor architecture are solved, and efficient data processing and simplified system development are achieved.

CN120492178AActive Publication Date: 2025-08-15SHANDONG INSPUR SCI RES INST CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510990073.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-08-15
Estimated Expiration
2045-07-18

AI Technical Summary

Technical Problem

In the traditional deep learning processor architecture, the host processor performs data splitting and attention score calculations, resulting in excessive load, separation of attention mechanisms from processor calculation units leads to inefficiency, and complex architecture design is difficult to widely use.

Method used

The attention mechanism control section and the calculation unit control section are adopted with an integrated design, including the attention mechanism parameter real-time calculation unit and calculation unit. It is connected through the AXI bus to realize hardware autonomy of attention score calculation, and unified management through the management information configuration module, reducing the load of the host processor, and building a closed-loop pipeline for computing-scheduling-execution.

Benefits of technology

It reduces the computing load of the host processor, improves the timeliness of data processing, reduces the complexity of system development, and allows non-hardware professional researchers to quickly deploy customized models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492178A_ABST
    Figure CN120492178A_ABST
Patent Text Reader

Abstract

The invention discloses a deep learning processor architecture system containing an attention mechanism and an application method, mainly relates to the technical field of processor architectures, and is used for solving the problems that a traditional architecture is executed by a host processor, so that the load is too high, the attention mechanism is separated from a processor computing unit, and the architecture design is complex. Comprising an input sequence module, a decomposition mode reading module connected with the input sequence module, an attention parameter generation module and an input sequence cache module which are connected with the decomposition mode reading module, and an attention score calculation control FSM state machine module connected with the attention parameter generation module; the attention score calculation control FSM state machine module is connected with the input sequence cache module, the attention score calculation module is connected with the attention score calculation control FSM state machine module, the MUX data merging module is connected with the input sequence cache module and the attention score calculation control FSM state machine module, and the optimal scheduling method selection setting module is connected with the MUX data merging module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of deep learning processor architecture, and in particular to a deep learning processor architecture system including an attention mechanism and an application method. Background Art

[0002] A processor, collectively referred to as a CPU (Central Processing Unit) or a GPU (Graphic Processing Unit), is a core component of modern computer systems, primarily responsible for parsing computer instructions and rapidly processing complex data. With the rapid advancement of technology, the demand for higher-performance processors for deep learning is growing. Researchers have proposed incorporating attention mechanisms into deep learning.

[0003] However, this design also has some problems in practical application, which are specifically reflected in: 1. The traditional architecture requires the host processor to perform the splitting of the data to be processed and the attention score calculation, and perform specific scheduling on the target computing unit for task allocation. These processes bring extremely high load to the processor, and excessive load will cause excessive delays in the calculation of other important processes of the processor; 2. Existing solutions usually adopt an organizational structure that separates the attention mechanism from the processor computing unit, that is, the user's instructions and the data to be processed need to first pass through the host processor to perform the attention score calculation, and the results are then transmitted to the processor computing unit section via the on-chip bus after scheduling to perform target data processing. This design is bound to reduce the data processing efficiency of important processes; 3. The existing architecture design is complex and extremely difficult to apply. Users need to master complex programming and development basic systems, which is not conducive to the application and development of deep learning systems by a wide range of researchers. Summary of the Invention

[0004] The present application provides a deep learning processor architecture system and application method including an attention mechanism to solve the problems of excessive load caused by execution of the traditional architecture by the host processor, separation of the attention mechanism from the processor computing unit, and complex architecture design.

[0005] In a first aspect, the present application provides a deep learning processor architecture system with an attention mechanism, the system comprising: The deep learning processor includes an attention mechanism control module, a computing unit control module, a processor configuration module, and an AXI (Advanced eXtensible Interface) control / data bus connecting these modules. The attention mechanism control module and the computing unit control module adopt an integrated design. Among them, the attention mechanism control module includes: attention mechanism parameter real-time calculation unit, implicit attention mechanism scheduling unit, and explicit attention mechanism scheduling unit; The real-time calculation unit of attention mechanism parameters includes: an input sequence module, a decomposition method reading module connected to the input sequence module, an attention parameter generation module and an input sequence cache module connected to the decomposition method reading module, an attention score calculation control FSM (Finite State Machine) state machine module connected to the attention parameter generation module, an attention score calculation module connected to the attention score calculation control FSM state machine module, a MUX (Multiplexer) data merging module connected to the input sequence cache module and the attention score calculation control FSM state machine module, an optimal scheduling method selection setting module connected to the MUX data merging module, an attention score output module connected to the optimal scheduling method selection setting module, and a management information configuration module connected to all modules; the management information configuration module is connected to the processor configuration panel, and the attention score output module is connected to the implicit attention mechanism scheduling unit and the explicit attention mechanism scheduling unit.

[0006] In one implementation of the present application, the system further includes: Host, the Host is connected to the deep learning processor and is used to transmit information data to the deep learning processor; wherein the information data includes preset necessary parameters, calculation instructions and data to be processed in the deep learning processor.

[0007] In one implementation of the present application, the deep learning processor further includes: DMA (Direct Memory Access) block, processor interrupt control block, and data storage block.

[0008] In a second aspect, the present application provides an application method of a deep learning processor architecture system including an attention mechanism. Based on the deep learning processor architecture system including an attention mechanism, the method includes: Obtain the preset necessary parameters, calculation instructions and data to be processed transmitted by the host through the deep learning processor; Based on the preset necessary parameters, the processor configuration module completes the configuration of the parameters of each module in the deep learning processor and the parameters of each unit in the attention mechanism control module. In addition, the management information configuration module completes the configuration of the parameters of each module in the attention mechanism parameter real-time calculation unit. After completing the configuration of each module parameter in the attention mechanism parameter real-time calculation unit, Receive calculation instructions and data to be processed through the input sequence module, and transmit the calculation instructions and data to be processed to the decomposition mode reading module; The decomposition method reading module determines the data decomposition format corresponding to the calculation instruction according to the preset input sequence decomposition method comparison table, and then obtains the decomposition data corresponding to the data to be processed; sends the decomposition data to the attention parameter generation module, and sends the calculation instruction, decomposition data and data to be processed to the input sequence buffer module; The attention parameter generation module performs a preset linear transformation on the decomposed data to generate the corresponding Q, K, and V vectors; the decomposed data and the corresponding Q, K, and V vectors are sent to the attention score calculation control FSM state machine; The attention score calculation module is controlled by the FSM state machine according to the preset steps to calculate the attention score of the Q, K, and V vectors corresponding to the decomposed data; the attention score is sent to the MUX data merging module; The MUX data merging module obtains the data to be processed and the calculation instructions transmitted by the input sequence buffer module, aggregates and splices the data to be processed and the attention score, and sends the calculation instructions and the aggregated splicing results to the optimal scheduling method selection and setting module; The optimal scheduling method selection setting module verifies whether the attention score meets the preset verification rules. If it meets the requirements, the scheduling mechanism to be adopted is determined according to the calculation instructions. The data packet containing the scheduling mechanism, the data to be processed, the attention score, and the calculation instructions is sent to the attention score output module. According to the scheduling mechanism, the attention score output module sends the data packet to the corresponding implicit attention mechanism scheduling unit or explicit attention mechanism scheduling unit; According to the preset allocation scheduling strategy in the implicit attention mechanism scheduling unit or the explicit attention mechanism scheduling unit, the data to be processed and the corresponding attention scores are sent to the computing unit control board in sequence; Use the computing unit and deep learning algorithm in the computing unit control module to obtain the calculation results and write the calculation results back to the host.

[0009] In one implementation of the present application, the decomposition method reading module determines the data decomposition format corresponding to the calculation instruction according to the preset input sequence decomposition method comparison table, and then obtains the decomposition data corresponding to the data to be processed, specifically including: Perform hash calculation on the calculation instruction to obtain the data decomposition format storage address of the current calculation instruction stored in the preset input sequence decomposition method comparison table, and then read the data decomposition format corresponding to the calculation instruction; adjust the format of the received data to be processed according to the data decomposition format to obtain the decomposed data.

[0010] In one implementation of the present application, the attention score calculation module is controlled by the attention score calculation control FSM state machine according to preset steps to calculate the attention scores of the Q, K, and V vectors corresponding to the decomposed data, specifically including: According to the order of input decomposition data, the preset linear transformation is performed on each decomposition data with the user target attention weight in the preset necessary parameters to generate the corresponding Q, K, and V vectors.

[0011] In one implementation of the present application, the attention score calculation module is controlled by the attention score calculation control FSM state machine according to preset steps to calculate the attention scores of the Q, K, and V vectors corresponding to the decomposed data, specifically for: Perform attention score calculation on the Q, K, and V vectors corresponding to each input decomposition data in 4 steps: S0, receives the Q, K, V vectors corresponding to the decomposed data; S1, transmit the Q, K, V vectors corresponding to the decomposed data to the attention score calculation module, control the target operator of the attention score calculation module to perform the dot product calculation of Q and K, and obtain the result R1; S2, controls the target operator of the attention score calculation module to perform the Softmax function normalization calculation and obtains the attention weight R2; S3, controls the target operator of the attention score calculation module to perform weighted sum calculation of attention weights, and performs weighted summation of attention weights and corresponding V vectors to obtain weighted sum vector R3; The data to be processed, R1, R2, and R3 are aggregated and spliced one by one as the attention score.

[0012] In one implementation of the present application, the Q, K, and V vectors corresponding to the decomposed data are transmitted to the attention score calculation module, and the target operator of the attention score calculation module is controlled to perform the dot product calculation of Q and K to obtain the result R1, which specifically includes: The similarity is calculated by the formula R1= ; Where i represents the number of columns of the Q vector, and j represents the number of rows of the K vector; The target operator of the control attention score calculation module performs the Softmax function normalization calculation to obtain the attention weight R2, which specifically includes: By formula: , normalized; among them, Indicates the dimensions of Q and K; By formula: , calculate the attention weight R2; The target operator of the control attention score calculation module performs the weighted sum calculation of the attention weight. The attention weight is weighted and summed with the corresponding V vector to obtain the weighted sum vector R3, which specifically includes: By formula: , and obtain the weighted sum vector R3.

[0013] In one implementation of the present application, according to the preset allocation scheduling strategy in the implicit attention mechanism scheduling unit or the explicit attention mechanism scheduling unit, the data to be processed and the corresponding attention scores are sent to the computing unit control module in sequence, specifically including: The implicit attention mechanism scheduling unit or the explicit attention mechanism scheduling unit sends the data to be processed and the corresponding attention scores to the target computing unit of the computing unit control module in sequence according to the preset allocation scheduling strategy, and the target computing unit performs the corresponding deep learning data processing and calculation.

[0014] In one implementation of the present application, the computing unit and deep learning algorithm in the computing unit control module are used to obtain the calculation results and write the output results back to the host, specifically including: After the target computing unit in the computing unit control module completes the computing task, the computing result is temporarily stored in the consistent cache of the computing unit control module and written back to the host terminal in the form of DMA through the PCIe interface (Peripheral Component Interconnect Express); The CPU of the host terminal merges and processes the calculation instructions and the calculation results of each target calculation unit corresponding to the attention score allocation, verifies the calculation results according to the preset verification rules, and confirms that the calculation is completed when the preset verification rules are met.

[0015] It can be seen from the above technical solutions that this application has the following advantages: 1. Reduce the host processor computing load The architecture involved in this application separates the attention score calculation function from the host processor to a dedicated hardware unit (the real-time calculation unit for attention mechanism parameters), so that the host processor only needs to handle initial configuration management (through the management information configuration module). The attention score calculation controls the collaborative work of the FSM state machine module and the calculation module, achieving hardware autonomy for the entire process from input sequence decomposition and parameter generation to score calculation. This design eliminates the burden of the host processor frequently intervening in data splitting, score calculation, and task scheduling in traditional architectures, fundamentally avoiding the problem of critical process delays caused by overload. The topology of the AXI bus directly connecting various functional modules further reduces the data handling pressure on the host processor.

[0016] 2. Improved data processing timeliness: The integrated design of the attention mechanism control module and the computation unit control module creates a closed-loop pipeline of computation, scheduling, and execution. The linkage between the MUX data merging module and the optimal scheduling method selection and setting module allows the output of attention scores to be directly allocated to computing resources through implicit / explicit scheduling units, eliminating the cross-processor data transmission required in traditional solutions. The introduction of the input sequence cache module enables preloading of computing resources. Combined with the parallel processing capabilities of the attention score calculation module, the entire data process from feature extraction to result output is completed on-chip.

[0017] 3. Reduced system development complexity: The unified management provided by the processor configuration module (through the management information configuration module) abstracts previously fragmented operations such as attention mechanism configuration and computational unit parameter setting into a standardized instruction set. Developers simply read the module's policy parameters in a decomposed manner through a high-level API, automatically triggering the subsequent attention parameter generation, score calculation, and scheduling allocation processes. The programmable interface for explicit and implicit scheduling units supports flexible switching between different algorithmic paradigms. This "configure and use" feature lowers the technical threshold for deep learning system development, enabling researchers without hardware expertise to quickly deploy customized models. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solution of the present invention, the following is a brief introduction to the drawings required for the description. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0019] Figure 1 This is a schematic diagram of the internal structure of a deep learning processor architecture system with an attention mechanism provided in an embodiment of the present application.

[0020] Figure 2 This is a schematic diagram of the internal structure of a real-time calculation unit for attention mechanism parameters provided in an embodiment of the present application.

[0021] Figure 3 This is a schematic diagram of the internal structure of a deep learning processor architecture system provided in an embodiment of the present application.

[0022] Figure 4 This is a detailed internal structure diagram of an efficient deep learning processor with an attention mechanism provided in an embodiment of the present application.

[0023] Figure 5This is a flow chart of a method for applying a deep learning processor architecture system with an attention mechanism provided in an embodiment of the present application.

[0024] Description of main reference numerals: 100. Deep learning processor; 110. Attention mechanism control module; 111. Attention mechanism parameter real-time calculation unit; 1. Input sequence module; 2. Decomposition method reading module; 3. Attention parameter generation module; 4. Input sequence cache module; 5. Attention score calculation control FSM state machine module; 6. Attention score calculation module; 7. MUX data merging module; 8. Optimal scheduling method selection and setting module; 9. Attention score output module; 10. Management information configuration module; 112. Implicit attention mechanism scheduling unit; 113. Explicit attention mechanism scheduling unit; 120, computing unit control block; 130, processor configuration block; 140, AXI control / data bus; 150, on-chip cache block; 160, DMA block; 170, processor interrupt control block; 180, data storage block; 200, host. DETAILED DESCRIPTION

[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0026] It should be understood by those skilled in the art that the embodiments described below are merely preferred embodiments of the present disclosure and do not imply that the present disclosure can only be implemented through these preferred embodiments. These preferred embodiments are merely intended to explain the technical principles of the present disclosure and are not intended to limit the scope of protection of the present disclosure. Based on the preferred embodiments provided by the present disclosure, all other embodiments obtained by those skilled in the art without creative effort should still fall within the scope of protection of the present disclosure.

[0027] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0028] The technical solutions proposed in the embodiments of the present application are described in detail below with reference to the accompanying drawings.

[0029] This application Figure 1-2 As shown in FIG, a deep learning processor architecture system with an attention mechanism is provided in an embodiment of the present application. Figure 1-2 As shown, the system provided in the embodiment of the present application mainly includes: Deep learning processor 100, such as Figure 1 As shown, the deep learning processor 100 includes: an attention mechanism control block 110, a computing unit control block 120, a processor configuration block 130, and an AXI control / data bus 140 connecting the various blocks; and the attention mechanism control block 110 and the computing unit control block 120 adopt an integrated design; The attention mechanism control module 110 includes: an attention mechanism parameter real-time calculation unit 111, an implicit attention mechanism scheduling unit 112, and an explicit attention mechanism scheduling unit 113; like Figure 2 As shown, the attention mechanism parameter real-time calculation unit 111 includes: an input sequence module 1, a decomposition method reading module 2 connected to the input sequence module 1, an attention parameter generation module 3 and an input sequence cache module 4 connected to the decomposition method reading module 2, an attention score calculation control FSM state machine module 5 connected to the attention parameter generation module 3, an attention score calculation module 6 connected to the attention score calculation control FSM state machine module 5, a MUX data merging module 7 connected to the input sequence cache module 4 and the attention score calculation control FSM state machine module 5, an optimal scheduling method selection setting module 8 connected to the MUX data merging module 7, an attention score output module 9 connected to the optimal scheduling method selection setting module 8, and a management information configuration module 10 connected to all modules; and the management information configuration module 10 is connected to the processor configuration panel 130, and the attention score output module 9 is connected to the implicit attention mechanism scheduling unit 112 and the explicit attention mechanism scheduling unit 113.

[0030] To further explain, Figure 2 The function of the attention mechanism control module 110 involved is to receive all deep learning instructions to be executed and related data to be processed, and calculate the attention scores corresponding to all data to be processed in sequence according to the input processing instruction category and related user configuration data; and package the input deep learning processing data and the corresponding attention scores, and send the data packets in sequence to the implicit attention mechanism scheduling unit 112 / explicit attention mechanism scheduling unit 113 according to the type that best adapts the attention score calculation result.

[0031] from Figure 2It can be seen that the architecture of the attention mechanism parameter real-time calculation unit 111 includes a management information configuration module 10, an input sequence module 1, a decomposition method reading module 2, an input sequence cache module 4, an attention parameter generation module 3, an attention score calculation control FSM state machine module 5, an attention score calculation module 6 (including a series of underlying computing function units such as ALU, FPU, SFU, etc.), an attention score output module 9, a MUX data merging module 7, and an optimal scheduling method selection setting module 8.

[0032] Among them, the management information configuration module 10 is responsible for completing the necessary parameter settings and initialization configurations for other modules, and configuring the information preset by the Host user; the input sequence module 1 is responsible for receiving the calculation instructions and related data to be processed from the on-chip cache block 150 in the order of the user's deep learning request; the decomposition method reading module 2 is used to read the data decomposition format corresponding to the instruction according to the data decomposition format storage address stored in the input sequence decomposition method comparison table of the current instruction, and adjust the format of the received raw data to be processed; the input sequence cache module 4 is used to temporarily store instructions and the decomposition results of the input data after the format adjustment; the attention parameter generation module 3 is responsible for performing a specific linear transformation on each input data based on the user configuration information in the order of the input decomposition data, and then generating the Q, K, and V vectors corresponding to each group of data; the attention The score calculation control FSM state machine module 5 is responsible for controlling the attention score calculation module 6 to perform attention score calculation on the Q, K, and V vectors corresponding to each input original decomposition data according to preset steps, and integrate the calculation results; the MUX data merging module 7 is responsible for reading the split input to be processed original data from the input sequence cache module 4, and completing data aggregation and splicing in one-to-one correspondence with the calculation results R1, R2, and R3 of the attention score calculation control FSM state machine module 5; the optimal scheduling method selection setting module 8 is responsible for secondary verification of the original decomposition data and the corresponding attention score calculation results (preset verification rules are provided, and technical personnel in this field can determine the specific rule content and scope according to actual needs), and then judge according to the input instructions whether the current task is allocated to the subsequent computing unit and is more suitable for implicit attention mechanism scheduling / explicit attention mechanism scheduling.

[0033] The deep learning processor 100 architecture system can be further specifically, as Figure 3 As shown, this application can be specifically: Figure 3 An efficient deep learning processor 100 architecture system with an attention mechanism is demonstrated. The server user node is connected to the Host 200 through the network, and the heterogeneous deep learning processor 100 communicates and exchanges information with the Host 200 in the form of DMA through the PCIe interface.

[0034] The deep learning processor 100 is mainly composed of a DMA data transfer module ( Figure 3 It is composed of the attention mechanism control module 110, the computing unit control module 120, the on-chip cache module 150 and other parts (not shown).

[0035] After the Host 200 receives the user's deep learning processing request, it sends the calculation instructions and the image / text data to be processed to the deep learning processor 100 through the PCIe interface. The attention mechanism control module 110 of the deep learning processor 100 completes the attention mechanism parameter calculation for the relevant data of the current calculation request and dispatches the specific calculation task to the target calculation unit of the calculation unit control module 120; after the calculation is completed, the calculation result is sent to the Host 200 in the form of DMA through the PCIe interface, and then fed back to the requesting user terminal.

[0036] In addition, the deep learning processor 100 further includes: DMA block 160 , processor interrupt control block 170 , and data storage block 180 .

[0037] The deep learning processor can be further specifically, Figure 4 As shown, the deep learning processor can be specifically: Figure 4 The detailed architecture of the efficient deep learning processor 100 with attention mechanism is shown, including the attention mechanism control block 110 (including the real-time calculation unit 111 of attention mechanism parameters, the implicit attention mechanism scheduling unit 112, and the explicit attention mechanism scheduling unit 113), the DMA block 160, the computing unit control block 120 (including all computing units deployed by the efficient deep learning processor 100), the processor interrupt control block 170, the processor configuration block 130, the data storage block 180, and the AXI control / data bus 140 connecting each block.

[0038] The processor interrupt control module 170 is responsible for processing the interrupt processing request from the host 200; The processor configuration module 130 is responsible for completing the host 200's configuration of parameters related to the attention mechanism control module 110, the computational unit control module 120, and the data storage module 180, such as the instruction TCM (tightly coupled memory), data TCM (data TCM refers to tightly coupled memory data), etc. The DMA module 160 deploys a DMA data transceiver engine, which is responsible for data exchange between the processor and the host memory, sending and receiving computing task-related information and transmitting computing results in real time; The attention mechanism control module 110 is equipped with a real-time attention mechanism parameter calculation unit 111, an input sequence decomposition method comparison table (which stores all user instructions and the decomposition format of the data to be processed corresponding to the instructions), an implicit attention mechanism scheduling unit 112, and an explicit attention mechanism scheduling unit 113. This module is responsible for completing the calculation of the attention mechanism parameters for the relevant data to be processed related to the current user's deep learning computing request and scheduling the specific computing tasks containing the attention mechanism parameters to the target computing unit of the computing unit control module 120. The attention mechanism control module 110 and the computing unit control module 120 adopt an integrated design. After completing the calculation of the attention mechanism parameters for the data to be processed and the attention mechanism parameters, the task scheduling can be directly executed on the original data without the tedious and time-consuming system bus transmission. The computing unit control module 120 deploys all deep learning computing units of the processor and is responsible for executing deep learning computing tasks on the data to be processed, including attention mechanism parameters. The deep learning computing units include scalar data computing functions, vector data computing functions, and tensor data computing functions required for deep learning data processing. The computing unit control module 120 is responsible for unified coordination and control of all computing units. The data storage module 180 includes instruction TCM storage and data TCM storage, which are responsible for storing relevant instruction data and calculation original input data; Different modules are connected through the AXI control bus and AXI data bus. The AXI control / data bus 140 carries the necessary instructions and related data for operation. All functional modules are uniformly controlled and deployed by the efficient deep learning processor 100 with an attention mechanism.

[0039] In addition, the embodiment provides a deep learning processor architecture system with an attention mechanism, such as Figure 5 As shown, the method provided in the embodiment of the present application mainly includes the following steps: Step 210: The deep learning processor obtains the preset necessary parameters, computation instructions, and intended processing data transmitted by the host. Based on the preset necessary parameters, the processor configuration module configures the parameters of each module in the deep learning processor and each unit in the attention mechanism control module. Furthermore, the management information configuration module configures the parameters of each module in the attention mechanism parameter real-time calculation unit.

[0040] It can be understood by those skilled in the art that the configuration of the parameters of each module in the real-time calculation unit of the attention mechanism parameters can be specifically as follows: the management information configuration module completes the corresponding necessary parameter settings (the necessary parameters are preset necessary parameters) for the input sequence module, the decomposition method reading module, the input sequence cache module, the attention parameter generation module, the attention score calculation control FSM state machine module, the attention score calculation module, the optimal scheduling method selection setting module, the attention score output and other modules.

[0041] Step 220: After completing the configuration of the parameters of each module in the real-time calculation unit of the attention mechanism parameters, the calculation instructions and the data to be processed are received through the input sequence module, and the calculation instructions and the data to be processed are transmitted to the decomposition reading module.

[0042] It should be noted that the above-mentioned data transmission method is: the calculation instructions and the data to be processed are sent to the decomposition reading module respectively according to the input order.

[0043] Step 230: The decomposition method reading module determines the data decomposition format corresponding to the calculation instruction according to the preset input sequence decomposition method comparison table, and then obtains the decomposition data corresponding to the data to be processed; sends the decomposition data to the attention parameter generation module, and sends the calculation instruction, decomposition data and data to be processed to the input sequence cache module.

[0044] The decomposition method reading module determines the data decomposition format corresponding to the calculation instruction according to the preset input sequence decomposition method comparison table, and then obtains the decomposition data corresponding to the data to be processed, which can be specifically: Perform hash calculation on the calculation instruction to obtain the data decomposition format storage address stored in the preset input sequence decomposition method comparison table of the current calculation instruction, and then read the data decomposition format corresponding to the calculation instruction; then the decomposition method reading module adjusts the format of the received data to be processed according to the data decomposition format (including the decomposition data bit width, decomposition format scalar, vector, matrix attributes, decomposition header and tail signals, etc.).

[0045] Step 240: Perform a preset linear transformation on the decomposed data through the attention parameter generation module to generate corresponding Q, K, and V vectors; send the decomposed data and the corresponding Q, K, and V vectors to the attention score calculation control FSM state machine.

[0046] Among them, the attention score calculation module is controlled by the FSM state machine according to the preset steps to calculate the attention score of the Q, K, and V vectors corresponding to the decomposed data. Specifically, it can be: According to the order of input decomposition data, the preset linear transformation is performed on each decomposition data with the user target attention weight in the preset necessary parameters to generate the corresponding Q, K, and V vectors.

[0047] It should be noted that the Q, K, and V vectors are the three elements for calculating attention parameters - Query, Key, and Value vectors.

[0048] Step 250: Control the attention score calculation module by the FSM state machine according to preset steps through attention score calculation to calculate the attention scores of the Q, K, and V vectors corresponding to the decomposed data; and send the attention scores to the MUX data merging module.

[0049] The attention score calculation module is controlled by the FSM state machine according to preset steps to calculate the attention score of the Q, K, and V vectors corresponding to the decomposed data. It is specifically used for: Perform attention score calculation on the Q, K, and V vectors corresponding to each input decomposition data in 4 steps: S0, receives the Q, K, V vectors corresponding to the decomposed data; S1: Transmit the Q, K, and V vectors corresponding to the decomposed data to the attention score calculation module, and control the target operator (ALU, FPU...the specific operator type is related to the data type) of the attention score calculation module to perform the dot product calculation of Q and K to obtain the result R1; S2, controls the target operator of the attention score calculation module to perform the Softmax function normalization calculation and obtains the attention weight R2; S3, controls the target operator of the attention score calculation module to perform weighted sum calculation of attention weights, and performs weighted summation of attention weights and corresponding V vectors to obtain weighted sum vector R3; The data to be processed, R1, R2, and R3 are aggregated and spliced one by one as the attention score.

[0050] More specifically, the Q, K, and V vectors corresponding to the decomposed data are transmitted to the attention score calculation module, and the target operator of the attention score calculation module is controlled to perform the dot product calculation of Q and K to obtain the result R1, which can be specifically: The similarity is calculated by the formula R1= ; Where i represents the number of columns of the Q vector, and j represents the number of rows of the K vector; The target operator of the control attention score calculation module performs the Softmax function normalization calculation, and the obtained attention weight R2 can be specifically: By formula: , normalized; among them, Indicates the dimensions of Q and K; By formula: , calculate the attention weight R2; The target operator of the control attention score calculation module performs the weighted sum calculation of the attention weight. The attention weight is weighted and summed with the corresponding V vector to obtain the weighted sum vector R3, which can be specifically: By formula: , and obtain the weighted sum vector R3.

[0051] Step 260: Obtain the data to be processed and the calculation instructions transmitted by the input sequence cache module through the MUX data merging module, aggregate and splice the data to be processed and the attention score, and send the calculation instructions and the aggregated splicing results to the optimal scheduling method selection and setting module.

[0052] Step 270: Verify whether the attention score meets the preset verification rules through the optimal scheduling method selection setting module. If it meets the requirements, determine the scheduling mechanism to be adopted according to the calculation instructions; send the data packet containing the scheduling mechanism, the data to be processed, the attention score, and the calculation instructions to the attention score output module.

[0053] Step 280: According to the scheduling mechanism, the attention score output module sends the data packet to the corresponding implicit attention mechanism scheduling unit or the explicit attention mechanism scheduling unit; according to the preset allocation scheduling strategy in the implicit attention mechanism scheduling unit or the explicit attention mechanism scheduling unit, the data to be processed and the corresponding attention score are sent to the computing unit control module in sequence.

[0054] According to the preset allocation scheduling strategy in the implicit attention mechanism scheduling unit or the explicit attention mechanism scheduling unit, the data to be processed and the corresponding attention scores are sent to the computing unit control section in sequence, including: The implicit attention mechanism scheduling unit or the explicit attention mechanism scheduling unit sends the data to be processed and the corresponding attention scores to the target computing unit of the computing unit control module in sequence according to the preset allocation scheduling strategy, and the target computing unit performs the corresponding deep learning data processing and calculation.

[0055] Step 290: Use the computing unit and deep learning algorithm in the computing unit control module to obtain the calculation results, and write the calculation results back to the Host.

[0056] Utilize the computing units and deep learning algorithms in the computing unit control module to obtain calculation results and write the output results back to the host. Specifically, this includes: After the target computing unit in the computing unit control module completes the computing task, the computing result is temporarily stored in the consistent cache of the computing unit control module and written back to the host terminal in the form of DMA through the PCIe interface; The CPU of the host terminal merges and processes the calculation instructions and the calculation results of each target calculation unit corresponding to the attention score allocation, verifies the calculation results according to the preset verification rules, and confirms that the calculation is completed when the preset verification rules are met.

[0057] Based on the above description, as an example, this embodiment may be: 1. The user sends a deep learning computing or application-related request to the host.

[0058] 2. The host CPU receives the currently requested instructions and the image / text data to be processed, and then sends the instructions and related data to the deep learning processor with the attention mechanism in the form of DMA through the PCIe interface.

[0059] 3. The attention mechanism parameter calculation unit included in the attention mechanism control module of the deep learning processor sequentially calculates the attention scores corresponding to all the data to be processed based on the input processing instructions and relevant user configuration data.

[0060] 4. After calculating the attention score, the attention mechanism parameter calculation unit packages the input pseudo-deep learning processing data and the corresponding attention score. Based on the best-suited type of attention score calculation result (implicit / explicit), the data packets are sent to the implicit attention mechanism scheduling unit or the explicit attention mechanism scheduling unit in sequence, which then generates the corresponding allocation scheduling strategy.

[0061] 5. The implicit attention mechanism scheduling unit / explicit attention mechanism scheduling unit sends the input data to be processed and the corresponding attention scores to the target computing unit of the computing unit control module in sequence according to the generated allocation scheduling strategy, and the target computing unit performs the corresponding deep learning data processing and calculation.

[0062] 6. After the target computing unit completes the computing task, the computing result is temporarily stored in the consistency cache of the top-level control module of the computing unit, and the computing result is written back to the host terminal in the form of DMA through the PCIe interface.

[0063] 7. The CPU merges the current instruction data and the calculation results of each execution calculation unit corresponding to the attention score allocation, and checks the deep learning calculation results twice to confirm that the calculation is completed.

[0064] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A deep learning processor architecture system with an attention mechanism, characterized in that: The system comprises: The deep learning processor includes: an attention mechanism control module, a computing unit control module, a processor configuration module, and an AXI control / data bus connecting the various modules; the attention mechanism control module and the computing unit control module adopt an integrated design; Among them, the attention mechanism control module includes: attention mechanism parameter real-time calculation unit, implicit attention mechanism scheduling unit, and explicit attention mechanism scheduling unit; The real-time calculation unit of attention mechanism parameters includes: an input sequence module, a decomposition method reading module connected to the input sequence module, an attention parameter generation module and an input sequence cache module connected to the decomposition method reading module, an attention score calculation control FSM state machine module connected to the attention parameter generation module, an attention score calculation module connected to the attention score calculation control FSM state machine module, a MUX data merging module connected to the input sequence cache module and the attention score calculation control FSM state machine module, an optimal scheduling method selection setting module connected to the MUX data merging module, an attention score output module connected to the optimal scheduling method selection setting module, and a management information configuration module connected to all modules; and the management information configuration module is connected to the processor configuration panel, and the attention score output module is connected to the implicit attention mechanism scheduling unit and the explicit attention mechanism scheduling unit.

2. The deep learning processor architecture system with attention mechanism according to claim 1, characterized in that The system further comprises: The host is connected to the deep learning processor and is used to transmit information data to the on-chip cache of the deep learning processor; the information data includes the preset necessary parameters, calculation instructions and data to be processed in the deep learning processor.

3. The deep learning processor architecture system with attention mechanism according to claim 1, characterized in that The deep learning processor further includes: DMA section, processor interrupt control section, and data storage section.

4. A method for applying a deep learning processor architecture system with an attention mechanism, based on the deep learning processor architecture system with an attention mechanism according to claim 1, characterized in that: The method comprises: Obtain the preset necessary parameters, calculation instructions and data to be processed transmitted by the host through the deep learning processor; Based on the preset necessary parameters, the processor configuration module completes the configuration of the parameters of each module in the deep learning processor and the parameters of each unit in the attention mechanism control module. In addition, the management information configuration module completes the configuration of the parameters of each module in the attention mechanism parameter real-time calculation unit. After completing the configuration of the parameters of each module in the real-time calculation unit of the attention mechanism parameters, the calculation instructions and the data to be processed are received through the input sequence module, and the calculation instructions and the data to be processed are transmitted to the decomposition mode reading module; The decomposition method reading module determines the data decomposition format corresponding to the calculation instruction according to the preset input sequence decomposition method comparison table, and then obtains the decomposition data corresponding to the data to be processed; sends the decomposition data to the attention parameter generation module, and sends the calculation instruction, decomposition data and data to be processed to the input sequence buffer module; The attention parameter generation module performs a preset linear transformation on the decomposed data to generate the corresponding Q, K, and V vectors; the decomposed data and the corresponding Q, K, and V vectors are sent to the attention score calculation control FSM state machine; The attention score calculation module is controlled by the FSM state machine according to the preset steps to calculate the attention score of the Q, K, and V vectors corresponding to the decomposed data; the attention score is sent to the MUX data merging module; The MUX data merging module obtains the data to be processed and the calculation instructions transmitted by the input sequence buffer module, aggregates and splices the data to be processed and the attention score, and sends the calculation instructions and the aggregated splicing results to the optimal scheduling method selection and setting module; The optimal scheduling method selection setting module verifies whether the attention score meets the preset verification rules. If it meets the requirements, the scheduling mechanism to be adopted is determined according to the calculation instructions. The data packet containing the scheduling mechanism, the data to be processed, the attention score, and the calculation instructions is sent to the attention score output module. According to the scheduling mechanism, the attention score output module sends the data packet to the corresponding implicit attention mechanism scheduling unit or explicit attention mechanism scheduling unit; according to the preset allocation scheduling strategy in the implicit attention mechanism scheduling unit or the explicit attention mechanism scheduling unit, the data to be processed and the corresponding attention score are sent to the computing unit control board in sequence; Use the computing unit and deep learning algorithm in the computing unit control module to obtain the calculation results and write the calculation results back to the host.

5. The method for applying a deep learning processor architecture system with an attention mechanism according to claim 4, wherein: The decomposition method reading module determines the data decomposition format corresponding to the calculation instruction according to the preset input sequence decomposition method comparison table, and then obtains the decomposition data corresponding to the data to be processed, specifically including: Perform hash calculation on the calculation instruction to obtain the data decomposition format storage address of the current calculation instruction stored in the preset input sequence decomposition method comparison table, and then read the data decomposition format corresponding to the calculation instruction; adjust the format of the received data to be processed according to the data decomposition format to obtain the decomposed data.

6. The method for applying a deep learning processor architecture system with an attention mechanism according to claim 4, wherein: The attention score calculation module is controlled by the FSM state machine according to the preset steps to calculate the attention score of the Q, K, and V vectors corresponding to the decomposed data. Specifically, the following steps are included: According to the order of input decomposition data, the preset linear transformation is performed on each decomposition data with the user target attention weight in the preset necessary parameters to generate the corresponding Q, K, and V vectors.

7. The method for applying a deep learning processor architecture system with an attention mechanism according to claim 4, wherein: The attention score calculation module is controlled by the FSM state machine according to preset steps to calculate the attention score of the Q, K, and V vectors corresponding to the decomposed data. It is specifically used for: Perform attention score calculation on the Q, K, and V vectors corresponding to each input decomposition data in 4 steps: S0, receives the Q, K, V vectors corresponding to the decomposed data; S1, transmit the Q, K, V vectors corresponding to the decomposed data to the attention score calculation module, control the target operator of the attention score calculation module to perform the dot product calculation of Q and K, and obtain the result R1; S2, controls the target operator of the attention score calculation module to perform the Softmax function normalization calculation and obtains the attention weight R2; S3, controls the target operator of the attention score calculation module to perform weighted sum calculation of attention weights, and performs weighted summation of attention weights and corresponding V vectors to obtain weighted sum vector R3; The data to be processed, R1, R2, and R3 are aggregated and spliced one by one as the attention score.

8. The method for applying a deep learning processor architecture system with an attention mechanism according to claim 7, wherein: The Q, K, and V vectors corresponding to the decomposed data are transmitted to the attention score calculation module, and the target operator of the attention score calculation module is controlled to perform the dot product calculation of Q and K to obtain the result R1, which specifically includes: The similarity is calculated by the formula R1= ; Where i represents the number of columns of the Q vector, and j represents the number of rows of the K vector; The target operator of the control attention score calculation module performs the Softmax function normalization calculation to obtain the attention weight R2, which specifically includes: By formula: , normalized; among them, Indicates the dimensions of Q and K; By formula: , calculate the attention weight R2; The target operator of the control attention score calculation module performs the weighted sum calculation of the attention weight. The attention weight is weighted and summed with the corresponding V vector to obtain the weighted sum vector R3, which specifically includes: By formula: , and obtain the weighted sum vector R3.

9. The method for applying a deep learning processor architecture system with an attention mechanism according to claim 4, wherein: According to the preset allocation scheduling strategy in the implicit attention mechanism scheduling unit or the explicit attention mechanism scheduling unit, the data to be processed and the corresponding attention scores are sent to the computing unit control section in sequence, including: The implicit attention mechanism scheduling unit or the explicit attention mechanism scheduling unit sends the data to be processed and the corresponding attention scores to the target computing unit of the computing unit control module in sequence according to the preset allocation scheduling strategy, and the target computing unit performs the corresponding deep learning data processing and calculation.

10. The method for applying a deep learning processor architecture system with an attention mechanism according to claim 4, wherein: Utilize the computing units and deep learning algorithms in the computing unit control module to obtain calculation results and write the output results back to the host. Specifically, this includes: After the target computing unit in the computing unit control module completes the computing task, the computing result is temporarily stored in the consistent cache of the computing unit control module and written back to the host terminal in the form of DMA through the PCIe interface; The CPU of the host terminal merges and processes the calculation instructions and the calculation results of each target calculation unit corresponding to the attention score allocation, verifies the calculation results according to the preset verification rules, and confirms that the calculation is completed when the preset verification rules are met.

Citation Information

Patent Citations

  • Graph attention network algorithm with optimized calculation sequence and hardware accelerator thereof

    CN118333116A

  • Accelerator supporting attention mechanism

    CN119670639A

  • Visual Transform accelerator implementation method and system

    CN119692408A

  • Attention fusion processing-in-memory architecture for transformer acceleration with triple sparsity-handling

    KR102802168B1

  • Device for accelerating self-attention operation in neural networks

    US20230161783A1