Matrix multiplication pipeline calculation method and device, AI chip, electronic equipment and medium

By triggering general computing components and dedicated computing components in parallel in the AI ​​chip, the parallel execution of matrix element type conversion and matrix multiplication calculation is achieved, which solves the problem of low efficiency of matrix multiplication calculation in the prior art and improves the performance of AI chip.

CN119988808APending Publication Date: 2025-05-13KUNWANG (SHANGHAI) TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510080918.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In matrix multiplication calculation, data type conversion and matrix multiplication calculation need to be performed serially, the hardware cannot achieve pipeline operation, resulting in waste of performance of special computing components and reduce computing efficiency and the performance of AI chips.

Method used

By triggering multiple general computing components and dedicated computing components in parallel in the AI ​​chip, the general computing components perform type conversion of matrix elements under each conversion round and identify the conversion completion status in the identification component; the dedicated computing components query the identification component under each calculation round and obtain the converted matrix elements for matrix multiplication calculation.

Benefits of technology

The parallel execution of general computing components and special computing components is realized, so that the data type conversion time matches the matrix multiplication calculation time, making full use of the performance of special computing components, significantly improving the speed and efficiency of matrix multiplication calculation, and improving the performance of AI chips.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988808A_ABST
    Figure CN119988808A_ABST
Patent Text Reader

Abstract

The invention provides a matrix multiplication pipeline calculation method and device, an AI chip, electronic equipment, a medium and a program product, and relates to the field of artificial intelligence, in particular to the field of chips. According to the specific implementation scheme, in response to a matrix multiplication request, a plurality of general calculation components and a plurality of special calculation components in a chip are triggered in parallel, and in each conversion round, the general calculation components obtain matrix element types required by conversion from a first matrix and a second matrix and identify a conversion completion state; and under each calculation round, when the special calculation part queries that each conversion matrix element needing to be calculated in the current calculation round is converted, each conversion matrix element is acquired from the matched general calculation part to execute matrix multiplication calculation. The general calculation part and the special calculation part are executed in parallel, so that the time consumption of data type conversion and the time consumption of matrix multiplication can be mutually masked, the performance of the special calculation part is fully exerted, and the efficiency of matrix multiplication calculation and the performance of an AI chip are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, in particular to technical fields such as artificial intelligence and chips, and specifically to a pipeline calculation method for matrix multiplication, a pipeline calculation device for matrix multiplication, an AI chip, an electronic device, a non-transitory computer-readable storage medium, and a computer program product. Background Art

[0002] AI (Artificial Intelligence) chips are equipped with specialized matrix multiplication components, which give them a natural advantage in accelerating matrix multiplication. However, these matrix multiplication components usually have specific input data type requirements. In practical applications, due to the variety of input data types, type conversion is often required before the matrix multiplication component can be started for calculation. Summary of the invention

[0003] The present disclosure provides a pipeline computing method for matrix multiplication, a pipeline computing device for matrix multiplication, an AI chip, an electronic device, a non-transitory computer-readable storage medium, and a computer program product.

[0004] According to one aspect of the present disclosure, a pipeline computing method for matrix multiplication is provided, which is executed by an AI chip and includes:

[0005] In response to a matrix multiplication request for a first matrix and a second matrix, a plurality of general-purpose computing components and a plurality of special-purpose computing components in the AI ​​chip are triggered in parallel to perform the following operations:

[0006] The general computing component obtains each matrix element to be converted from the first matrix and the second matrix in each conversion round, performs type conversion, obtains each conversion matrix element, and identifies the conversion completion status of each conversion matrix element in the identification component in the AI ​​chip;

[0007] In each calculation round, if the conversion matrix elements required for calculation in the current calculation round are found in the identification component and the conversion is completed, the conversion matrix elements are obtained from the matching general calculation component to perform matrix multiplication calculation.

[0008] According to another aspect of the present disclosure, there is provided an AI chip, comprising: a plurality of general-purpose computing components, a plurality of special-purpose computing components, and an identification component;

[0009] The general computing component is used to implement type conversion of input matrices in matrix multiplication computing scenarios;

[0010] The dedicated computing component is used to implement matrix multiplication calculation on the input matrix in the matrix multiplication calculation scenario;

[0011] The identification component is used to identify the conversion status of each matrix element in the input matrix by the general computing component in the matrix multiplication calculation scenario;

[0012] Through the cooperation of the general computing component and the special computing component, the pipeline computing method of matrix multiplication as described in any one of the embodiments of the present disclosure is jointly implemented.

[0013] According to another aspect of the present disclosure, an electronic device is provided, comprising the AI ​​chip as described in any one of the embodiments of the present disclosure.

[0014] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the pipeline calculation method of matrix multiplication according to any one of the embodiments of the present disclosure.

[0015] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0017] Figure 1 is a schematic diagram of a pipeline calculation method for matrix multiplication provided according to an embodiment of the present disclosure;

[0018] Figure 2 is a schematic diagram of another pipeline calculation method for matrix multiplication provided according to an embodiment of the present disclosure;

[0019] Figure 3 is a schematic diagram of another pipeline calculation method for matrix multiplication provided according to an embodiment of the present disclosure;

[0020] Figure 4 is a schematic diagram of another pipeline calculation method for matrix multiplication provided according to an embodiment of the present disclosure;

[0021] Figure 5 is a schematic diagram of a pipeline calculation of matrix multiplication using C0 and S0 as an example applicable to an embodiment of the present disclosure;

[0022] Figure 6 is a schematic diagram of a pipeline computing device for matrix multiplication provided according to an embodiment of the present disclosure;

[0023] Figure 7 is a schematic diagram of the structure of an AI chip provided according to an embodiment of the present disclosure;

[0024] Figure 8It is a block diagram of an electronic device used to implement the pipeline calculation method of matrix multiplication according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0025] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0026] In the related art, when it is necessary to perform matrix multiplication calculation on two input matrices, each matrix in the input matrix needs to be converted into a specific data type adapted by the dedicated computing component through each dedicated computing component in the AI ​​chip, and after completing the data type conversion, each dedicated computing component performs the corresponding matrix multiplication calculation.

[0027] The data type conversion of matrix elements and matrix multiplication calculations usually need to be executed serially. Data type conversion mainly involves frequent memory accesses and is a memory-intensive task. Its speed is limited by memory bandwidth and access latency, and it usually takes a long time. Matrix multiplication calculations are computationally intensive tasks that mainly rely on the computing power of the processor and are executed quickly. This dependency makes it impossible for the hardware to achieve pipeline operation, resulting in performance waste of dedicated computing components, reducing computing efficiency and the performance of AI chips.

[0028] Figure 1 It is a schematic diagram of a pipeline calculation method for matrix multiplication provided according to an embodiment of the present disclosure. The embodiment of the present disclosure can be applied to the case where the type conversion of the input matrix and the matrix multiplication calculation are jointly realized through the pipeline cooperation of various general-purpose computing components and various special-purpose computing components in the AI ​​chip. The method can be executed by a pipeline calculation device for matrix multiplication, which can be implemented in hardware and / or software and can generally be configured in the AI ​​chip.

[0029] Correspondingly, such as Figure 1 As shown, the method may specifically include:

[0030] S110. In response to a matrix multiplication request for a first matrix and a second matrix, trigger multiple general-purpose computing components and multiple special-purpose computing components within the AI ​​chip in parallel.

[0031] In this embodiment, the matrix multiplication request can be understood as a request for calculating the matrix multiplication between the first matrix and the second matrix. In the matrix multiplication request, the identification information of the first matrix and the second matrix is ​​generally carried, which is used to locate the first matrix and the second matrix in the storage unit of the AI ​​chip.

[0032] The first matrix and the second matrix can be specifically understood as: two matrix data input into the matrix multiplication operation, where the first matrix is ​​the matrix on the left (also called the left operand) and the second matrix is ​​the matrix on the right (also called the right operand). For example, if you want to calculate the multiplication of matrix A by matrix B, then matrix A is the first matrix and matrix B is the second matrix.

[0033] The general computing component can be specifically understood as: a hardware computing unit in an AI chip that performs various types of computing tasks (for example, data operations or logical judgments, etc.), such as a CPU (Central Processing Unit) or an FPGA (Field Programmable Gate Array). Correspondingly, the dedicated computing component can be specifically understood as a hardware unit in an AI chip that is specially designed for specific computing tasks and has higher performance and efficiency when performing these computing tasks. For example, a TPU (Tensor Processing Unit) or an NPU (Neural Processing Unit). In the calculation scenario of matrix multiplication, the dedicated computing component is used to perform the core computing operations of matrix multiplication, such as multiplication and accumulation operations.

[0034] Specifically, when the AI ​​chip receives a matrix multiplication calculation request, this request may come from external software or application, and the request specifies the two matrices that need to be calculated (the first matrix and the second matrix). In order to speed up the calculation process of matrix multiplication, the AI ​​chip will simultaneously start multiple general-purpose computing components and multiple special-purpose computing components to process this request. The general-purpose computing components are responsible for tasks such as data preprocessing and type conversion, and the special-purpose computing components are responsible for performing the core computing operations of matrix multiplication.

[0035] In this embodiment, the general computing components that respond to the matrix multiplication request may be all or part of the general computing components in the AI ​​chip, and the special computing components that respond to the matrix multiplication request may be all or part of the special computing components in the AI ​​chip. The number of general computing components and special computing components that respond to the matrix multiplication request may be dynamically determined in real time according to the computing tasks currently required to be performed in the AI ​​chip, or the number of general computing components and special computing components used to respond to the matrix multiplication request may be fixedly written in a pre-built configuration file.

[0036] Furthermore, the number of general computing components and special computing components triggered in parallel may be the same or different.

[0037] S120. In each conversion round, the general computing component obtains each matrix element to be converted from the first matrix and the second matrix, performs type conversion, obtains each conversion matrix element, and identifies the conversion completion status of each conversion matrix element in the identification component in the AI ​​chip.

[0038] Since the matrix may contain a large number of elements, it is necessary to gradually complete the conversion of all elements in multiple rounds. The conversion round can be specifically understood as: during the matrix multiplication calculation process, the general computing component performs multiple stages or batches of type conversion on the matrix elements.

[0039] Specifically, before starting the matrix multiplication calculation, the identification component in the AI ​​chip is initialized, and an identification bit or status code is prepared for each matrix element or a group of matrix elements to be converted in a conversion round. The initial state indicates that the conversion is not completed. According to the size of the matrix and the processing power of the chip, the matrix elements are divided into multiple batches or rounds.

[0040] In each round, each general computing component that is triggered to execute reads a certain number of matrix elements from the first matrix and the second matrix, and converts the obtained elements from the original type to the target type according to the data type requirements of the special computing component. For example, the integer type is converted to the floating point type, or a floating point number of one precision is converted to a floating point number of another precision. After the type conversion, the conversion matrix elements that meet the requirements of the subsequent matrix multiplication calculation and can be directly used for the matrix multiplication operation are obtained. At the same time, if the type conversion of a matrix element or a group of matrix elements is completed, the general computing component will notify the identification component to update the conversion completion status of the matrix element or the group of matrix elements. The identification component sets the corresponding identification bit or status code to a state indicating that the conversion has been completed. The general computing component that completes this round of data conversion will continue to perform type conversion for the next conversion round until all matrix elements have completed type conversion.

[0041] S130. In each calculation round, if the conversion matrix elements required for the current calculation round are found in the identification component to have completed the conversion, the conversion matrix elements are obtained from the matching general calculation component to perform matrix multiplication calculation.

[0042] Specifically, each dedicated computing component may be matched with one or more general computing components, and these general computing components are responsible for providing the corresponding dedicated computing components with transformed matrix elements, that is, transformation matrix elements.

[0043] Accordingly, the matrix multiplication calculation can be divided into multiple calculation rounds according to the size of the matrix and the processing power of the chip. In each calculation round, each dedicated computing component needs to obtain all the conversion matrix elements required for this calculation from one or more matching general computing components and perform matrix multiplication operations. First, the dedicated computing component will query the identification component in real time, such as reading the corresponding identification bit or status code in the identification component to check whether the conversion matrix elements required to be calculated in the current calculation round have completed the conversion. If the query result shows that the conversion matrix elements required to be calculated have completed the conversion, the dedicated computing component will obtain these conversion matrix elements from the matching general computing component, perform multiplication and accumulation operations on these elements, and obtain the calculation result of the current calculation round.

[0044] At the same time, after completing the current calculation round, the dedicated calculation component can update the calculation progress information, record the number of elements and results that have been calculated, and then continue to the next calculation round until all matrix elements are calculated and the final matrix multiplication result is obtained.

[0045] Based on this, the technical solution of the embodiment of the present disclosure triggers multiple general-purpose computing components and multiple special-purpose computing components in the AI ​​chip in parallel in response to a matrix multiplication request for the first matrix and the second matrix, and performs the following operations: through the general-purpose computing component, in each conversion round, each matrix element required for conversion is obtained from the first matrix and the second matrix to perform type conversion, and each conversion matrix element is obtained, and the conversion completion status of each conversion matrix element is identified in the identification component in the AI ​​chip; through the special-purpose computing component, in each calculation round, if the conversion matrix elements required to be calculated in the current calculation round are queried in the identification component and the conversion is completed, then each conversion matrix element is obtained from the matching general-purpose computing component to perform matrix multiplication calculation. Through the parallel execution of multiple general-purpose computing components and special-purpose computing components, the data conversion time of the general-purpose computing components and the matrix multiplication calculation time of the special-purpose computing components are overlapped to a certain extent in a hardware pipeline manner, achieving the effect of performing type conversion while performing matrix multiplication calculations. This implementation method of using only special-purpose computing components to perform special matrix multiplication calculations can make full use of the performance of special-purpose computing components, significantly improve the speed and efficiency of matrix multiplication calculations, and can efficiently complete large-scale data conversion and matrix multiplication calculations, thereby improving the performance of AI chips.

[0046] Figure 2It is a schematic diagram of another pipeline calculation method for matrix multiplication provided according to an embodiment of the present disclosure. This embodiment refines the operation of "in response to a matrix multiplication request for a first matrix and a second matrix, triggering multiple general computing components and multiple special computing components in the AI ​​chip in parallel" in the above embodiment, specifically, before the parallel triggering of multiple general computing components and multiple special computing components in the AI ​​chip, it can also include: in the matrix multiplication calculation request, obtaining the first matrix size of the first matrix and the second matrix size of the second matrix; according to the first matrix size, the second matrix size and the performance parameters of the AI ​​chip, determining the matrix block strategy for the first matrix and the second matrix; according to the matrix block strategy and the number of each general computing component involved in the type conversion, determining the matrix elements that each general computing component needs to convert in each conversion round; according to the matrix block strategy and the number of each special computing component involved in the matrix multiplication calculation, determining the conversion matrix elements that each special computing component needs to calculate in each calculation round.

[0047] Correspondingly, such as Figure 2 As shown, the method may specifically include:

[0048] S210, responding to a matrix multiplication request for a first matrix and a second matrix.

[0049] S220. In the matrix multiplication calculation request, obtain a first matrix size of the first matrix and a second matrix size of the second matrix.

[0050] The first matrix size and the second matrix size can be specifically understood as: the number of rows and columns corresponding to the first matrix and the second matrix, respectively. By obtaining these two size information and comparing whether the number of columns of the first matrix is ​​equal to the number of rows of the second matrix, it can be determined whether matrix multiplication is feasible and the size of the result matrix can be calculated. For example, if the size of the first matrix is ​​m*k and the size of the second matrix is ​​k*n, then they can perform matrix multiplication calculations and a result matrix of size m*n can be calculated. Among them, m, k, and n are all positive integers.

[0051] S230. Determine a matrix blocking strategy for the first matrix and the second matrix according to the first matrix size, the second matrix size, and performance parameters of the AI ​​chip.

[0052] The performance parameters of AI chips can specifically include: memory bandwidth and parallel computing capability. Among them, memory bandwidth limits the speed of data transmission. Larger bandwidth can support larger data block transmission. Parallel computing capability determines the number of data blocks that can be processed simultaneously. Larger blocks can reduce the number of data transmissions, but may increase the complexity of calculations.

[0053] Optionally, the result matrix size can be determined according to the first matrix size and the second matrix size, and based on the first matrix size, the second matrix size and the result matrix scale, a test program is used to evaluate the performance of the first matrix and the second matrix of different block sizes. By testing the data type conversion time, data transmission time and matrix multiplication calculation time, the best balance point is found so that the sum of the data type conversion time and data transmission time of a single general computing component is close to or equal to the matrix multiplication calculation time of a single dedicated computing component. The closest division method between the two is used as the final matrix block strategy to maximize the efficiency of parallel computing and achieve a balance between computing and data transmission.

[0054] S240. Determine the matrix elements that need to be converted by each general computing component in each conversion round according to the matrix partitioning strategy and the number of the general computing components involved in the type conversion.

[0055] S250. Determine, according to the matrix partitioning strategy and the number of the dedicated computing components involved in the matrix multiplication calculation, the transformation matrix elements that each dedicated computing component needs to calculate in each calculation round.

[0056] Specifically, in the matrix block strategy, it is assumed that the result matrix is ​​divided into multiple sub-blocks, and the size of each sub-block is b*b, where b is the block size. For example, if the size of the result matrix is ​​m*n, it can be divided into (m / b)*(n / b) sub-blocks. According to the number of dedicated computing components involved in the calculation, the number of result matrix sub-blocks that each dedicated computing component is responsible for calculating can be determined.

[0057] For example, if there are P dedicated computing components, the sub-blocks of the result matrix can be evenly distributed to these dedicated computing components for calculation, that is, each dedicated computing component is responsible for processing (m / b)*(n / b) / P computing tasks.

[0058] The division method of the first matrix and the second matrix needs to match the division method of the result matrix. For the first matrix (size is m*k), it can be divided into (m / b)*(k / b) sub-blocks; for the second matrix (size is k*n), it can be divided into (k / b)*(n / b) sub-blocks. Similarly, according to the number of each of the general computing components involved in the type conversion, the number of sub-blocks for which each general computing component is responsible for the data type conversion of the first matrix and the second matrix can be determined.

[0059] For example, if there are Q general computing units, the sub-blocks of the first matrix and the second matrix can be evenly distributed to these general computing units, that is, each general computing unit is responsible for processing ((m / b)*(k / b)+(k / b)*(n / b)) / Q data type conversion tasks, for example, the first general computing unit is responsible for converting one or more sub-blocks of the first matrix and one or more sub-blocks of the second matrix, the second general computing unit is responsible for converting the remaining one or more sub-blocks of the first matrix and the remaining one or more sub-blocks of the second matrix, and so on. The general computing unit performs data type conversion on each sub-block in sequence, and the dedicated computing unit calculates the matrix multiplication of each sub-block of the first matrix processed by the general computing unit and the sub-block corresponding to the second matrix in sequence.

[0060] S260. Trigger multiple general computing components and multiple special computing components in the AI ​​chip in parallel.

[0061] S270. In each conversion round, the general computing component obtains each matrix element to be converted from the first matrix and the second matrix, performs type conversion, obtains each conversion matrix element, and identifies the conversion completion status of each conversion matrix element in the identification component in the AI ​​chip.

[0062] S280. In each calculation round, if the conversion matrix elements required for the current calculation round are found in the identification component to have completed the conversion, the conversion matrix elements are obtained from the matching general calculation component to perform matrix multiplication calculation.

[0063] Based on this, the technical solution of the embodiment of the present disclosure determines the matrix partitioning strategy for the first matrix and the second matrix through the first matrix size, the second matrix size and the performance parameters of the AI ​​chip, and determines the matrix elements that each general computing component needs to convert in each conversion round in combination with the number of each general computing component involved in the type conversion. Combined with the number of each special computing component involved in the matrix multiplication calculation, determine the conversion matrix elements that each special computing component needs to calculate in each calculation round. Through a reasonable partitioning strategy, the computing resources and memory bandwidth of the AI ​​chip can be better utilized, and the time overhead of data transmission and conversion can be reduced, thereby improving the overall computing efficiency, reducing the delay of data conversion and transmission, making the calculation process of matrix multiplication smoother, and reducing the idle time caused by waiting for data. Through parallel processing and reasonable resource allocation, the general computing components and special computing components in the AI ​​chip can work together, improving the utilization of hardware resources and improving the performance of the AI ​​chip.

[0064] In an optional implementation of this embodiment, determining the matrix partitioning strategy for the first matrix and the second matrix according to the first matrix size, the second matrix size, and the performance parameters of the AI ​​chip may include:

[0065] Determining a third matrix size of a matrix multiplication result matrix according to the first matrix size and the second matrix size;

[0066] Determining a block strategy for the matrix multiplication result matrix according to a data loading speed parameter of the general computing component, a type conversion speed parameter, a calculation speed parameter of the special computing component, and a size of the third matrix;

[0067] According to the blocking strategy of the matrix resulting from the matrix multiplication, a matrix blocking strategy for the first matrix and the second matrix is ​​determined.

[0068] Among them, the size of the third matrix can be specifically understood as: the size of the matrix of the matrix multiplication result after the matrix multiplication operation is performed on the first matrix and the second matrix. The data loading speed parameters of the general computing component, such as bandwidth, affect the transmission speed of data from the memory to the computing unit. The type conversion speed parameters, such as the computing power of the general computing component, affect the processing speed of the type conversion of data before entering the computing unit. The calculation speed parameters of the special computing component, such as the computing power of the special computing component, affect the speed of the actual matrix multiplication calculation. According to the above parameters and the size of the third matrix, select a suitable block size to improve the matrix calculation efficiency. For example, if the calculation speed is fast but the data loading speed is slow, you can choose a block strategy with larger blocks to reduce the number of loading times.

[0069] Use simulation tools or test programs to evaluate the performance of different blocking strategies. Test the data type conversion time, transmission time, and calculation time to find the best balance point so that the sum of the calculation time, data type conversion time, and transmission time is close to or equal, that is, the calculation speed of the dedicated computing component matches the data loading speed and type conversion speed of the general computing component, and the corresponding division method is used as the blocking strategy of the matrix multiplication result matrix. Then, based on the blocking strategy of the matrix multiplication result matrix, the matrix blocking strategy of the first matrix and the second matrix can be obtained.

[0070] By considering the data loading speed, type conversion speed and computing speed of the general computing components, the relationship between data transmission and the computing power of each component can be balanced when determining the block strategy, thereby improving the overall computing efficiency. By selecting an appropriate block strategy, the number of data transmissions from the memory to the computing unit can be reduced, the delay in data loading can be reduced, and the data type conversion can be matched with the processing power of the computing unit to avoid idle computing units due to slow preprocessing speed. In addition, the block strategy can also be adjusted according to the computing speed of the dedicated computing component to ensure that computing resources are fully utilized, avoid frequent waiting for data loading and preprocessing to be completed due to too fast computing speed, achieve efficient collaboration between general computing components and dedicated computing components, reduce overall computing delays, and improve the performance and resource utilization of AI chips.

[0071] In an optional implementation of this embodiment, determining the block strategy of the matrix multiplication result matrix according to the data loading speed parameter of the general computing component, the type conversion speed parameter, the calculation speed parameter of the special computing component and the size of the third matrix may specifically include:

[0072] The objective function is to take the total time of data loading and type conversion of a single general computing component in each conversion round and the matrix multiplication calculation time of a single special computing component in each calculation round as the consistency;

[0073] Under the objective function, the blocking strategy of the matrix multiplication result matrix is ​​determined according to the data loading speed parameter of the general computing component, the type conversion speed parameter, the calculation speed parameter of the special computing component and the third matrix size.

[0074] Specifically, the objective function can use the absolute value of the difference between the matrix multiplication calculation time of a single dedicated computing component in each calculation round and the total time of data loading and type conversion of a single general computing component in each conversion round to minimize the absolute value of the difference, and ensure that the total time of data loading and type conversion of a single general computing component in each conversion round is consistent with the matrix multiplication calculation time of a single dedicated computing component in each calculation round. For example, assuming that the computing power of p general computing components (taking into account the data transmission and data type conversion speed) is x, the computing power of p dedicated computing components is y, and the block strategy is b, that is, each sub-block is b*b in size, then the loading time required for the general computing component is ((m / b)*(k / b)+(k / b)*(n / b)) / (p*x), and the computing time required for the dedicated computing component is (m / b)*(n / b) / (p*y). When the absolute value of the two time differences is minimized, the corresponding block strategy is the optimal strategy.

[0075] In addition, other forms can be used. For example, a ratio form can be used to make the ratio of the total time of data loading and type conversion to the computing time as close to 1 as possible, thereby ensuring that the time of the two is as equal as possible. Alternatively, a weighted sum form can be used to assign different weight coefficients to data loading, type conversion, and computing time, and adjust their importance according to actual conditions to achieve an overall balance, thereby ensuring that the resources of general computing components and special computing components are fully utilized, avoiding resource waste caused by data transmission or computing imbalance, reducing the total time of data transmission and processing, thereby reducing the overall computing delay, achieving efficient collaboration between general computing components and special computing components, and improving the performance and resource utilization of AI chips.

[0076] In an optional implementation of this embodiment, the number of the general computing components participating in the type conversion is the same as or different from the number of the special computing components participating in the matrix multiplication calculation; one of the special computing components obtains the conversion matrix elements required to be calculated from one or more of the general computing components to perform the matrix multiplication calculation;

[0077] The identification component includes a counter, and the counter includes a plurality of counting units, and one counting unit corresponds to a general calculation component involved in type conversion.

[0078] Specifically, in an optional embodiment, the number of general computing components involved in type conversion in the AI ​​chip may be the same as or different from the number of special computing components involved in matrix multiplication calculations, depending on the design of the chip and the requirements of the computing task. For example, if the amount of matrix multiplication calculation is very large, more special computing components are required to process in parallel. Similarly, the number of general computing components is determined according to the requirements of data type conversion. When performing matrix multiplication calculations, the special computing component can obtain the required conversion matrix elements from one or more corresponding general computing components. Multiple general computing components can perform type conversion at the same time and provide the converted data to the corresponding special computing component. The identification component includes a counter, which includes multiple counting units, each counting unit corresponding to a general computing component participating in the type conversion. The counting unit is used to record the number of matrix elements for which each general computing component completes the type conversion.

[0079] By flexibly configuring the number of general-purpose computing components and special-purpose computing components, and allowing special-purpose computing components to obtain data from multiple general-purpose computing components, efficient use of computing resources can be achieved, waiting time and data transmission overhead can be reduced, thereby improving overall computing efficiency, and enabling AI chips to flexibly adjust resource allocation according to different computing tasks and data characteristics, adapt to computing needs of various scales and complexities, and enhance the adaptability and versatility of the chip. The counters and counting units of the identification components provide a real-time monitoring and coordination mechanism for the data processing process, ensuring that data conversion and the calculation process are carried out synchronously. Special-purpose computing components can obtain the required elements from the general-purpose computing components in a timely manner for matrix multiplication calculations, optimizing the entire data processing process, avoiding computing interruptions and resource waste due to data mismatch or waiting for data, thereby improving the reliability of the chip system.

[0080] In another optional implementation of this embodiment, the number of the general computing components participating in the type conversion is the same as the number of the special computing components participating in the matrix multiplication calculation; one special computing component only obtains the conversion matrix elements required to be calculated from one general computing component to perform the matrix multiplication calculation;

[0081] The identification component includes a first counter and a second counter, the first counter includes multiple first counting units, one first counting unit corresponds to a general computing component involved in type conversion, and the second counter includes multiple second counting units, one second counting unit corresponds to a special computing component involved in matrix multiplication calculation.

[0082] Specifically, in another optional embodiment, the number of general computing components involved in type conversion is the same as the number of special computing components involved in matrix multiplication calculation, which means that each general computing component corresponds to a special computing component. The one-to-one mapping relationship simplifies the data flow and control logic, and each special computing component obtains the required conversion matrix elements from only one general computing component. Each first counting unit in the first counter corresponds to a general computing component, and is used to record the number of matrix elements for which the general computing component completes type conversion. Each second counting unit in the second counter corresponds to a special computing component, and is used to record the number of matrix elements for which the special computing component completes matrix multiplication calculation.

[0083] Since each dedicated computing component only obtains data from one general computing component, the source and flow of the data are clear, ensuring data consistency and traceability, reducing calculation errors caused by data errors or inconsistencies, and the one-to-one mapping relationship simplifies system design and scheduling logic, reduces system complexity, makes data flow and control flow clearer and easier to manage, avoids repeated transmission and storage of data between multiple components, reduces data redundancy and storage overhead, and improves the storage efficiency of the system. Through real-time monitoring of the first counter and the second counter, the working status of the general computing component and the dedicated computing component can be timely understood. The parallel and collaborative execution of the general computing component and the dedicated computing component matches the data conversion time of the general computing component with the matrix multiplication calculation time of the dedicated computing component, and the two are roughly synchronized, making full use of the performance of the dedicated computing component and improving the performance of the AI ​​chip.

[0084] Figure 3 It is a schematic diagram of another pipeline calculation method for matrix multiplication provided according to an embodiment of the present disclosure. This embodiment is an optional implementation method when "the number of each general computing component participating in the type conversion is the same or different from the number of each special computing component participating in the matrix multiplication calculation". It is a refinement of the operation of "obtaining each matrix element required to be converted from the first matrix and the second matrix in each conversion round through the general computing component to perform type conversion, obtain each conversion matrix element, and identify the conversion completion status of each conversion matrix element in the identification component in the AI ​​chip" in the above embodiment. Specifically, it may include: after the current general computing component completes the type conversion of each matrix element required to be converted in the current conversion round, the count value in the counting unit in the counter that matches the current general computing component is updated.

[0085] Correspondingly, such as Figure 3 As shown, the method may specifically include:

[0086] S310. In response to a matrix multiplication request for a first matrix and a second matrix, trigger multiple general computing components and multiple special computing components within the AI ​​chip in parallel.

[0087] S320, after completing the type conversion of each matrix element to be converted in the current conversion round through the current general computing component, update the count value in the counting unit in the counter that matches the current general computing component.

[0088] Specifically, in the current conversion round, when the general computing component completes the type conversion of the matrix element to be converted, the count value in the counting unit matching the current general computing component in the counter is increased, and each time the data type conversion of an element is completed, the count value is increased by 1 to reflect that the current general computing component has successfully converted a certain number of matrix elements. The current count value in the counting unit represents the number of matrix elements that the current general computing component has successfully converted.

[0089] S330. In each calculation round, if the conversion matrix elements required for the current calculation round are found in the identification component to have completed the conversion, the conversion matrix elements are obtained from the matching general calculation component to perform matrix multiplication calculation.

[0090] Based on this, the technical solution of the embodiment of the present disclosure can accurately grasp the number of matrix elements that have completed the type conversion by updating the count value in the counting unit that matches the current general computing component, so as to reasonably arrange the subsequent matrix multiplication calculation tasks. When a sufficient number of elements have completed the conversion, the special computing component can obtain the required data and perform calculations, reducing the idle time caused by waiting for the data conversion to be completed. The general computing component and the special computing component can work together better. The parallel execution of the general computing component and the special computing component makes the data conversion time of the general computing component match the matrix multiplication calculation time of the special computing component. The two are roughly synchronized, and the performance of the special computing component can be fully utilized. The speed and efficiency of the matrix multiplication calculation can be significantly improved, and the conversion and matrix multiplication calculation of large-scale data can be efficiently completed, which improves the performance of the AI ​​chip.

[0091] In an optional implementation of this embodiment, in each calculation round, if the dedicated calculation component finds in the identification component that each conversion matrix element required to be calculated in the current calculation round is converted, then obtaining each of the conversion matrix elements from the matching general calculation component to perform matrix multiplication calculation may include:

[0092] Determine the target counting unit to be queried and the expected counting value of the target counting unit in the counter according to each conversion matrix element to be calculated in the current calculation round by the current dedicated calculation component;

[0093] querying the target counting unit through the current dedicated computing component, and acquiring a target general-purpose computing component matching the target counting unit when determining that the current count value of the target counting unit is greater than or equal to the expected count value;

[0094] Each conversion matrix element required to be calculated in the current calculation round is obtained in the target general-purpose calculation component through the current dedicated calculation component, and matrix multiplication calculation is performed.

[0095] Specifically, when the current dedicated computing component performs the current computing round, it needs to determine which matrix elements have been converted and are ready for computing, and determine the corresponding target counting units in the counter according to the required conversion matrix elements. Each target counting unit corresponds to a general computing component and records the number of matrix elements that the component has completed type conversion.

[0096] The expected count value can be specifically understood as: the number of matrix elements that have completed type conversion recorded in the target counting unit that the dedicated computing component expects before starting the calculation. This value is set based on the number of matrix elements required for the current calculation round and the total number of elements used previously, ensuring that all required elements have been converted when acquiring data. The current dedicated computing component actively queries the target counting unit to obtain its current count value, understands the progress of the target general computing component in completing type conversion in real time, and determines whether the data is ready. When the query result shows that the current count value of the target counting unit is greater than or equal to the expected count value, it means that all required conversion matrix elements are ready. At this time, the dedicated computing component can obtain the target general computing component that matches the target counting unit, and obtain the conversion matrix elements required to be calculated in the current calculation round from the component, and use the obtained conversion matrix elements to perform the matrix multiplication calculation of the current calculation round.

[0097] By querying the target counting unit and judging the relationship between the current counting value and the expected counting value, the dedicated computing component can ensure that all required transformation matrix elements are ready before starting the calculation, avoiding calculation errors or interruptions caused by incomplete data. The dedicated computing component can obtain the required data in a timely manner, reducing waiting time, and improving overall computing efficiency, making the coordination between general computing components and dedicated computing components closer and more efficient, thereby improving the performance of the AI ​​chip.

[0098] Further, based on the above embodiments, after querying the target counting unit through the current dedicated computing component, the method may further include:

[0099] When the current dedicated computing component determines that the current count value of the target counting unit is less than the expected count value, it returns to execute the operation of querying the target counting unit after a preset waiting time interval until it is determined that the current count value of the target counting unit is greater than or equal to the expected count value.

[0100] Specifically, when the dedicated computing component queries the target counting unit in the current dedicated computing component, if it is determined that the current count value of the target counting unit is less than the expected count value, it means that the required conversion matrix elements are not yet fully prepared, and the dedicated computing component will wait for a preset duration. This waiting duration can be set according to the actual situation, which can range from a few milliseconds to a few seconds, so as to give the general computing component enough time to complete the remaining type conversion tasks and avoid resource waste and system burden caused by frequent queries. After the waiting time is over, the dedicated computing component will query the target counting unit again to check whether its current count value has reached the expected count value. This process will be repeated continuously, that is, if it is found that the count value still does not meet expectations after each query, it will wait for the preset duration again, and then continue to query until the current count value of the target counting unit is greater than or equal to the expected count value. If the current count value of the target counting unit meets the conditions, it means that the required conversion matrix elements are all ready. At this time, the dedicated computing component can obtain these elements and perform the matrix multiplication calculation of the current calculation round.

[0101] By setting a preset waiting time at intervals, the dedicated computing component will not frequently query the target counting unit, thereby avoiding unnecessary resource consumption, such as processor resource occupation and memory resource occupation, improving the overall efficiency and performance of the system, and ensuring that all required conversion matrix elements are ready before the dedicated computing component starts calculating, avoiding calculation errors or interruptions caused by incomplete data, and ensuring the accuracy and reliability of the calculation results. Reasonable waiting time and query mechanism enable the system to run more stably when facing inconsistent or delayed data conversion progress, making the coordination between general computing components and dedicated computing components closer and more efficient, thereby improving the performance of AI chips.

[0102] Figure 4It is a schematic diagram of another pipeline calculation method for matrix multiplication provided according to an embodiment of the present disclosure. This embodiment is another optional implementation method when "the number of each general computing component participating in the type conversion is the same as the number of each special computing component participating in the matrix multiplication calculation". It is a refinement of the operation of "obtaining each matrix element to be converted from the first matrix and the second matrix in each conversion round by the general computing component to perform type conversion to obtain each conversion matrix element, and identifying the conversion completion status of each conversion matrix element in the identification component in the AI ​​chip" in the above embodiment, and specifically may include: after the current general computing component completes the type conversion of each matrix element to be converted in the current conversion round, updating the count value in the current first counting unit in the first counter that matches the current general computing component; after the current general computing component determines that the type conversion of all matrix elements to be converted has been completed, determining the associated general computing component and determining the associated first counting unit corresponding to the associated general computing component in the first counter; and obtaining each conversion matrix element to be loaded from the associated general computing component for local storage and updating the count value of the current first counting unit whenever the current general computing component detects that the count value in the associated first counting unit meets the loading condition.

[0103] Correspondingly, such as Figure 4 As shown, the method may specifically include:

[0104] S410. In response to a matrix multiplication request for a first matrix and a second matrix, trigger multiple general computing components and multiple special computing components within the AI ​​chip in parallel.

[0105] S420: After completing the type conversion of each matrix element to be converted in the current conversion round through the current general computing component, update the count value in the current first counting unit in the first counter that matches the current general computing component.

[0106] The first counter includes a plurality of first counting units, each of which corresponds to a general computing component participating in the type conversion. The current first counting unit is a counting unit that matches the general computing component currently executing the type conversion task, and records the number of matrix elements that have completed the type conversion of the current general computing component.

[0107] Specifically, in the current conversion round, when the general computing component completes the type conversion of each matrix element that needs to be converted, the system will increase the count value in the current first counting unit in the first counter that matches the current general computing component to reflect that the current general computing component has successfully converted a certain number of matrix elements.

[0108] S430: After determining that type conversion of all matrix elements to be converted is completed through the current general computing component, determine an associated general computing component, and determine an associated first counting unit corresponding to the associated general computing component in the first counter.

[0109] In AI chips, multiple general-purpose computing components can work together. Associated general-purpose computing components can be specifically understood as other general-purpose computing components that have data dependencies or collaborations with the current general-purpose computing components. For example, in the process of matrix multiplication, multiple general-purpose computing components may be required to perform type conversion on different matrix blocks respectively, and then aggregate or pass all or part of the converted data to the corresponding dedicated computing components for matrix multiplication calculations.

[0110] Specifically, after the current general computing component determines that all matrix elements that the component is responsible for have been converted from the original data type to a data type suitable for subsequent calculations, it needs to determine the associated general computing component through the communication mechanism or scheduling algorithm inside the chip, for example, to determine the association relationship based on data flow, task allocation or pre-set collaboration rules. After determining the associated general computing component, the current general computing component needs to find the associated first counting unit corresponding to the associated general computing component in the first counter, and the associated first counting unit records the number of matrix elements for which the associated general computing component has completed type conversion.

[0111] S440, whenever the current general computing component detects that the count value in the associated first counting unit meets the loading condition, each conversion matrix element required to be loaded is obtained from the associated general computing component for local storage, and the count value of the current first counting unit is updated.

[0112] The loading condition can be specifically understood as: the count value reaches a preset threshold. For example, when the count value in the associated first counting unit reaches a preset threshold, it means that the associated general computing component has completed a sufficient number of matrix element type conversions, and the current general computing component can start to obtain data from it.

[0113] Specifically, the current general computing component will periodically or in real time detect the count value in the associated first counting unit to determine whether the loading condition is met. Once it is detected that the count value meets the loading condition, the current general computing component will obtain the required loaded conversion matrix elements from the associated general computing component. Among them, these elements have completed type conversion and meet the requirements of subsequent calculations. The obtained conversion matrix elements will be stored in the local storage space of the current general computing component for subsequent calculations. The local storage can be a cache or dedicated storage area inside the chip, which has a faster access speed and helps to improve computing efficiency.

[0114] In addition, after acquiring and storing the conversion matrix elements, the current general computing component needs to update the count value of the current first counting unit corresponding to itself in the first counter. The amount of increase in the count value corresponds to the number of acquired conversion matrix elements to accurately reflect the number of conversion matrix elements prepared by the current general computing component. The system can accurately track the progress of each general computing component to ensure the coordination and consistency of data processing and calculation processes.

[0115] S450. In each calculation round, if the conversion matrix elements required for the current calculation round are found in the identification component to have completed the conversion, the conversion matrix elements are obtained from the matching general calculation component to perform matrix multiplication calculation.

[0116] Based on this, the technical solution of the embodiment of the present disclosure allows the dedicated computing component to obtain the required data from the corresponding general-purpose computing unit or the associated general-purpose computing unit. The multi-path data acquisition method increases the flexibility of data acquisition, so that the dedicated computing component can more conveniently obtain the required data. It allows data to be obtained from associated units, which can better adapt to the distribution of data and avoid the inability of some dedicated computing components to obtain data in time due to uneven data distribution. The dedicated computing component can obtain the required data faster, reducing the idle time caused by waiting for data, making the calculation process more continuous, and making full use of the performance of the dedicated computing component. The speed and efficiency of matrix multiplication calculations are significantly improved, and large-scale data conversion and matrix multiplication calculations can be efficiently completed, thereby improving the performance of the AI ​​chip.

[0117] In an optional implementation of this embodiment, in each calculation round, if the dedicated calculation component finds in the identification component that each conversion matrix element required to be calculated in the current calculation round is converted, then obtaining each of the conversion matrix elements from the matching general calculation component to perform matrix multiplication calculation may include:

[0118] Determine, by means of a current dedicated computing component, an associated general computing component matching the current dedicated computing component, and obtain, in the first counter, a first current value in a current first counting unit matching the associated general computing component;

[0119] Acquire, by the current dedicated computing component, in the second counter, a second current counting value in a current second counting unit matching the current dedicated computing component;

[0120] When the current dedicated computing component determines that the second current count value is less than the first current value, the conversion matrix elements required to be calculated in the current calculation round are obtained from the associated general computing component, matrix multiplication is performed, and the second current value in the current second counting unit is updated.

[0121] The first current value can be specifically understood as the number of matrix elements that have completed type conversion by the associated general computing component, reflecting the data preparation status. The second current count value can be specifically understood as the number of matrix elements that have completed calculation by the current dedicated computing component, reflecting the calculation progress.

[0122] Specifically, the current dedicated computing component compares the second current count value with the first current value. If the second current count value is less than the first current value, it means that the associated general computing component has prepared enough data, but the current dedicated computing component has not yet completed the corresponding calculation. The current dedicated computing component obtains the elements of the transformation matrix required to be calculated in the current calculation round from the associated general computing component. After obtaining the data, the matrix multiplication calculation is performed. After the calculation is completed, the current dedicated computing component updates the second current value in the current second counting unit and increases the corresponding count value to reflect the completed matrix multiplication calculation.

[0123] By comparing the count values ​​to determine whether the data is ready, it is ensured that the dedicated computing component has available data when performing calculations, avoiding calculation delays or interruptions caused by insufficient data. The dedicated computing component can obtain the required data and perform calculations in a timely manner, reducing waiting time and improving overall computing efficiency. The coordination between general computing components and dedicated computing components is closer and more efficient, making full use of the computing resources of the dedicated computing components and improving the performance of the AI ​​chip.

[0124] Further, on the basis of the above embodiments, after obtaining the second current count value in the current second counting unit matching the current dedicated computing component in the second counter through the current dedicated computing component, the method may further include:

[0125] When the current dedicated computing component determines that the second current count value is equal to the first current value, after a preset waiting period, it returns to execute the operation of obtaining the first current value in the current first counting unit that matches the associated general computing component in the first counter until the second current count value is less than the first current value.

[0126] Specifically, when the second current count value is equal to the first current value, it means that the dedicated computing component has used all available data for calculations, and the general computing component has not yet provided new data, and the dedicated computing component needs to wait for a preset period of time. This waiting period is to give the general computing component enough time to complete more type conversion tasks and thus provide new data. After the waiting period is over, the dedicated computing component will re-query the general computing component to see if it has completed more data conversions. After each query, if it is found that the second current count value is still equal to the first current value, it will wait for the preset period of time again, and then continue to query until the second current count value is less than the first current value, that is, the general computing component has provided new data, and the dedicated computing component can continue to perform calculations.

[0127] By setting a preset waiting time at intervals, the dedicated computing component will not query data frequently and continuously, thereby avoiding unnecessary resource consumption, such as processor occupancy and frequent memory access, improving the overall efficiency and performance of the system, and ensuring that the dedicated computing component can continuously obtain new data when performing calculations, avoiding computing stagnation caused by data interruptions. Reasonable waiting time and query mechanism enable the system to run more stably when facing inconsistent or delayed data conversion progress, making the coordination between dedicated computing components and general computing components more efficient, making full use of the computing performance of dedicated computing components, thereby improving the performance of AI chips.

[0128] For ease of understanding, the specific application scenarios applicable to each embodiment of the present disclosure are now described. AI chips have dedicated matrix multiplication components, so AI chips have certain advantages in accelerating matrix multiplication. The matrix multiplication component has specific input type requirements. However, in actual application scenarios, due to the diversity of input data types, type conversion is often required before the matrix multiplication component can be started for matrix multiplication. Data type conversion is usually a memory-intensive task, and its speed is limited by memory bandwidth and access latency, and it usually takes a long time, while matrix multiplication calculation is a computationally intensive task, which mainly depends on the computing power of the processor and has a faster execution speed. When the data type conversion has not been completed, the matrix multiplication component has to be in a waiting state and cannot start the calculation in time, resulting in the inability of the hardware to operate in a pipeline and low computing efficiency. The relevant data type conversion technical solutions are all serial execution solutions. After completing the data type conversion, the matrix multiplication component relies on the converted result to implement the matrix multiplication operation. Although the operation is simple and more intuitive, it causes poor performance, the hardware cannot operate in a pipeline, and the calculation and memory access time cannot cover each other, resulting in a large amount of computing resource waste.

[0129] To solve the above problems, the embodiments of the present disclosure propose a pipeline calculation method for matrix multiplication, which can be used for the simultaneous synchronization of multiple streams of general-purpose computing components and special-purpose computing components on AI chips. Among them, the general-purpose computing component is component A, which is mainly responsible for memory-intensive operations, and the special-purpose computing component is component B, which is mainly responsible for computationally intensive operations, and is particularly good at matrix multiplication operations. Component A performs data type conversion, and at the same time, another computing flow is started to start component B to perform matrix multiplication operations, so as to achieve a state where multiple streams work simultaneously. At the same time, a reasonable data blocking strategy is applied so that the time consumption of data type conversion is basically the same as the time consumption of matrix multiplication, so as to achieve an optimal state in which data type conversion and matrix multiplication cover each other, thereby effectively improving performance. This method can be specifically understood as:

[0130] Assume that the AI ​​chip has p general computing components, denoted as C0, C1, C2... and Cp; and p dedicated computing components, denoted as S0, S1, S2,... and Sp. The scale of the input matrix x is m*k, the scale of the input matrix w is k*n, and the scale of the output matrix y is m*n. A space of length p (hereinafter referred to as counter num_A) is opened up to record the completion of data type conversion by each general computing component, and a space of length p (hereinafter referred to as counter num_B) is opened up to record the matrix multiplication of each dedicated computing component.

[0131] Taking C0 and S0 as an example, C0 is responsible for the type conversion of (m*k+k*n) / p data, and S0 is responsible for the calculation of the matrix multiplication result of m*n / p data. Divide the y matrix into blocks, and the final result is recorded as y0, y1, y2,... and y_(m-1)*n / p; divide the x and w matrices into corresponding blocks. C0 starts to execute the operation, C0 reads the segmented x data and w data, and performs data type conversion. After the data type conversion is completed, C0 triggers the corresponding counting unit num_A[0] in the counter num_A to increase by 1, indicating that there is a block of data that has completed the data type conversion. S0 constantly reads the count values ​​of the corresponding counting units num_B[0] and num_A[0] in the counter num_B0. If the count value of num_B[0] is less than num_A[0], it proves that there is data to be read. S0 then reads the x data and w data that have completed the data type conversion, performs the corresponding matrix multiplication operation, and triggers the count value of num_B[0] to increase by 1 at the same time, proving that the loading of x data and w data has been completed, and starts to execute the corresponding matrix multiplication operation.

[0132] Since the general computing component C0 and the special computing component S0 do not interfere with each other, C0 continues to perform its responsible work and completes the type conversion of the remaining data. Whenever new data completes the type conversion, C0 will trigger num_A[0] to increment by 1; at the same time, S0 continues to read the count values ​​of num_B[0] and num_A[0]. As long as the two are inconsistent, it means that there is new data that has been converted. S0 will read the data, update num_B[0], and complete the corresponding matrix multiplication operation. After S0 completes the calculation of all assigned data conversion tasks, the current num_A[0] records the number of all type conversion rounds completed by C0. However, generally speaking, when C0 completes the data conversion task, S0's matrix multiplication calculation task is generally not completed. For example, it continues to need the data converted by C1 and C2 to perform subsequent matrix multiplication calculations. Assume that the data completed by the first and second rounds of conversion by C1 is required. Furthermore, C0 can detect whether the subsequent calculation data required by S0 is ready by querying the corresponding counting unit num_A[1] in the counter num_A, or querying whether the corresponding counting unit num_B[1] in the counter num_B has been incremented to 1 or incremented to 2.

[0133] For example, if C0 finds that num_B[1]=2, it determines that the two rounds of calculation data that S0 needs to obtain from C1 have been converted. At this time, C0 can obtain all the converted data from C1 at one time for local storage, and trigger num_A[0] to increment by 2. Since S0 reads the count values ​​of num_B[0] and num_A[0] in real time, the difference in the count values ​​between the two can be immediately found. Then, two sets of converted data can be obtained from C0 at one time to continue the matrix multiplication calculation, and num_B[0] can be incremented by 2.

[0134] Based on the same implementation principle as above, the general computing components and the special computing components cooperate in parallel pipelines, and finally, the matrix multiplication operation of the x matrix and the w matrix is ​​completed.

[0135] The above steps can be performed in multiple streams at the same time without any conflicts. Assuming that the more data component C converts each time, the more data component S can calculate. Matrix multiplication can be performed after only some data types are converted, which optimizes a lot of data type conversion time. The memory access and calculation process is masked, and the calculation efficiency is improved.

[0136] If the data is divided reasonably, for example, the data handling and data type conversion computing power of the general computing components in the AI ​​chip are combined with the matrix multiplication computing power of the special computing components. After calculating the total amount of data required to be converted by each general computing hardware and the total amount of data required for matrix multiplication calculation by each special computing component by combining the number of general and special computing components and the input matrix size, the optimal data division strategy is determined with the goal of making the data handling and data conversion time of a single general computing component consistent with the matrix multiplication calculation time of a single special computing component, which can ensure that the time consumption of data type conversion and matrix multiplication is basically consistent, achieving an optimal state of mutual concealment, and the overall performance can be greatly improved. For example, suppose the computing power of p general-purpose computing components (taking into account the speed of data transmission and data type conversion) is x, the computing power of q special-purpose computing components is y, and the block strategy is b, that is, each sub-block is b*b in size. Then the loading time required for the general-purpose computing component is ((m / b)*(k / b)+(k / b)*(n / b)) / (p*x), and the computing time required for the special-purpose computing component is (m / b)*(n / b) / (q*y). When the absolute value of the two time differences is the smallest, the corresponding block strategy is the optimal strategy.

[0137] In addition, in the above method, only one set of counters num_A can be maintained to store the data type conversion status of each general computing component. Before each special computing component implements matrix multiplication calculation in each round, it can first query whether the value of the counting unit of the counter at the corresponding position has reached a preset value (determine the currently required matrix elements and the total matrix elements already used according to the current calculation progress, and obtain the corresponding preset value). If it has reached the value, it indicates that the conversion is completed, and the required data can be obtained from the general computing component corresponding to the counting unit to perform matrix multiplication calculation. If it has not reached the value, the operation of waiting for the preset waiting time and querying whether the value of the counting unit at the corresponding position has reached the preset value is executed in a loop until the preset value is reached, so that a special computing component can obtain the converted data from multiple general computing components.

[0138] Correspondingly, a dedicated computing component can also only obtain the data required for calculation from a general computing component for matrix multiplication calculation. After the general computing component completes all the calculations it is responsible for, it can check whether the data required for subsequent calculations of the corresponding dedicated computing component has been converted by other general computing components (also determined by querying the counting units of num_A). If the conversion is completed, the data can be directly obtained from the corresponding general computing component, and the corresponding count value can be directly updated. For example, C0 can obtain all the conversion completion data from C1 at one time for local storage, and then trigger num_A[0] to increase by 2. Since S0 reads the count value of num_A[0] in real time, it can immediately find that the count value reaches the preset value. Furthermore, two sets of conversion completion data can be obtained from C0 at one time to continue to perform matrix multiplication calculations, and the corresponding preset value can be updated according to the calculation requirements, and the count value of num_A[0] can continue to be read in real time.

[0139] In addition, in order to more conveniently implement parallel pipelining, a set of counters can be maintained to record the matrix multiplication calculation rounds performed by each dedicated computing component. By comparing the difference between the count value of the counter and the count value of the counter of the matching general computing component, it can be determined whether the calculation data required for the current calculation of the dedicated computing component is ready.

[0140] In a specific example, assume that the AI ​​chip has three general computing components, denoted as C0, C1 and C2; and three special computing components, denoted as S0, S1 and S2. The scale of the input matrix x is 6*2000, the scale of the input matrix w is 2000*3, and the scale of the output matrix y is 6*3. A space of length 3 (hereinafter referred to as counter num_A) is opened up to record the completion of data type conversion by each general computing component, and a space of length 3 (hereinafter referred to as counter num_B0) is opened up to record the matrix multiplication of each special computing component. The following takes C0 and S0 as examples. C0 is responsible for the type conversion of 6*2000 / 3+2000 data, and S0 is responsible for the calculation of the matrix multiplication results of 2*3 data. The y matrix is ​​divided into blocks, and the final results are recorded as y0, y1, y2, y3, y4 and y5; the x data is divided into blocks, recorded as x00, x01, x10 and x11; the w data is divided into blocks, recorded as w00 and w10. Figure 5 It is a schematic diagram of a pipeline calculation of matrix multiplication using C0 and S0 as an example applicable to an embodiment of the present disclosure. Figure 5 Here, num_A[0] is the counting unit corresponding to C0, and num_B from left to right are num_B[0], num_B[1], and num_B[2], corresponding to the counting units of S0, S1, and S2, respectively.

[0141] A0 starts to execute the operation. A0 reads x00 and w00 to convert the data type. After the data type conversion is completed, the data becomes casted_x00 and casted_w00, and the counter num_A[0] increments by 1, indicating that there is a piece of data that has completed the data type conversion. Counter num_B[0] reads the value of counter num_A[0] at all times. If counter num_B[0] is less than counter num_A[0], it proves that there is data to be read. Component S0 reads casted_x00 and casted_w00 and performs the corresponding matrix multiplication operation. After the calculation, num_B[0] increments by 1, proving that a partial matrix multiplication calculation has been performed. Since the two components C0 and S0 do not interfere with each other, C0 continues to perform its responsible work and completes the data type conversion of x01, w10, x10 and x11, obtaining casted_x01, casted_w10, casted_x10 and casted_x11. Whenever new data is converted, the counter num_A[0] will increment by 1; meanwhile, num_B[0] will continue to read num_A[0]. As long as there is new data, S0 will read the data and complete the corresponding matrix multiplication operation. When all the data have been read, S0 has completed the calculation of y0 and y3 assigned to it. The counter num_B[0] records the number of matrix multiplication results completed by S0, which is 2. At this time, since the w matrix data required to calculate y1 and y4 is read by S1, and the data required to calculate y2 and y5 is read by S2, it is only necessary to determine the value of the corresponding position on num_B[1] to know whether the corresponding w data exists. For example, num_B[1] of S1 is 2, which proves that S1 has also completed the calculation of the two data. It can be obtained that w10 and w11 have been converted. Therefore, S0 can directly read casted_w10, casted_w11 and casted_x00, casted_x01, casted_x10 and casted_x11 at the corresponding positions to complete the corresponding matrix multiplication and obtain the results y1 and y4.

[0142] Similarly, S0 will also read the w data required for the matrix multiplication that S2 is responsible for according to num_B[2], complete the corresponding matrix multiplication, and obtain the results y2 and y5. Components S1 and S2 also perform the same operation to complete the matrix multiplication operation of the responsible data. Finally, the matrix multiplication operation of the x matrix and the w matrix is ​​completed. By using the corresponding block strategy, the time consumption of C for data type conversion is controlled to be close to the time consumption of S calculation, an optimal state in which type conversion and calculation cover each other can be achieved, which effectively improves the utilization rate of hardware and completes performance optimization.

[0143] A pipeline calculation method for matrix multiplication proposed in the embodiments of the present disclosure can realize the parallel execution of general-purpose computing components and special-purpose computing components, achieve the technical effect that the time consumption of data type conversion and the time consumption of matrix multiplication can cover each other, maximize the computing performance of special-purpose computing components, improve the efficiency of matrix multiplication calculation, and can efficiently complete the conversion and matrix multiplication calculation of large-scale data, especially the conversion and matrix multiplication calculation tasks of large-scale data such as mapping of 3D point cloud data, thereby improving the performance of AI chips.

[0144] As an implementation of the pipeline calculation methods for the above-mentioned matrix multiplications, the present disclosure also provides an optional embodiment of an execution device for implementing the pipeline calculation methods for the above-mentioned matrix multiplications.

[0145] Figure 6 is a schematic diagram of a pipeline computing device for matrix multiplication provided according to an embodiment of the present disclosure, such as Figure 6 As shown, the device includes: a parallel trigger module 610, a type conversion module 620 and a matrix multiplication calculation module 630, wherein:

[0146] The parallel trigger module 610 is used to trigger multiple general computing components and multiple special computing components in the AI ​​chip in parallel in response to a matrix multiplication request for the first matrix and the second matrix, and execute the following modules:

[0147] A type conversion module 620 is used to obtain each matrix element to be converted from the first matrix and the second matrix in each conversion round through the general computing component to perform type conversion, obtain each conversion matrix element, and identify the conversion completion status of each conversion matrix element in the identification component in the AI ​​chip;

[0148] The matrix multiplication calculation module 630 is used to obtain the transformation matrix elements required to be calculated in the current calculation round from the matching general calculation component to perform matrix multiplication calculation in each calculation round through the dedicated calculation component if the transformation matrix elements required to be calculated in the current calculation round are completed in the identification component.

[0149] The technical solution of the disclosed embodiment triggers multiple general computing components and multiple special computing components in the AI ​​chip in parallel in response to the matrix multiplication request for the first matrix and the second matrix, and performs the following operations: the general computing component obtains the matrix elements required to be converted from the first matrix and the second matrix in each conversion round to perform type conversion, obtains each conversion matrix element, and identifies the conversion completion status of each conversion matrix element in the identification component in the AI ​​chip; the special computing component obtains each conversion matrix element from the matching general computing component to perform matrix multiplication calculation if the conversion matrix elements required to be calculated in the current calculation round are queried in the identification component in each calculation round. Through the parallel execution of the general computing component and the special computing component, the data conversion time of the general computing component matches the matrix multiplication calculation time of the special computing component, and the two are roughly synchronized, making full use of the performance of the special computing component, significantly improving the speed and efficiency of the matrix multiplication calculation, and being able to efficiently complete the conversion and matrix multiplication calculation of large-scale data, thereby improving the performance of the AI ​​chip.

[0150] Further, based on the above embodiments, the pipeline computing device for matrix multiplication may further include: a size acquisition module, a strategy determination module, a conversion determination module and a calculation determination module, wherein:

[0151] a size acquisition module, configured to acquire, in the matrix multiplication calculation request, a first matrix size of the first matrix and a second matrix size of the second matrix before the parallel triggering of the plurality of general computing components and the plurality of special computing components in the AI ​​chip;

[0152] A strategy determination module, configured to determine a matrix partitioning strategy for the first matrix and the second matrix according to the first matrix size, the second matrix size, and the performance parameters of the AI ​​chip;

[0153] A conversion determination module, used to determine the matrix elements that each general computing component needs to convert in each conversion round according to the matrix block strategy and the number of each general computing component participating in the type conversion;

[0154] The calculation determination module is used to determine the transformation matrix elements that each dedicated computing component needs to calculate in each calculation round according to the matrix blocking strategy and the number of each dedicated computing component participating in the matrix multiplication calculation.

[0155] On the basis of the above embodiments, the strategy determination module is specifically used to:

[0156] Determining a third matrix size of a matrix multiplication result matrix according to the first matrix size and the second matrix size;

[0157] Determining a block strategy for the matrix multiplication result matrix according to a data loading speed parameter of the general computing component, a type conversion speed parameter, a calculation speed parameter of the special computing component, and a size of the third matrix;

[0158] According to the blocking strategy of the matrix resulting from the matrix multiplication, a matrix blocking strategy for the first matrix and the second matrix is ​​determined.

[0159] Further, based on the above embodiments, the strategy determination module may also be specifically used for:

[0160] The objective function is to take the total time of data loading and type conversion of a single general computing component in each conversion round and the matrix multiplication calculation time of a single special computing component in each calculation round as the consistency;

[0161] Under the objective function, the blocking strategy of the matrix multiplication result matrix is ​​determined according to the data loading speed parameter of the general computing component, the type conversion speed parameter, the calculation speed parameter of the special computing component and the third matrix size.

[0162] On the basis of the above embodiments, the number of the general computing components participating in the type conversion is the same as or different from the number of the special computing components participating in the matrix multiplication calculation; one special computing component obtains the conversion matrix elements required to be calculated from one or more general computing components to perform the matrix multiplication calculation;

[0163] The identification component includes a counter, and the counter includes a plurality of counting units, and one counting unit corresponds to a general calculation component involved in type conversion.

[0164] On the basis of the above embodiments, the type conversion module is specifically used for:

[0165] After the current general computing component completes the type conversion of each matrix element to be converted in the current conversion round, the count value in the counting unit in the counter that matches the current general computing component is updated.

[0166] Based on the above embodiments,

[0167] The matrix multiplication calculation module is specifically used to: determine the target counting unit to be queried and the expected counting value of the target counting unit in the counter according to each conversion matrix element to be calculated in the current calculation round by the current dedicated calculation component;

[0168] querying the target counting unit through the current dedicated computing component, and acquiring a target general-purpose computing component matching the target counting unit when determining that the current count value of the target counting unit is greater than or equal to the expected count value;

[0169] Each conversion matrix element required to be calculated in the current calculation round is obtained in the target general-purpose calculation component through the current dedicated calculation component, and matrix multiplication calculation is performed.

[0170] Further, based on the above embodiments, the pipeline computing device for matrix multiplication may further include: a first waiting module, wherein:

[0171] The first waiting module is used to return to the operation of querying the target counting unit after the target counting unit is queried through the current dedicated computing component, when it is determined through the current dedicated computing component that the current counting value of the target counting unit is less than the expected counting value, after a preset waiting time, until it is determined that the current counting value of the target counting unit is greater than or equal to the expected counting value.

[0172] Based on the above embodiments, the number of the general computing components participating in the type conversion is the same as the number of the special computing components participating in the matrix multiplication calculation; one special computing component only obtains the conversion matrix elements required to be calculated from one general computing component to perform the matrix multiplication calculation;

[0173] The identification component includes a first counter and a second counter, the first counter includes multiple first counting units, one first counting unit corresponds to a general computing component involved in type conversion, and the second counter includes multiple second counting units, one second counting unit corresponds to a special computing component involved in matrix multiplication calculation.

[0174] Further, based on the above embodiments, the type conversion module may be specifically used for:

[0175] After the current general computing component completes the type conversion of each matrix element to be converted in the current conversion round, the count value in the current first counting unit in the first counter that matches the current general computing component is updated;

[0176] After determining that the type conversion of all matrix elements to be converted is completed by the current general computing component, determining an associated general computing component, and determining an associated first counting unit corresponding to the associated general computing component in the first counter;

[0177] Whenever the current general computing component detects that the count value in the associated first counting unit meets the loading condition, each conversion matrix element required to be loaded is obtained from the associated general computing component for local storage, and the count value of the current first counting unit is updated.

[0178] Further, based on the above embodiments, the matrix multiplication calculation module may be specifically used for:

[0179] Determine, by means of a current dedicated computing component, an associated general computing component matching the current dedicated computing component, and obtain, in the first counter, a first current value in a current first counting unit matching the associated general computing component;

[0180] Acquire, by the current dedicated computing component, in the second counter, a second current counting value in a current second counting unit matching the current dedicated computing component;

[0181] When the current dedicated computing component determines that the second current count value is less than the first current value, the conversion matrix elements required to be calculated in the current calculation round are obtained from the associated general computing component, matrix multiplication is performed, and the second current value in the current second counting unit is updated.

[0182] Further, based on the above embodiments, the pipeline computing device for matrix multiplication may further include: a second waiting module, wherein:

[0183] A second waiting module is used for, after obtaining the second current count value in the current second counting unit that matches the current dedicated computing component in the second counter through the current dedicated computing component, when the current dedicated computing component determines that the second current count value is equal to the first current value, after a preset waiting time interval, returning to execute the operation of obtaining the first current value in the current first counting unit that matches the associated general computing component in the first counter until the second current count value is less than the first current value.

[0184] The above-mentioned product can execute the method provided by any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method.

[0185] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0186] Figure 7 Schematic diagram of the structure of an AI chip provided according to an embodiment of the present disclosure. Figure 7As shown, the AI ​​chip includes: a plurality of general computing components 710, a plurality of special computing components 720 and an identification component 730, wherein:

[0187] The general computing component 710 is used to implement the type conversion of the input matrix in the matrix multiplication computing scenario;

[0188] The dedicated computing component 720 is used to implement matrix multiplication calculation on the input matrix in the matrix multiplication calculation scenario;

[0189] The identification component 730 is used to identify the conversion status of each matrix element in the input matrix by the general calculation component 710 in the matrix multiplication calculation scenario;

[0190] Through the cooperation of the general computing component 710 and the special computing component 720, the pipeline computing method of matrix multiplication as described in any embodiment of the present disclosure is jointly implemented.

[0191] Specifically, before performing matrix multiplication, the data type of the input matrix may need to be converted to meet the needs of subsequent calculations. For example, converting an integer type to a floating point type, or converting from one floating point precision to another precision (such as converting from a 32-bit floating point to a 16-bit floating point), the general computing component can realize the type conversion of the input matrix to ensure that the input data meets the calculation requirements, thereby ensuring the correctness and accuracy of the calculation results. The dedicated computing component is a hardware unit designed specifically for matrix multiplication, which usually adopts a parallel computing architecture and can simultaneously process the multiplication and accumulation operations of multiple matrix elements to achieve efficient matrix multiplication operations. During the matrix multiplication calculation process, the identification component is responsible for tracking and recording the type conversion status of each matrix element. For example, it can record how many times a general computing unit in the general computing component has completed the type conversion of the element, or how many times a special computing unit in the special computing component has completed the matrix multiplication operation of the element. By identifying the conversion status and the matrix multiplication operation process, it can be ensured that when performing matrix multiplication, all matrix elements involved in the calculation are in the correct data type state, thereby avoiding calculation errors and data inconsistency problems.

[0192] Among them, the number of each general computing component participating in the type conversion is the same as or different from the number of each special computing component participating in the matrix multiplication calculation; one special computing component obtains each conversion matrix element required to be calculated from one or more general computing components to perform matrix multiplication calculation; the identification component includes a counter, and the counter includes multiple counting units, and one counting unit corresponds to a general computing component participating in the type conversion.

[0193] Alternatively, the number of the general computing components participating in the type conversion is the same as the number of the special computing components participating in the matrix multiplication calculation; one special computing component only obtains the required conversion matrix elements from one general computing component to perform the matrix multiplication calculation; the identification component includes a first counter and a second counter, the first counter includes a plurality of first counting units, one first counting unit corresponds to a general computing component participating in the type conversion, and the second counter includes a plurality of second counting units, one second counting unit corresponds to a special computing component participating in the matrix multiplication calculation.

[0194] Through the cooperation of the general computing component and the special computing component, the pipeline computing method of matrix multiplication as described in any embodiment of the present disclosure is jointly implemented, that is:

[0195] In response to a matrix multiplication request for a first matrix and a second matrix, a plurality of general-purpose computing components and a plurality of special-purpose computing components in the AI ​​chip are triggered in parallel to perform the following operations:

[0196] The general computing component obtains each matrix element to be converted from the first matrix and the second matrix in each conversion round, performs type conversion, obtains each conversion matrix element, and identifies the conversion completion status of each conversion matrix element in the identification component in the AI ​​chip;

[0197] In each calculation round, if the conversion matrix elements required for calculation in the current calculation round are found in the identification component and the conversion is completed, the conversion matrix elements are obtained from the matching general calculation component to perform matrix multiplication calculation.

[0198] The AI ​​chip of the disclosed embodiment includes multiple general-purpose computing components, multiple special-purpose computing components, and an identification component, which can be used to respond to matrix multiplication requests for a first matrix and a second matrix, trigger multiple general-purpose computing components and multiple special-purpose computing components in the AI ​​chip in parallel, perform pipeline calculations of matrix multiplication, and realize parallel execution of the general-purpose computing components and the special-purpose computing components, so that the data conversion time of the general-purpose computing components matches the matrix multiplication calculation time of the special-purpose computing components, and the two are roughly synchronized, thereby making full use of the performance of the special-purpose computing components, significantly improving the speed and efficiency of the matrix multiplication calculation, and being able to efficiently complete the conversion of large-scale data and the matrix multiplication calculation work, thereby improving the performance of the AI ​​chip.

[0199] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0200] Figure 8 Schematic block diagram of an example electronic device that can be used to implement an embodiment of the present disclosure is shown. Figure 8 As shown, the electronic device includes the AI ​​chip 810 as described in any embodiment of the present disclosure.

[0201] Among them, the AI ​​chip 810 may include a general computing component, a special computing component and an identification component. Among them, the general computing component is used to realize the type conversion of the input matrix in the matrix multiplication calculation scenario; the special computing component is used to realize the matrix multiplication calculation of the input matrix in the matrix multiplication calculation scenario; the identification component is used to identify the conversion state of each matrix element in the input matrix by the general computing component in the matrix multiplication calculation scenario. Through the joint cooperation of the general computing component and the special computing component, the pipeline calculation method of matrix multiplication is jointly realized, that is:

[0202] In response to a matrix multiplication request for a first matrix and a second matrix, a plurality of general-purpose computing components and a plurality of special-purpose computing components in the AI ​​chip are triggered in parallel to perform the following operations:

[0203] The general computing component obtains each matrix element to be converted from the first matrix and the second matrix in each conversion round, performs type conversion, obtains each conversion matrix element, and identifies the conversion completion status of each conversion matrix element in the identification component in the AI ​​chip;

[0204] In each calculation round, if the conversion matrix elements required for calculation in the current calculation round are found in the identification component and the conversion is completed, the conversion matrix elements are obtained from the matching general calculation component to perform matrix multiplication calculation.

[0205] In some embodiments, the pipeline computing method for matrix multiplication may be implemented as a computer program product tangibly embodied in a computer-readable storage medium, such as a memory unit.

[0206] The computer program for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These computer programs can be provided to the general computing components and the dedicated computing components in the AI ​​chip to perform the pipeline computing method of matrix multiplication.

[0207] In the context of the present disclosure, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by general-purpose computing components and special-purpose computing components in an AI chip. Computer-readable storage media may include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media may include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0208] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0209] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions provided by this disclosure can be achieved, and this document does not limit this.

[0210] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A pipeline calculation method for matrix multiplication, executed by an AI chip, comprising: In response to a matrix multiplication request for a first matrix and a second matrix, a plurality of general-purpose computing components and a plurality of special-purpose computing components in the AI ​​chip are triggered in parallel to perform the following operations: The general computing component obtains each matrix element to be converted from the first matrix and the second matrix in each conversion round, performs type conversion, obtains each conversion matrix element, and identifies the conversion completion status of each conversion matrix element in the identification component in the AI ​​chip; In each calculation round, if the conversion matrix elements required for calculation in the current calculation round are found in the identification component and the conversion is completed, the conversion matrix elements are obtained from the matching general calculation component to perform matrix multiplication calculation.

2. The method according to claim 1, before the parallel triggering of the plurality of general computing components and the plurality of special computing components in the AI ​​chip, further comprises: In the matrix multiplication calculation request, obtaining a first matrix size of the first matrix and a second matrix size of the second matrix; Determining a matrix block strategy for the first matrix and the second matrix according to the first matrix size, the second matrix size, and the performance parameters of the AI ​​chip; According to the matrix partitioning strategy and the number of the general computing components involved in the type conversion, determining the matrix elements that each general computing component needs to convert in each conversion round; According to the matrix partitioning strategy and the number of the dedicated computing components participating in the matrix multiplication calculation, the transformation matrix elements that each dedicated computing component needs to calculate in each calculation round are determined.

3. The method according to claim 2, wherein: The determining, according to the first matrix size, the second matrix size, and the performance parameters of the AI ​​chip, a matrix partitioning strategy for the first matrix and the second matrix includes: Determining a third matrix size of a matrix multiplication result matrix according to the first matrix size and the second matrix size; Determining a block strategy for the matrix multiplication result matrix according to a data loading speed parameter of the general computing component, a type conversion speed parameter, a calculation speed parameter of the special computing component, and a size of the third matrix; According to the matrix blocking strategy of the matrix multiplication result, a matrix blocking strategy for the first matrix and the second matrix is ​​determined.

4. The method according to claim 3, wherein: Determining the block strategy of the matrix multiplication result matrix according to the data loading speed parameter of the general computing component, the type conversion speed parameter, the calculation speed parameter of the special computing component and the size of the third matrix specifically includes: The objective function is to take the total time of data loading and type conversion of a single general computing component in each conversion round and the matrix multiplication calculation time of a single special computing component in each calculation round as the consistency; Under the objective function, the blocking strategy of the matrix multiplication result matrix is ​​determined according to the data loading speed parameter of the general computing component, the type conversion speed parameter, the calculation speed parameter of the special computing component and the third matrix size.

5. The method according to any one of claims 2 to 4, wherein: The number of the general computing components participating in the type conversion is the same as or different from the number of the special computing components participating in the matrix multiplication calculation; one special computing component obtains each conversion matrix element required to be calculated from one or more general computing components to perform the matrix multiplication calculation; The identification component includes a counter, and the counter includes a plurality of counting units, and one counting unit corresponds to a general calculation component involved in type conversion.

6. The method according to claim 5, wherein: The general computing component obtains each matrix element to be converted from the first matrix and the second matrix in each conversion round, performs type conversion, obtains each conversion matrix element, and identifies the conversion completion state of each conversion matrix element in the identification component in the AI ​​chip, including: After the current general computing component completes the type conversion of each matrix element to be converted in the current conversion round, the count value in the counting unit in the counter that matches the current general computing component is updated.

7. The method according to claim 6, characterized in that The dedicated computing component, in each computing round, if the conversion matrix elements required to be calculated in the current computing round are found in the identification component to complete the conversion, then the conversion matrix elements are obtained from the matching general computing component to perform matrix multiplication calculation, including: Determine the target counting unit to be queried and the expected counting value of the target counting unit in the counter according to each conversion matrix element to be calculated in the current calculation round by the current dedicated calculation component; querying the target counting unit through the current dedicated computing component, and acquiring a target general computing component matching the target counting unit when determining that the current count value of the target counting unit is greater than or equal to the expected count value; Each conversion matrix element required to be calculated in the current calculation round is obtained in the target general-purpose calculation component through the current dedicated calculation component, and matrix multiplication calculation is performed.

8. The method according to claim 7, after querying the target counting unit through the current dedicated computing component, further comprises: When the current dedicated computing component determines that the current count value of the target counting unit is less than the expected count value, it returns to query the target counting unit after a preset waiting time until it is determined that the current count value of the target counting unit is greater than or equal to the expected count value.

9. The method according to any one of claims 2 to 4, wherein: The number of the general computing components participating in the type conversion is the same as the number of the special computing components participating in the matrix multiplication calculation; one special computing component only obtains the conversion matrix elements required for calculation from one general computing component to perform the matrix multiplication calculation; The identification component includes a first counter and a second counter, the first counter includes multiple first counting units, one first counting unit corresponds to a general computing component involved in type conversion, and the second counter includes multiple second counting units, one second counting unit corresponds to a special computing component involved in matrix multiplication calculation.

10. The method according to claim 9, characterized in that The general computing component obtains each matrix element to be converted from the first matrix and the second matrix in each conversion round, performs type conversion, obtains each conversion matrix element, and identifies the conversion completion state of each conversion matrix element in the identification component in the AI ​​chip, including: After the current general computing component completes the type conversion of each matrix element to be converted in the current conversion round, the count value in the current first counting unit in the first counter that matches the current general computing component is updated; After determining that type conversion of all matrix elements to be converted is completed by the current general-purpose computing component, determining an associated general-purpose computing component, and determining an associated first counting unit corresponding to the associated general-purpose computing component in the first counter; Whenever the current general computing component detects that the count value in the associated first counting unit meets the loading condition, each conversion matrix element required to be loaded is obtained from the associated general computing component for local storage, and the count value of the current first counting unit is updated.

11. The method according to claim 10, wherein: The dedicated computing component, in each computing round, if the conversion matrix elements required to be calculated in the current computing round are found in the identification component to complete the conversion, then the conversion matrix elements are obtained from the matching general computing component to perform matrix multiplication calculation, including: Determine, by means of a current dedicated computing component, an associated general computing component matching the current dedicated computing component, and obtain, in the first counter, a first current value in a current first counting unit matching the associated general computing component; Acquire, by the current dedicated computing component, in the second counter, a second current counting value in a current second counting unit matching the current dedicated computing component; When the current dedicated computing component determines that the second current count value is less than the first current value, the conversion matrix elements required to be calculated in the current calculation round are obtained from the associated general computing component, matrix multiplication is performed, and the second current value in the current second counting unit is updated.

12. The method according to claim 11, characterized in that After acquiring, in the second counter, a second current count value in a current second counting unit matching the current dedicated computing component, the method further includes: When the current dedicated computing component determines that the second current count value is equal to the first current value, after a preset waiting period, it returns to execute the operation of obtaining the first current value in the current first counting unit that matches the associated general computing component in the first counter until the second current count value is less than the first current value.

13. A pipeline computing device for matrix multiplication, comprising: A parallel trigger module, configured to trigger, in response to a matrix multiplication request for the first matrix and the second matrix, a plurality of general computing components and a plurality of special computing components in the AI ​​chip in parallel to execute the following modules: A type conversion module, configured to obtain each matrix element to be converted from the first matrix and the second matrix in each conversion round through the general computing component to perform type conversion, obtain each conversion matrix element, and identify the conversion completion status of each conversion matrix element in the identification component in the AI ​​chip; The matrix multiplication calculation module is used to obtain each conversion matrix element from the matching general computing component to perform matrix multiplication calculation in each calculation round if the conversion matrix elements required for the current calculation round are completed in the identification component through the dedicated computing component.

14. The apparatus according to claim 13, further comprising: a size acquisition module, configured to acquire, in the matrix multiplication calculation request, a first matrix size of the first matrix and a second matrix size of the second matrix before the parallel triggering of the plurality of general computing components and the plurality of special computing components in the AI ​​chip; A strategy determination module, configured to determine a matrix partitioning strategy for the first matrix and the second matrix according to the first matrix size, the second matrix size, and the performance parameters of the AI ​​chip; A conversion determination module, used to determine the matrix elements that each general computing component needs to convert in each conversion round according to the matrix block strategy and the number of each general computing component participating in the type conversion; The calculation determination module is used to determine the transformation matrix elements that each dedicated computing component needs to calculate in each calculation round according to the matrix blocking strategy and the number of each dedicated computing component participating in the matrix multiplication calculation.

15. The device according to claim 14, wherein: The strategy determination module is specifically used to: Determining a third matrix size of a matrix multiplication result matrix according to the first matrix size and the second matrix size; Determining a block strategy for the matrix multiplication result matrix according to a data loading speed parameter of the general computing component, a type conversion speed parameter, a calculation speed parameter of the special computing component, and a size of the third matrix; According to the matrix blocking strategy of the matrix multiplication result, a matrix blocking strategy for the first matrix and the second matrix is ​​determined.

16. The device according to claim 15, wherein: The strategy determination module is further specifically used for: The objective function is to take the total time of data loading and type conversion of a single general computing component in each conversion round and the matrix multiplication calculation time of a single special computing component in each calculation round as the consistency; Under the objective function, the blocking strategy of the matrix multiplication result matrix is ​​determined according to the data loading speed parameter of the general computing component, the type conversion speed parameter, the calculation speed parameter of the special computing component and the third matrix size.

17. The device according to claims 14-16, wherein: The number of the general computing components participating in the type conversion is the same as or different from the number of the special computing components participating in the matrix multiplication calculation; one special computing component obtains each conversion matrix element required to be calculated from one or more general computing components to perform the matrix multiplication calculation; The identification component includes a counter, and the counter includes a plurality of counting units, and one counting unit corresponds to a general calculation component involved in type conversion.

18. The device according to claim 17, wherein: Type conversion module, specifically used for: After the current general computing component completes the type conversion of each matrix element to be converted in the current conversion round, the count value in the counting unit in the counter that matches the current general computing component is updated.

19. The device according to claim 18, wherein: Matrix multiplication calculation module, specifically used for: Determine the target counting unit to be queried and the expected counting value of the target counting unit in the counter according to each conversion matrix element to be calculated in the current calculation round by the current dedicated calculation component; querying the target counting unit through the current dedicated computing component, and acquiring a target general computing component matching the target counting unit when determining that the current count value of the target counting unit is greater than or equal to the expected count value; Each conversion matrix element required to be calculated in the current calculation round is obtained in the target general-purpose calculation component through the current dedicated calculation component, and matrix multiplication calculation is performed.

20. The apparatus according to claim 19, further comprising: The first waiting module is used to return to the operation of querying the target counting unit after the target counting unit is queried through the current dedicated computing component, when it is determined through the current dedicated computing component that the current counting value of the target counting unit is less than the expected counting value, after a preset waiting time, until it is determined that the current counting value of the target counting unit is greater than or equal to the expected counting value.

21. The device according to claims 14-16, wherein: The number of the general computing components participating in the type conversion is the same as the number of the special computing components participating in the matrix multiplication calculation; one special computing component only obtains the conversion matrix elements required for calculation from one general computing component to perform the matrix multiplication calculation; The identification component includes a first counter and a second counter, the first counter includes multiple first counting units, one first counting unit corresponds to a general computing component involved in type conversion, and the second counter includes multiple second counting units, one second counting unit corresponds to a special computing component involved in matrix multiplication calculation.

22. The device according to claim 21, wherein The type conversion module is further specifically used for: After the current general computing component completes the type conversion of each matrix element to be converted in the current conversion round, the count value in the current first counting unit in the first counter that matches the current general computing component is updated; After determining that type conversion of all matrix elements to be converted is completed by the current general-purpose computing component, determining an associated general-purpose computing component, and determining an associated first counting unit corresponding to the associated general-purpose computing component in the first counter; Whenever the current general computing component detects that the count value in the associated first counting unit meets the loading condition, each conversion matrix element required to be loaded is obtained from the associated general computing component for local storage, and the count value of the current first counting unit is updated.

23. The device according to claim 22, wherein: The matrix multiplication calculation module is further specifically used for: Determine, by means of a current dedicated computing component, an associated general computing component matching the current dedicated computing component, and obtain, in the first counter, a first current value in a current first counting unit matching the associated general computing component; Acquire, by the current dedicated computing component, in the second counter, a second current counting value in a current second counting unit matching the current dedicated computing component; When the current dedicated computing component determines that the second current count value is less than the first current value, the conversion matrix elements required to be calculated in the current calculation round are obtained from the associated general computing component, matrix multiplication is performed, and the second current value in the current second counting unit is updated.

24. The apparatus according to claim 23, further comprising: A second waiting module is used for, after obtaining the second current count value in the current second counting unit that matches the current dedicated computing component in the second counter through the current dedicated computing component, when the current dedicated computing component determines that the second current count value is equal to the first current value, after a preset waiting time interval, returning to execute the operation of obtaining the first current value in the current first counting unit that matches the associated general computing component in the first counter until the second current count value is less than the first current value.

25. An AI chip, characterized in that: include: A plurality of general computing components, a plurality of special computing components and an identification component; The general computing component is used to implement type conversion of input matrices in matrix multiplication computing scenarios; The dedicated computing component is used to implement matrix multiplication calculation on the input matrix in the matrix multiplication calculation scenario; The identification component is used to identify the conversion status of each matrix element in the input matrix by the general computing component in the matrix multiplication calculation scenario; The method according to any one of claims 1 to 12 is implemented by the cooperation of the general computing component and the special computing component.

26. An electronic device, characterized in that: Comprising the AI ​​chip as described in claim 25.

27. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable the general computing component and the special computing component in the AI ​​chip to execute the method according to any one of claims 1-12.

28. A computer program product, comprising a computer program, wherein when the computer program is executed by a general-purpose computing component and a special-purpose computing component in an AI chip, the steps of the method described in any one of claims 1 to 12 are implemented.

Citation Information

Cited By

  • Fine-grained quantization matrix multiplication device and method based on systolic array

    CN121479115A

  • Fine-grained quantized matrix multiplication apparatus and method based on systolic array

    CN121479115B